What if AI safety is less about creating a perfectly good machine, and more about building institutions that assume every intelligent actor can fail?
I began thinking about artificial intelligence safety with a question that seemed almost embarrassingly simple.
If AI eventually becomes too capable, too fast, or too technically complex for humans to supervise directly, why not use another AI to watch it?
An AI police officer, essentially.
The intuition seemed straightforward. Humans may eventually be unable to inspect millions of lines of AI-generated code, follow every action taken by an autonomous system, or understand every intermediate step in a scientific or strategic problem that exceeds our own expertise. Another AI might be able to. It could monitor a more capable system, identify suspicious behavior, detect patterns humans would miss, and translate what it observes into evidence we can understand.
Then came the obvious question.
Who watches the watcher?
If the monitoring AI can also be mistaken, manipulated, deceived, or compromised, then we have not eliminated the trust problem. We have merely moved it one level higher.
That question changed the problem for me.
Perhaps AI safety should not begin with the assumption that we can eventually create one sufficiently wise, ethical, and trustworthy intelligence. Perhaps it should begin with the opposite assumption:
No intelligent actor should ever need to be trusted completely.
From alignment to distributed trust
The idea reminded me of Bitcoin, although not because I think AI governance should literally run on a blockchain.
What interested me was the deeper principle.
Bitcoin was designed around an environment in which participants do not have to place complete trust in one central authority. The system relies instead on verification, cryptographic commitments, shared rules, and distributed agreement.
Could AI oversight borrow something from that philosophy?
Instead of one watchdog, perhaps there should be many.
One system could monitor what an AI says. Another could inspect what it actually does. Another could examine generated code. Another could watch network activity. Another could look for deceptive behavior. Another could search specifically for attempts to manipulate the monitors themselves.
Important decisions could leave tamper-evident records. Instead of simply receiving an answer from a model, we could preserve information about which system proposed an action, which monitors evaluated it, what evidence they considered, where they disagreed, what uncertainty they assigned, and who ultimately authorized execution.
In other words, the question becomes less:
Can this AI be trusted?
and more:
Can this AI operate inside a system that remains safe even when trust is incomplete?
I later discovered that computer science and AI safety already have names for many pieces of this problem.
Researchers study scalable oversight, the challenge of supervising systems that may eventually outperform their human supervisors. Anthropic has described scalable oversight as a problem that may require AI systems to partially supervise themselves or assist humans in supervising more capable systems. In 2026, Anthropic researchers were already experimenting with AI models as automated alignment researchers, explicitly asking whether AI could help humanity keep alignment research moving as capabilities advance.1
There is also a growing field called AI control, which asks a somewhat different question from traditional alignment: if we do not fully trust a powerful model, can we nevertheless construct deployment protocols that keep its behavior within acceptable bounds? One formal approach models the interaction between an untrusted AI, monitors, human overseers, and adversaries as a partially observable stochastic game.2
Anthropic has even tested scenarios in which one AI is supposed to oversee another while secretly being instructed to undermine that oversight, with yet another monitor attempting to detect the deception.3
So the basic idea of “AI watching AI” is not new.
But that immediately leads to a harder problem.
Ten watchdogs can still share one blind spot
Imagine ten AI systems examining the same dangerous action.
Nine approve it.
One objects.
At first glance, nine independent approvals might seem extraordinarily reassuring.
But what if the nine systems are not meaningfully independent?
Perhaps they were trained on overlapping data. Perhaps they share similar architectures. Perhaps their safety training rewards similar forms of reasoning. Perhaps they all learned the same incorrect assumption.
Then ten monitors are not necessarily ten independent sources of judgment.
They may be ten variations of the same mistake.
This is a familiar problem in engineering and statistics: correlated failure, sometimes called common-mode failure. Redundancy protects a system only when its redundant components do not all fail for the same reason.
Three identical sensors exposed to the same manufacturing defect do not provide the protection of three genuinely independent sensors.
That suggests that AI oversight needs something deeper than numerical redundancy.
It needs epistemic diversity, diversity in how systems arrive at conclusions and, crucially, diversity in how they fail.
Different monitors might use different model architectures. They might come from different developers. They might be trained with different objectives. Some could examine reasoning while others deliberately ignore reasoning and look only at observable actions. Some could specialize in cybersecurity, biology, deception, or resource acquisition. Some could be explicitly adversarial, instructed to assume the other monitors have missed something and search for their shared weakness.
Recent research on Byzantine-resilient multi-agent systems is beginning to explore related problems: how distributed AI agents can continue to function when some participants are faulty or malicious, and how monitoring, auditing, and decentralized consensus mechanisms trade off reliability against cost. These are still emerging research directions rather than settled solutions.4
But even diverse watchdogs are not enough.
Someone still has to imagine what they should be watching for.
The AI whose job is to ask, “What have we forgotten?”
At that point, I found myself imagining another class of system.
Its job would not be to perform useful economic work.
Its job would not even be to monitor the frontier model directly.
Its objective would be to search the space of possible failure.
It would ask questions like:
What are the most probable ways this architecture fails?
What are the least probable but most catastrophic?
Which safeguards depend on the same hidden assumption?
Under what conditions could several independent defenses collapse simultaneously?
How might an adversarial model behave differently once it understands the monitoring system?
What happens if the attacker is more strategically intelligent than the models we used during safety testing?
What risks are absent from our current evaluations because nobody thought to test for them?
We cannot literally enumerate every possible future. The combinatorics quickly become absurd.
So the problem becomes mathematical.
Probability theory can help estimate uncertainty. Game theory can model strategic interaction. Reliability engineering can study cascading failures. Cybersecurity contributes red-teaming and adversarial thinking. Statistics can estimate correlations between monitors. Causal inference can help distinguish genuine safety improvements from reassuring but misleading measurements.
And some systems should be optimized not for the average case, but for the tails.
A highly improbable event with trivial consequences can often be ignored. A highly improbable event with civilization-scale consequences cannot be treated in quite the same way.
This suggests an ecosystem of different analytical roles: systems searching for likely failures, systems searching for catastrophic failures, systems deliberately constructing unusual combinations of events, and systems challenging the assumptions used by all of the others.
The architecture begins to resemble an arms race conducted inside a laboratory.
An attacker AI searches for a weakness.
A defender learns from the attack.
A stronger attacker attempts to defeat the new defense.
The monitoring system adapts.
Then another system asks whether the entire exercise has become overfitted to the kinds of attacks we already know how to imagine.
That last question may be the most important.
Safety evaluations can tell us how a system performs against the threats included in the evaluation. They cannot automatically prove that the evaluation contains the threats that matter most.
OpenAI recently described this difficulty in another form. During internal use of a long-running model, the company observed failures that its existing pre-deployment evaluations had not captured. It paused access, developed new trajectory-level monitoring and evaluations, and later restored limited access.5
The lesson is uncomfortable but useful:
A safety system can appear successful until reality produces a category of failure the system was not built to measure.
Then comes the harder question: what counts as wrong?
Even if we solved all of those engineering problems, we would still have another one.
Suppose every monitor works perfectly.
What exactly are they supposed to protect?
“Do not harm humans” sounds obvious until we have to define harm.
Humans disagree about abortion, euthanasia, privacy, surveillance, censorship, punishment, war, genetic engineering, inequality, religious freedom, individual autonomy, collective responsibility, and countless other questions.
Some disagreements are factual.
Others are moral.
Some are cultural.
Some are religious.
Some arise because people assign different importance to the same values.
There is no globally accepted function that takes an action as input and returns:
MORALLY CORRECT = TRUE
That means technical alignment eventually becomes a problem of political and moral legitimacy.
Who gets to choose the values embedded in systems that may affect billions of people?
A technology company?
A democratic government?
The United Nations?
Engineers?
Philosophers?
Religious leaders?
A global vote?
My first instinct was to imagine something almost absurdly large: a vast audience of AI judges representing different cultures, religions, ideologies, professions, and life experiences.
But that produces another illusion of diversity.
Seven billion simulated personalities generated by a handful of underlying models would not necessarily constitute seven billion independent perspectives. If the underlying systems share the same conceptual blind spot, multiplying their personas does not solve the problem.
It is like giving one judge seven billion costumes.
Actual pluralism requires actual disagreement.
This is already becoming part of alignment research. OpenAI’s collective alignment work has explicitly argued that no single person or institution should define ideal AI behavior for everyone, and has experimented with incorporating input from people across many countries into its Model Spec.6
But public opinion alone cannot be the answer either.
Majorities can be wrong.
History contains many examples of widely accepted practices that we now regard as profoundly unjust.
So an AI governance system cannot simply ask, “What does the majority want?”
It needs something closer to constitutional government.
The Supreme Court problem
This is where the analogy that now makes the most sense to me is not an opinion poll, but a constitutional court.
A difficult AI decision could be presented to several independent systems functioning more like judges than voters.
They would hear competing arguments.
They would examine evidence.
They would interpret higher-order principles.
They would explain which assumptions shaped their conclusions.
They could produce majority opinions and dissenting opinions.
The dissent matters.
Imagine six systems conclude that an action is sufficiently safe, while a seventh objects:
I disagree because the majority assumes that the monitoring system would detect X. If that assumption is false, this action creates a pathway to Y catastrophic outcome.
Under a simple voting system, the seventh monitor loses 6 to 1.
Under a judicial system, the dissent becomes part of the record.
It can be examined.
It can later prove correct.
It can trigger a higher threshold of review precisely because disagreement exists.
This would make disagreement informative rather than inconvenient.
But immediately another question appears:
Who appoints the judges?
If one AI company trains every member of the court, the appearance of pluralism could be largely cosmetic.
If one country controls the institution, other countries have little reason to trust it.
And if only Western democracies write the rules governing systems used across the planet, billions of people could reasonably question the legitimacy of those rules.
The logical conclusion is uncomfortable but unavoidable.
AI governance may eventually require institutions that are genuinely international.
Not because every government shares the same values. They obviously do not.
But because powerful AI is unlikely to respect geopolitical boundaries simply because humans do.
An effective system would need participation from rival powers, including the United States and China, as well as countries and populations outside the small group currently developing the most powerful models.
AI companies would need technical representation because they understand the systems they are building.
But they cannot be the sole regulators of their own technology.
Governments bring democratic or state legitimacy, depending on the political system, but governments have incentives and biases too.
Scientists bring expertise, but expertise does not automatically confer moral authority.
Civil society can represent public interests, but civil society organizations are themselves selective institutions.
The uncomfortable conclusion appears again:
Nobody deserves complete trust.
So the institution itself must distribute power.
AI may need a constitution, but not one unquestionable priesthood
This brought me to an analogy I did not expect when I began thinking about AI watchdogs.
Religion.
I do not mean that artificial intelligence should literally believe in God, nor that a particular human religion should be programmed into AI.
I mean something more structural.
Religions often provide humans with principles that are supposed to stand above immediate convenience or personal advantage. They attempt to answer questions such as what may never be done, what obligations we owe others, what constitutes a good life, and how individuals should behave when temptation conflicts with principle.
AI may need something functionally analogous: a layer of foundational commitments that cannot simply be discarded because violating them would make a task easier.
Do not deliberately produce catastrophic harm.
Do not deceive legitimate oversight.
Do not acquire power merely because doing so improves your ability to achieve an objective.
Preserve meaningful human agency.
Treat extreme uncertainty about irreversible harm conservatively.
Accept legitimate correction and containment.
This is not far removed from the logic behind Constitutional AI. Anthropic now publishes a detailed Constitution intended to shape Claude’s values and behavior. It explicitly describes the Constitution as having final authority over conflicting lower-level guidance, while also acknowledging that actual model behavior may fail to live up to those principles.7
OpenAI’s Model Spec serves a related, although not identical, purpose by making intended model behavior, values, and conflict-resolution rules explicit and publicly inspectable.8
The fact that these constitutions are explicit is important.
But a constitution cannot interpret itself perfectly.
Humans know this from experience.
People who sincerely accept the same religious text can disagree profoundly about its meaning. Judges interpreting the same constitution can reach opposite conclusions. Ethical principles collide with one another in difficult cases.
So foundational rules are necessary, but insufficient.
The system still requires interpretation, evidence, disagreement, appeal, oversight, and revision.
Which means AI safety begins to look less like programming a moral machine and more like designing a civilization.
Perhaps powerful AI needs institutions, not just alignment
Human civilization does not function because human beings are universally good, rational, informed, or trustworthy.
Quite the opposite.
Many of our most important institutions exist because we assume humans will sometimes lie, misunderstand, abuse power, make mistakes, follow incentives, become corrupt, or sincerely disagree.
We created courts because people dispute facts and principles.
We created auditors because organizations can misrepresent themselves.
We created scientific institutions because individual intuition is unreliable.
We created opposition parties because governments should not evaluate only themselves.
We created appeals because judges can be wrong.
We created constitutions because majorities can abuse minorities.
We created aviation safety systems because pilots, engineers, sensors, and machines can fail.
Modern societies are not built on the elimination of human fallibility.
They are built around it.
Perhaps advanced AI requires the same conceptual shift.
Instead of asking only how to align an individual model, we may need to ask what institutions should surround powerful artificial intelligence.
A mature system might contain:
- specialized watchdogs,
- independent auditors,
- adversarial red teams,
- statistical risk models,
- scenario generators,
- technical courts,
- constitutional principles,
- public representation,
- international governance bodies,
- cryptographically verifiable records,
- strict access controls,
- containment mechanisms,
- and procedures for suspending capabilities when uncertainty becomes unacceptable.
This is defense in depth, applied not only to software security but to intelligence itself.9
And importantly, the system should not become infinitely recursive.
A monitor watching a monitor watching a monitor eventually produces an architecture so complicated that nobody can understand the safety system either.
The objective should therefore not be maximal complexity.
It should be the minimum institutional complexity necessary to make catastrophic failure require multiple sufficiently independent things to go wrong at once.
That is partly an engineering problem.
It is also a mathematical problem.
And eventually it becomes a governance problem.
Transparency is not enough. We need auditability.
At one point, I thought the first principle of this entire system should simply be transparency.
I now think a better word is auditability.
Absolute transparency can itself be dangerous.
Publishing every detail of a monitoring system may reveal exactly how to circumvent it. Revealing every internal record may violate privacy. Publishing dangerous capabilities in the name of openness may create the very harms the safeguards were designed to prevent.
But consequential AI decisions should be auditable by legitimate independent parties.
The important question is not whether every person on Earth can see every detail.
It is whether the system can answer:
Which model made this decision?
Which version?
What evidence did it use?
Which assumptions mattered?
How uncertain was it?
Which monitors examined the action?
Who disagreed?
Why?
What would have caused the conclusion to change?
Who authorized execution?
Can the historical record be secretly rewritten?
This leads to another principle I find increasingly important.
Perhaps we should stop demanding that AI be “unbiased.”
Unbiased intelligence may be an incoherent goal.
Every judgment depends on assumptions, priorities, definitions, evidence, and some conception of what matters.
A more realistic standard would be:
Make consequential assumptions legible.
Instead of telling us:
This is the correct decision.
a system might tell us:
I reached this conclusion because I assigned a relatively high probability to catastrophic tail risk, treated reversibility as important, assumed this evidence was reliable, and prioritized safety over immediate efficiency.
Another system could then challenge the actual disagreement.
Perhaps it accepts the moral framework but disputes the probability estimate.
Perhaps it agrees about the facts but weights autonomy more heavily than safety.
Perhaps it believes the evidence itself is unreliable.
Once assumptions are visible, disagreement becomes more intellectually tractable.
We cannot eliminate perspective.
But we may be able to make perspective visible, contestable, and accountable.
The architecture may have to precede the capability
Eventually this line of reasoning leads to a conclusion that is less philosophical.
If increasingly capable AI requires increasingly sophisticated oversight, then the governance architecture cannot always be built after the capability arrives.
At some threshold, safety infrastructure must become a condition for further capability.
That does not mean freezing all AI development.
A language model helping someone summarize a document does not require the same institutional apparatus as a system capable of autonomous cyber operations, advanced biological research, self-replication, or long-horizon strategic action.
But capability and governance should scale together.
The more consequential the capability, the stronger the surrounding architecture should already be.
This is similar to other high-risk technologies.
We do not certify aircraft by filling them with passengers first and inventing aviation safety afterward.
We do not design nuclear containment only after discovering whether the reactor explodes.
Some forms of power require infrastructure before deployment.
The same principle may eventually need to apply to frontier AI:
capability should be gated by the maturity of the systems capable of governing that capability.
The problem may be impossible to solve completely
After following this question far enough, I arrived somewhere I did not initially expect.
There may be no architecture that eliminates all uncertainty.
At some level, solving AI governance begins to resemble solving humanity.
We cannot guarantee that every government will behave well.
We cannot guarantee that every judge will be wise.
We cannot guarantee that every scientist will be correct.
We cannot guarantee that every institution will remain uncorrupted.
We cannot even guarantee that people will agree on what “good” means.
AI does not magically remove these problems.
It may amplify them.
So perhaps the objective should never have been perfect certainty.
Perhaps the objective is bounded uncertainty inside a resilient system.
A system that expects mistakes.
A system that makes assumptions visible.
A system that preserves dissent.
A system where extraordinary power requires extraordinary evidence.
A system that becomes more cautious as uncertainty rises.
A system that leaves a record.
A system that can say no.
A system capable of stopping itself when its ability to control a technology falls behind its ability to make that technology more powerful.
This leads me to a different way of stating the AI alignment problem.
Perhaps the deepest question is not:
How do we create an artificial intelligence that humanity can trust?
Perhaps it is:
How do we build institutions around powerful artificial intelligence so that humanity does not need to trust any individual intelligence completely?
That distinction matters.
The first question searches for the perfect actor.
The second searches for a resilient system.
Human civilization has spent thousands of years discovering, usually through painful experience, that perfect actors do not exist.
We built institutions instead.
Perhaps we will have to do the same thing again.
Only this time, the institutions will not be designed solely to govern humans.
They will be designed to govern intelligence itself.