Start with the disagreement
There is no consensus on how dangerous today's AI really is. Some argue that capabilities are advancing faster than anyone's ability to understand or control them, and that catastrophe is a real near-term possibility. Others argue the opposite: that these systems are overstated, brittle and oversold, that warnings of imminent superhuman AI are unfounded, and that the alarm is partly a bid for public money and protection from competition.
The safety decision, however, does not have to wait for that argument to be settled. A system designed so that its worst credible action cannot cause serious harm is the right design under either view. If the pessimists are right, hard limits reduce how much damage a highly capable system can do. If the skeptics are right, the same limits stop unreliable, overpromoted systems from being handed authority they have not earned. Designing for bounded consequences increases safety, resilience and accountability regardless of who is correct. That is the case this statement makes.
The common weakness in today's proposals
The leading responses to AI risk are to slow development, embed outside evaluators, adopt shared standards and coordinate across companies and countries. Each is worth doing. But each rests on the same hidden assumption: that safety can be achieved by getting some actor — the developer, the reviewer, the standards body, the coalition — to behave correctly or to catch the problem in time. None of them changes what the system is actually able to cause. They also rely on the unlikely prospect that all actors will fulfill their commitments and obligations – something that the record does not support.
Slowing down buys calendar time, not understanding or necessarily risk reduction. The binding limit is not the schedule; it is that no person can fully review a complex, changing system faster than it acts — and it is the same overloaded people doing the reviewing. A slower program can simply deliver the same flawed system later.
Outside evaluators improve visibility, but aviation learned long ago that bringing in reviewers has limits: the system is too complex for any outsider to keep pace, and an embedded reviewer tends to inherit the same information, assumptions and blind spots as the team under review. A reviewer is a warning system, not a barrier.
Common standards can spread a common mistake. If a shared rule carries a flawed assumption or misses a hazard, every company that follows it faithfully reproduces the same unsafe condition. Different company names are not independence when everyone shares the same methods.
Coordination cannot bind the actor most likely to cause trouble — a competitor, a new entrant or another country that declines to take part. Cooperation must be treated as something that can be incomplete or withdrawn, not assumed.
Control what you can: design out the loss
The alternative is to stop trying to guarantee good behavior and instead design the larger system so that foreseeable unsafe decisions — by an AI, a person or an organization — cannot become unacceptable losses. This is an established idea in system safety, and in the systems-theoretic security approach developed by William Young and Nancy Leveson: define the losses that must never happen, trace every path that could lead to them, and place enforceable limits where a digital decision becomes a real-world effect. Assume the worst credible action succeeds, and ensure that even then an independent control prevents or contains the loss.
It is the same principle that protects a building from an electrical fault. No one promises the wiring will never fail. A circuit breaker sits where a fault would become a fire and cuts the power regardless of the cause. For AI, that means placing independent limits at the points where an AI's output becomes consequential — moving money, changing systems, issuing instructions, acting through other software or through people — so a wrong decision is stopped before it does harm.
A test to apply before giving any AI more authority
For every serious loss, ask one question: if the AI acts wrongly, the evaluator misses it, the governing standard has a gap and the deploying organization proceeds anyway, what still prevents the loss? If the only answer is that someone was supposed to notice, or that a company promised not to proceed, the loss is not yet controlled. A real answer names a protection that acts before the loss, sits at the boundary where an AI output becomes an effect, and still works in the situation the unsafe action has created.
That protection must be independent in more than one sense: the AI cannot alter or bypass it; the commercial organization cannot quietly waive it; it does not simply repeat the same model, data or assumptions that might be wrong together; and it acts before the hazard outruns the response.
If everyone else misses it, what still prevents the loss?
What this looks like in practice
Controlling consequences means:
- Separating capability from authority, so a system can recommend or plan without unilateral power to execute high-consequence actions;
- Bounding what the AI and the organization can reach, change and cause, including combined and indirect effects;
- Enforcing decisive limits through components the AI cannot modify and the release chain cannot casually waive;
- Preserving cancellation, rollback, containment and recovery that still work after an unsafe action has begun;
- Giving human overseers the information, time, authority and independent evidence to actually intervene; and
- Narrowing authority automatically when the system, its environment or the supporting evidence changes.
An honest boundary
No single organization can stop every company or country from building or deploying AI. The honest task is to control the interfaces through which AI gains consequential authority — credentials, networks, transaction systems, data-release paths, physical equipment, operating permissions and protected infrastructure — so that a system outside any cooperation still cannot use those pathways to cause an unacceptable loss. Where an effect can bypass every controlled interface, that limitation should be stated plainly and the design reconsidered, rather than resting on an unstated assumption that everyone will cooperate.
This is not a promise of zero risk, and not a claim that any AI is safe in the abstract. It is a claim about engineering: the safety of a consequential system should rest on what it can be allowed to cause, verified for the actual deployment — not on the hope that its builders, reviewers, standards and competitors all behave as intended.
Related reading
William Young and Nancy G. Leveson, “An Integrated Approach to Safety and Security Based on Systems Theory,” Communications of the ACM 57(2), 31–35 (2014).
About Endrisk
Endrisk develops collaborative system-safety software for complex, software-intensive systems. Its platform supports System-Theoretic Process Analysis and Causal Analysis based on System Theory and connects hazards, safety constraints, requirements, assumptions, evidence and change in one defensible record. Endrisk's mission is to modernize system safety so organizations can maintain control of consequential systems.
For further details, contact Endrisk: contact@endrisk.io.