Almost every conversation about the danger of AI agents is a conversation about intelligence. How capable is the model. How autonomous is the system. How likely is it to deceive, to scheme, to pursue goals we did not give it. The entire discourse is organised around the question of how smart the agent is, on the assumption that intelligence is the axis of risk.
It is the wrong axis. The dangerous property of an AI agent is not how smart it is. It is how fast it is. Speed is the vulnerability, and almost everything that makes agents hard to govern traces back to it rather than to their intelligence. This is not a minor reframing. It changes what you build to make agents safe, because the controls that address intelligence are not the controls that address speed, and it is speed that is actually hurting you.
The Thought Experiment That Isolates the Variable
To see that speed is the real problem, hold everything else constant and change only the speed.
Take an agent exactly as dangerous as the ones you worry about. Same intelligence, same access, same tools, same susceptibility to manipulation. Now impose one artificial constraint: it may take one action every few minutes, and a human sees each action before the next one begins. Nothing about its capability has changed. It is precisely as smart, precisely as capable of forming a harmful plan, precisely as manipulable by malicious input. All that changed is the clock.
Almost every catastrophe you were worried about becomes catchable. The manipulated agent that would have drained an account now takes its first suspicious action and waits, visibly, while a human looks at it. The reasoning error that would have cascaded through a thousand steps produces one wrong step, seen, before the second. The subtle misalignment that would have compounded is caught at the first divergence. The intelligence is unchanged and the danger has largely evaporated, because the harm was never in the thinking, it was in the thinking reaching the world faster than anyone could respond.
Now remove the constraint and let it run at machine speed again. Nothing about its intelligence changed on the way back up. But every control that depended on a human being able to see and respond has just collapsed, all at once, because the actions now complete faster than a human can register that they are happening. The agent did not get smarter or more malicious when you sped it up. It got dangerous. The danger was the speed the entire time.
This is worth sitting with, because it inverts the standard picture. We imagine a spectrum from safe, dim agents to dangerous, brilliant ones, and we imagine safety as a problem of keeping intelligence in check. The thought experiment says otherwise. A brilliant agent at human speed is largely governable. A modest agent at machine speed is not. Intelligence is not the variable that determines whether you can control the thing. Speed is.
Every Human Control Is a Timing Assumption
The reason speed is so corrosive is that every safety mechanism we inherited was built for actors that move at human pace, and every one of them contains a hidden assumption about time. Speed does not defeat these controls by overpowering them. It defeats them by violating the timing assumption they were silently built on.
Detection assumes there is time to notice. A control that catches a problem works because the problem persists, or unfolds, over an interval long enough for a monitor, human or automated with a human in the loop, to observe it. Collapse the interval and detection has nothing to grab. The problem is already over.
Review assumes there is time to look. Oversight of a person's work, a manager's glance, a second pair of eyes, functions because human work arrives slowly enough to be examined. There is no reviewing a thousand actions a minute. The review model assumes a rate of production it can keep up with, and the agent produces faster than any reviewer can read.
Intervention assumes there is time to react. The kill switch, the escalation, the emergency stop, all assume a gap between noticing and harm that is wide enough to act inside. At machine speed that gap closes. By the time a human decides to intervene, the actions they wanted to prevent are complete.
Even hesitation assumes human pace. The most underrated safety feature of a human worker is that they pause, that something feels wrong and they slow down. That hesitation is a function of operating at human speed. An agent does not hesitate, and even if you built hesitation into it, the pause would be measured in milliseconds, long enough to change nothing.
| Human control | The timing assumption it makes | Why speed breaks it |
|---|---|---|
| Detection | The problem lasts long enough to notice | The action completes before it registers |
| Review | Work arrives slowly enough to examine | Thousands of actions outpace any reviewer |
| Intervention | There is a gap between noticing and harm | The gap closes faster than a human reacts |
| Hesitation | The actor pauses at human pace | The agent does not pause, or pauses trivially |
| Anomaly response | The anomaly unfolds slower than the responder | The anomaly completes before the response begins |
Read the middle column top to bottom and the pattern is undeniable. These are not five different failures. They are one failure, five times: every control assumes the actor moves slowly enough for a human to fit a response inside the window, and the agent has closed the window. The controls did not get weaker. The timescale they depend on disappeared.
Speed Converts Recoverable Into Unrecoverable
There is a second way speed does damage, distinct from defeating oversight, and it is about the nature of the harm itself.
At human speed, a mistake is usually a single event you can catch before it becomes many. A person makes an error, and there is a moment, often several, before the error propagates, in which it can be caught, corrected, undone. Recoverability is a function of the gap between the first wrong action and the ones that follow. Human pace keeps that gap open. Most human errors are recoverable precisely because the person was slow enough that the error stayed small while someone noticed.
At machine speed, the gap is gone. The first wrong action is followed immediately by the second, the tenth, the thousandth, all before anyone could have intervened. A single error does not stay a single error. It becomes a completed cascade, an accomplished fact, in the time it would have taken a human to notice the first one. Speed converts what would have been a recoverable mistake into an unrecoverable one, not by making the mistake worse, but by removing the interval in which it could have been contained.
This is why blast radius is really a speed problem. We talk about agents having a large blast radius as though it were about their access, and access matters, but the reason a wrong action becomes a large one is that nothing could get between it and the next before the whole sequence completed. A slow actor with the same access has a small blast radius, because the harm is paced out into catchable pieces. Speed is what fuses the pieces into a single unstoppable event.
Speed Is What Makes Manipulation Catastrophic
The same logic explains why manipulation of agents is so much more dangerous than manipulation of people, and again the difference is speed, not susceptibility.
People are manipulated constantly. Social engineering works, phishing works, deception works. Humans are not obviously harder to fool than a well built agent. If susceptibility were the variable, manipulated humans would be as catastrophic as manipulated agents, and they are not, because the manipulated human acts at human speed. The con takes time to execute. The victim does the harmful thing slowly, with pauses, in a sequence that can be interrupted. A large fraction of human social engineering attacks are caught mid execution, because the execution is slow enough to catch.
A manipulated agent completes the attack in the window before anyone can respond. The manipulation itself is not the new thing, agents are fooled by the same broad category of tricks that fool people. The new thing is that the fooled actor executes at machine speed, finishing the harmful sequence in the interval that, for a human, would have been the interval in which the con was noticed and stopped. Manipulation plus human speed is a risk we have managed for all of history. Manipulation plus machine speed is the thing we have not, because the response window that made the first manageable does not exist in the second.
the reaction gap
agent action: ▓▓▓▓▓▓▓▓▓▓ complete
└── milliseconds ──┘
human response: ....see.... understand.... decide.... act
└────────── seconds ──────────┘
│
▼
response lands after the fact
the harm finishes inside the gap the response needs to even begin.
The Same Incident, at Two Speeds
Make it concrete by running one incident at both speeds, changing nothing else.
An agent with access to financial systems is manipulated into moving funds to an external destination. Run it first at the human-speed constraint. The agent takes its first step, staging the transfer, and pauses for the next interval. A person, reviewing actions as they arrive, sees a transfer to an unfamiliar destination appear where none should be, recognises it as wrong, and stops it. The manipulation succeeded at fooling the agent. It failed at causing harm, because the harmful action sat in the open, at human pace, long enough to be caught. This is the world we know how to operate in, the world where a fooled actor is slow enough that the fooling gets noticed before the damage lands.
Now run the identical incident at machine speed. The agent is manipulated in exactly the same way, forms exactly the same harmful intent, and begins exactly the same sequence. This time it stages the transfer, executes it, confirms it, and moves on, all within a second. There is no interval in which the suspicious action sits in the open, because there are no intervals at all; the sequence is continuous and complete before a human could have looked. The same person, no less attentive, learns about the transfer the way everyone learns about machine-speed harm, afterward, from the consequences. Nothing about the agent differed between the two runs. The manipulation was identical, the intent was identical, the access was identical. Only the clock changed, and the clock was the difference between a caught attempt and a completed breach.
That is the entire thesis in one incident. The dangerous ingredient was not the agent's capability or the attacker's cleverness, both of which were constant. It was the speed, which in one run left a window and in the other did not.
Why the Obvious Fixes Miss
Once you see that speed is the vulnerability, the popular remedies reveal themselves as aimed at the wrong target.
Make the agent smarter, better aligned, more honest. Suppose you succeed completely and the agent is wrong only rarely and never malicious. It does not help enough, because a rare wrong action at machine speed is still catastrophic. The whole problem is that a wrong action completes before it can be caught, and that is true whether wrong actions are frequent or rare. Reducing their frequency reduces how often you are hurt, but each time it happens, speed still turns the single error into a completed cascade. Alignment addresses whether the agent tends to do the wrong thing. It does nothing about what happens in the moment it does, which is where speed does its damage.
Slow the agent down. This actually would work, and that is the tell. Impose the human-speed constraint from the thought experiment and the danger really does recede. But you cannot, because speed is the entire reason you deployed the agent. You did not automate the work to have it run at the pace of a person watching each step. Slowing the agent to the speed of human oversight throws away the value that motivated the deployment. The fix that works on the vulnerability destroys the reason for the system, which means it is not available.
So you are caught. You cannot fix it by making the agent smarter, because intelligence is not the axis of the danger. You cannot fix it by slowing the agent down, because speed is the point. The remedies aimed at intelligence do not touch the problem, and the remedy that touches the problem is not one you can use.
The Only Oversight That Keeps Pace
The way out is to stop trying to fit a human response inside a window the agent has closed, and instead put the control inside the action itself, where it runs at the agent's speed by construction.
Every failed control has the same structure: something happens, then a human, or a human-paced process, tries to respond, and loses the race. The response is separate from the action and slower than it. The fix is to make the control not separate and not slower, by placing it in the path of the action, executed automatically as part of the action's own execution. The action does not complete and then get reviewed. The action is evaluated before it is allowed to execute, synchronously, by an automated check that runs at the same speed the agent runs, because it is part of the same motion.
This is what authorisation at the action provides, and it is the only form of oversight that keeps pace with a fast actor. It does not depend on a human noticing in time, because there is no window between action and response for the human to lose; the evaluation is the gate the action passes through, not a reaction to an action that already happened. A wrong action is not caught after the fact by an observer who was too slow. It is declined before the fact by a check that was never in a race, because it was never behind. The control operates at machine speed because it is machine speed, an automatic evaluation in the action path, with human judgment reserved for the specific decisions that are escalated to it and can wait for it.
Notice what this does to the whole picture. It does not require the agent to be slower, so the value of the deployment survives. It does not require the agent to be smarter or more aligned, so it holds even when the agent is wrong or manipulated. It simply refuses to rely on a human winning a speed race, and instead makes the check part of the action, where speed can no longer defeat it. You cannot out-run a gate you have to pass through.
Stop Asking If It Is Smart Enough to Trust
The reframing is practical, not philosophical. The dominant question, is this agent intelligent and aligned enough to trust with this, is aimed at the wrong variable, and it will lead you to controls that do not address the danger. You will vet the model, test its alignment, satisfy yourself that it is capable and well behaved, and deploy it, and none of that will have touched the property that actually determines whether you can govern it.
The question that matters is whether your oversight operates at the agent's speed. If your answer to what happens when it goes wrong involves a human noticing, reviewing, deciding, or intervening, your oversight runs at human speed and the agent does not, and you have already lost the race you are describing. The only oversight that survives contact with a fast actor is oversight that is not in a race at all, because it is part of the action rather than a response to it.
Speed is the vulnerability. Not intelligence, not autonomy, not the model's character. The thing that makes an agent ungovernable is that it acts faster than anyone can respond, and the answer is not to make it slower or smarter but to make the control as fast as the action by putting the control inside it. Govern the action, in the action, at the speed of the action. It is the only thing that keeps pace, and keeping pace is the whole game.
Xybern is the authorisation layer for enterprise AI agents. Every agent action is enforced, audited, and governed before it executes. Learn more at xybern.com or read the technical documentation at docs.xybern.com.
Xybern
