Remember the open letter in 2023 that asked labs to pause training runs for six months? It landed with a lot of noise and almost no mechanism. There was no shared definition of what counted as a pause, no way to verify one, and no agreement on what the pause was supposed to produce. The training runs continued. The letter became a reference point in slide decks rather than a change in engineering practice.
So when Anthropic CEO Dario Amodei published an essay this past Saturday urging AI companies to deliberately slow the rate at which they push model capabilities forward, my first reaction was not skepticism about the motive. It was a question about the unit of measurement. Reuters reported that Amodei laid out a three-step framework meant to pace capability advances, and that other AI leaders, Sam Altman among them, backed the call. The stated reason is familiar to anyone who works on evaluation: capability is moving faster than the safety work meant to keep up with it.
That diagnosis matches what I see in practice. My concern is that “slow down” is being framed as a property of model training, when the thing actually accelerating in the field I work in is something else entirely.
Capability is not only a pretraining variable
If you measure AI progress by base model quality, then slowing the pace means longer gaps between frontier training runs, more careful scaling decisions, and more evaluation time before release. Reasonable. But agent systems have decoupled deployed capability from model capability in a way that makes that framing incomplete.
The same weights, wrapped in different scaffolding, are a different system. Give a model tool access, a persistent memory store, a planning loop, retry logic, and the ability to spawn subagents, and you have changed what it can accomplish in the world without touching a single parameter. Add a browser and a shell and you have changed its blast radius. The capability curve that matters for risk includes:
- Action space — what tools and side effects the system can reach
- Autonomy horizon — how many steps it takes before a human checks the work
- Composition — how many agents coordinate, and whether they can create new ones
- Persistence — whether state and learned strategy carry across sessions
- Integration depth — how far into production systems the agent is wired
None of those are gated by a training schedule. They ship in application code, often weekly, often by teams that are not the frontier lab. A pacing agreement between labs that leaves the scaffolding layer untouched slows the slowest-moving part of the stack.
What a three-step framework would need to do
I have not read the specifics of Amodei’s three steps beyond what has been reported, so I will not characterize them. What I can say is what any credible pacing proposal has to solve, based on the failure mode of every prior attempt.
It needs a measurable trigger
“Slower” is not a threshold. Something has to be counted, and the count has to be checkable by someone outside the organization doing the counting. Evaluation suites that labs run on themselves are useful engineering artifacts and weak governance instruments.
It needs to cover deployment, not just training
The riskiest thing an agent does is usually not reasoning. It is acting. A framework that constrains capability growth but says nothing about autonomy limits, tool permissions, or the conditions under which an agent gets write access to real systems has drawn its boundary in the wrong place.
It needs to survive competitive pressure
Public support from other executives, including Altman, is genuinely meaningful — coordination problems require coordination. But voluntary restraint holds exactly as long as restraint is not obviously costly. The 2023 letter is the control experiment here. Whatever mechanism follows has to work when a competitor decides not to participate.
Why this still matters
None of this is an argument against the call. The underlying claim, that safety work is losing a race against capability work, is one I would defend on technical grounds. Interpretability is improving but still cannot tell you why a long-horizon agent chose a particular plan. Evaluations for multi-step tool use are immature compared to benchmarks for single-turn reasoning. We are better at building agents than at auditing them, and the gap is widening.
A CEO of a frontier lab saying so publicly changes the conversation, because it makes deliberate pacing a defensible position rather than a competitive concession. That is worth something.
What would be worth more is a version of the framework aimed at the layer where capability is actually compounding. The agent stack is where the pace problem lives now. Slowing pretraining while shipping ever more autonomous scaffolding on top of existing models would satisfy the letter of the proposal and miss its point. The wheel is still turning; the question is what the brakes are attached to.
đź•’ Published: