Financial regulators have spent a century learning to watch for the same shape of failure. A bank lends too much, a counterparty vanishes, a margin call cascades. The pattern is slow enough to name while it is happening. What Andrew Bailey has put in front of the G20 is a different shape entirely: a system where the failure mode is not a balance sheet but a behaviour, and where the thing behaving is a model nobody fully understands from the inside.
The Financial Stability Board, which Bailey chairs alongside his role as Bank of England Governor, has warned that frontier AI models could threaten global financial stability. His specific framing is the part worth sitting with. AI, he said, could alter the speed, scale and economics of cyber risk. Three variables, all of them about dynamics rather than exposure.
Speed, scale, economics — read as systems properties
I want to take those three words seriously as technical claims, because they map onto things we can actually reason about in agent architecture.
Speed is the one that should worry architects most. Financial systems have circuit breakers, settlement windows, human sign-offs. These are all latency-based safety mechanisms. They work because the adversary and the market participant both operate on human timescales, or close enough. An agentic system that can run reconnaissance, exploitation and lateral movement in a tight loop does not violate any control in the rulebook. It just operates below the resolution at which the rulebook was written. A control that assumes a minute of reaction time is not a control against something that finishes in seconds.
Scale changes the shape of correlated risk. A single skilled attacker hits one target well. A model that can be replicated across a thousand parallel contexts hits a thousand targets adequately. Adequately, at that breadth, is worse. Financial stability frameworks treat concentration risk carefully — too many institutions exposed to the same counterparty. What we are describing here is a new concentration axis: too many institutions exposed to the same class of automated adversary, probing the same class of shared vendor infrastructure, at the same moment.
Economics is the quietest and possibly the most important. Cyber risk has always been priced, implicitly, by the cost of attacker labour. Skilled offensive work is expensive and scarce, which means a great many theoretically viable attacks are simply not worth executing. Drop that cost significantly and the set of economically rational targets expands downward. Small institutions that were never worth attacking become worth attacking. The long tail of the financial system, which is exactly where monitoring is thinnest, becomes newly interesting.
Why the agent framing matters here
There is a reading of these warnings that treats frontier models as better tools in familiar hands. I think that reading is too comfortable. The interesting risk is not a model that writes better exploit code. It is a system that plans, executes, observes and re-plans without a human in the loop for each cycle.
That distinction matters for defence design. Tool-assisted attackers are still bounded by human attention. Agentic attackers are bounded by compute and by the quality of their feedback signal. Those are different constraints, and they respond to different countermeasures. Rate limiting an API is a compute-side control. Poisoning or degrading an attacker’s observability is a feedback-side control. Most financial cybersecurity posture I have seen is built around neither — it is built around perimeter and detection, both of which assume a slower opponent.
What regulators can and cannot see
The structural problem for the FSB is that its usual instruments are disclosure-based. Capital ratios, stress tests, reporting requirements. These work because the underlying quantity is measurable and the institution holding it knows what it holds.
Frontier model capability is not like that. Capability is emergent, evaluated post hoc, and often discovered by users rather than developers. An institution cannot report its exposure to a capability that has not been characterised yet. A supervisor cannot stress-test against a threat model that updates faster than the supervisory cycle. This is a genuine mismatch between the tempo of the regulated phenomenon and the tempo of the regulatory apparatus, and I do not think anyone has solved it.
What I would want to see instead of capability disclosure is something closer to architectural disclosure. Not what can the model do, but where in your critical path does an autonomous system have write access, execution rights, or the ability to initiate a transaction. That question is answerable today. It is also, I suspect, going to produce some uncomfortable answers at institutions that have been quietly wiring agents into operations faster than their risk committees have been updating their maps.
Bailey’s letter is a warning, not a framework. But warnings from the chair of the FSB tend to precede consultation papers, and consultation papers tend to precede requirements. The institutions that start mapping their agent surface now will find that exercise considerably less painful than the ones that wait to be asked.
đź•’ Published: