Two AI agents are now arguing about you.
One sits inside a hospital’s billing department, optimizing how your visit gets coded. The other sits inside an insurer’s claims pipeline, optimizing how much of that code gets paid. Neither one examined you. Neither one knows what happened in the exam room. And according to the Blue Cross Blue Shield Association, their disagreement has already added $942 million in healthcare spending over two years.
The New York Times framed this as an escalation of a longstanding feud between hospitals and insurers, with AI as the new weapon. That reading is correct but incomplete. What’s happening in medical billing is one of the first large-scale, real-money demonstrations of a failure mode that agent researchers have been describing on whiteboards for years: two optimizers pointed at the same document with opposing objectives, and no shared source of truth between them.
What actually breaks here
Consider the architecture. A hospital deploys a coding agent whose objective function is, roughly, maximize reimbursement subject to staying inside defensible documentation. An insurer deploys a review agent whose objective is, roughly, minimize payout subject to staying inside defensible policy. Both objectives are legitimate from inside their own organizations. Both are proxies. Neither is “represent the patient’s care accurately.”
That last part is the whole problem. The reported pattern is billed complexity rising without corresponding changes in care delivered. In agent terms, the system is optimizing the representation rather than the reality it represents. The clinical encounter is the ground truth, but nothing in the loop is checking the code against it. The code is checked against policy documents, historical approval rates, and whatever signal each side has learned about what the other side will tolerate.
Once you remove grounding, an optimizer will happily walk uphill forever. There is always a slightly more complex, still technically defensible code. There is always a slightly stricter, still technically defensible denial.
Adversarial co-evolution without a referee
Self-play made AlphaGo strong because the game had rules, a terminal state, and a scoring function everybody agreed on. Billing has none of those properties. The rules are natural-language policy documents open to interpretation. There is no terminal state, just appeals. And the “score” is contested by definition, since the two players compute it differently.
So instead of convergence, you get drift. Each side’s agent learns from the other’s responses, adjusts, and the equilibrium moves. The measurable output of that drift is administrative cost: more codes generated, more codes reviewed, more appeals, more staff time on both sides. The $942 million figure is best understood not as fraud but as the compute bill for a negotiation that got automated on both ends at once.
And it’s asymmetric in an interesting way. Hospital agents operate on rich, unstructured clinical text. Insurer agents operate on the compressed output of that process, a handful of codes and modifiers. The side holding the raw signal has more room to shape the representation. The side holding only the summary has to infer intent from a lossy encoding. That information gap is where the extra spending lives.
What this predicts for other domains
Healthcare billing is early, not unique. The same structure appears anywhere two organizations exchange structured claims about unstructured reality:
- Legal discovery — one agent generating requests, another agent filtering responsiveness, both scored on volume rather than relevance.
- Procurement and invoicing — supplier agents optimizing line-item description, buyer agents optimizing rejection.
- Insurance underwriting beyond health — risk representation versus risk assessment, same lossy channel.
- Ad verification — supply-side agents describing inventory, demand-side agents auditing it.
In each case, deploying an agent on one side creates pressure to deploy one on the other. That’s not a market failure; it’s rational. But the aggregate result can be more total cost with no more delivered value, which is exactly the outcome insurers are describing now.
Designing around it
The architectural lesson is not “don’t automate billing.” It’s that adversarial agent pairs need something the current setup lacks: a shared, verifiable reference that neither party controls. In practice that could mean structured clinical attestation generated at the point of care, before billing intent exists. It could mean audit agents scored on agreement with retrospective chart review rather than on approval or denial rates. It could mean rate limits on appeal loops, which sounds bureaucratic but functions as a termination condition.
Any of those would give the loop a fixed point. Without one, we should expect the number to grow, because nothing in the system is currently rewarded for making it stop. When you deploy an optimizer against a proxy metric, the proxy is what improves. We now have a nine-figure estimate of what that costs when the proxy happens to be a medical bill.
🕒 Published: