Over 12,000 inboxes compromised across more than 10,000 organizations. Two arrests. Those two numbers, sitting next to each other, describe the entire problem with AI-assisted crime services better than any threat report could.
On 22 September 2026, Microsoft’s Digital Crimes Unit announced it had disrupted EvilTokens, a subscription-based cybercrime service that used AI to automate phishing. The Metropolitan Police Service had arrested two men, aged 32 and 38, on 11 September in connection with the operation. The operators numbered in the single digits. The victims numbered in the tens of thousands. That ratio is the story.
Read the numbers as an architecture diagram
Roughly 12,000 compromised inboxes spread across roughly 10,000 organizations works out to a little more than one inbox per organization. That distribution is worth sitting with, because it is not what mass phishing looks like. Spray-and-pray campaigns cluster: one organization gets hit, the attacker harvests the address book, and the compromise spreads laterally inside that tenant. You end up with deep, narrow damage.
A flat distribution across ten thousand organizations points somewhere else entirely. It suggests a pipeline optimized for breadth of selection rather than depth of exploitation, one that could identify a worthwhile target inside an unfamiliar organization, get in, and move on. Doing that well across ten thousand different environments is a targeting problem, not an exploitation problem.
The intelligence went where the bottleneck was
Reporting on the disruption describes an AI chatbot that helped attackers decide which victims to pursue and how to exploit them. Note the placement. The model was not writing novel exploits or discovering vulnerabilities. It sat in the decision layer, above the tooling, answering the question that has always throttled criminal operations: of all the people I could attack, which ones are worth my next hour?
This is, architecturally, the same insight that makes agent systems useful in legitimate work. The expensive part of most pipelines is rarely the execution step. It is the judgment step that decides what to execute against, in what order, with what framing. Automating execution gives you volume. Automating selection gives you yield. EvilTokens appears to have been built by people who understood that distinction.
The technique itself, device-code phishing, reinforces the point. Device-code flows exist so that input-constrained devices can authenticate by having a user approve a code elsewhere. Abusing that flow yields tokens rather than passwords, which is a meaningfully different prize. The attacker’s job becomes convincing one person to approve one code in a plausible context. That is a social and contextual challenge, and contextual plausibility at scale is exactly what language models are good at manufacturing.
Subscription is the part defenders should worry about
The service model deserves more attention than the model behind it. A subscription-based platform decouples capability from competence. The person paying for access does not need to understand token flows, tenant discovery, or how to write a convincing internal IT notice. They need a credit card and a target list, and the platform supplies the judgment.
That has two consequences worth thinking through:
- The talent floor drops to zero. Historically, the scarcity of skilled operators limited how many organizations could be attacked competently in a given week. Packaging judgment as a service removes that limit.
- Attribution gets harder, not easier. Two administrators can serve an unknown number of customers. Disrupting the platform is high-value precisely because the customer base is diffuse and largely invisible.
Which is why a coordinated takedown, with Microsoft working alongside industry partners and law enforcement, is a reasonable response to this class of threat. You cannot arrest a customer base. You can remove the infrastructure that made the customer base dangerous.
What I would want to know next
The public account leaves the interesting engineering questions open, and I would rather flag them than guess at answers. We do not know from what has been published how much autonomy the system actually had, whether the chatbot recommended actions to a human operator or drove tooling directly, what model or models sat underneath, or how target prioritization was scored. Those details separate a well-built assistant from something closer to an autonomous agent, and they matter enormously for predicting what the next version looks like.
For defenders, the practical takeaway is less about AI than about assumptions. Controls designed around the premise that sophisticated targeting is rare and expensive are built on a premise that a subscription fee now undermines. Authentication flows that leak tokens rather than credentials deserve the same scrutiny as password entry points, because that is where the volume went. And detection logic tuned to spot clustered, repetitive campaigns will struggle against a pipeline whose entire design goal is to look like one unremarkable request per organization.
Twelve thousand inboxes, two administrators. The disruption is genuinely good news. The ratio is the warning.
đź•’ Published: