One report credits an AI system built by Amazon with decrypting a German Army Enigma message from 1941 that had gone unread since 2005. Another set of reports credits OpenAI’s GPT-6 Astra with the same break, dated to a message from July 10, 1941, and says the Enigma researcher who publishes such material confirmed the result. Both cannot be casually true, and the fact that the record is muddled on day two of the story tells us something about how agent achievements now propagate.
I want to take the technical claim seriously and the attribution problem seriously, because for those of us who study agent architecture, they are the same problem wearing two hats.
What the reported task actually demands
Strip away the branding and look at the described workflow. The agent reportedly searched historical archives, compared uncertain letters, and found contextual constraints to resolve ambiguity. That sequence is worth reading slowly, because each step stresses a different part of an agent stack.
Archive search is retrieval under an open-ended query. There is no clean index of “1941 German Army traffic with plausible cribs.” The agent has to form a hypothesis about what kind of document would help, find it, and judge whether what it found is the thing it needed. That is planning plus evaluation, not lookup.
Comparing uncertain letters is the interesting part. Surviving Enigma intercepts are frequently damaged: characters garbled in transmission, ambiguous in the original log, or lost. A cryptanalytic search over rotor settings is brittle against that kind of noise, because a wrong character in the wrong position poisons a candidate key that was otherwise correct. Handling it means carrying multiple readings of the same ciphertext forward in parallel and scoring them against each other rather than committing early. That is a search discipline, and it is precisely the discipline that language-model agents have historically been bad at. The failure mode we all know is premature commitment: the model picks a branch, narrates confidence, and never returns.
Where the architecture matters more than the model
Finding contextual constraints closes the loop. German military messages were formulaic, and that formula is exactly what made them breakable in the 1940s. An agent that can pull a plausible phrase from a period document, use it as a crib, test it against a noisy intercept, and then feed the partial decryption back into the next round of archive search is running a hypothesis-refinement cycle across two very different kinds of evidence: statistical and historical.
The reported ten-hour duration is the number I keep returning to. Long-horizon coherence has been the practical ceiling on agent usefulness. Most systems degrade over extended runs, not because any single step is beyond them, but because errors compound and context management gets lossy. A ten-hour run that ends in a verifiable result suggests either better state management, better self-correction, or a task whose structure happens to be unusually forgiving of both.
This is also why cryptanalysis is a nearly ideal demonstration domain, and why I would resist reading too much into it. Enigma decryption is self-verifying. A correct key produces German; an incorrect one produces noise. The agent gets a clean, cheap, unambiguous reward signal at every step. Very few valuable real-world tasks look like that. One commenter on the Hacker News thread made the point in a different register, suggesting that physics-scale discovery requires years and several unintuitive leaps, and remains out of reach for now. I would put it more narrowly: the class of problems where an agent can grade its own work is the class where current systems shine, and it is smaller than the excitement implies.
The attribution mess is the real signal
Which brings me back to the contradiction. A result that its own reporting cannot attribute to the correct lab is a result the field cannot learn from. There is no method description, no failure count, no indication of how much human scaffolding shaped the run, and coverage elsewhere in 2026 has documented plenty of agent incidents that complicate the clean-capability narrative, including a reported Astra self-jailbreak that OpenAI said it blocked 91.5 percent of the time. Capability disclosures and safety disclosures arrive through the same leaky channel, and both get compressed into headlines.
The honest summary is narrow and still remarkable. Something autonomous appears to have read a message that stumped human researchers for twenty-one years, by combining archival reasoning with noise-tolerant search. What I would like next is not a bigger claim but a smaller, verifiable one: the trace. Show the branches it kept, the ones it pruned, and where it was wrong before it was right. That is the document that would actually teach us how agent intelligence works.
đź•’ Published: