A broken door still tells you the truth. It sticks, it groans, it swings the wrong way, and anyone standing in front of it can form an accurate theory of what is wrong within about four seconds. The failure is legible. You can even laugh at it, which is why doors have been reliable comedy props for a century. A broken AI service offers none of that. It returns something. The something is shaped correctly. It is wrong in a way that no one in the room can explain, and by the time anyone notices, the wrongness has been written to three downstream tables.
That gap between legible failure and illegible failure is, I think, the defining engineering problem of 2026. We have not gotten worse at building systems. We have gotten worse at building systems that can explain themselves when they break, and then we have quietly agreed to stop asking them to.
The shrug as a design decision
Software engineering right now is full of failures that get noticed and then dropped. Something didn’t run. Something ran twice. A sync job produced a number nobody can reconcile. The ticket gets filed, the ticket gets closed with a retry, and the organization moves on. Nobody decided to accept this. It accumulated.
From an architecture standpoint, though, an unexplained failure that gets tolerated is not neutral. It is a decision you have made about your system’s error surface. You have declared that this class of fault is beneath investigation, which means you have also declared that you will not learn anything from the next hundred instances of it. Tolerance compounds. Every shrug widens the region of your system where behavior is unmodeled, and unmodeled regions are exactly where the expensive surprises live.
The failures I find most interesting are the moderate ones, because they are the most common and the least studied. A total outage gets a postmortem. A moderate project failure — late, over budget, half the promised functionality, customers visibly annoyed — gets absorbed into the general cost of doing business. The money and the time are real. The institutional learning is close to zero.
Why agents make illegibility worse
Deployment trouble in AI and ERP systems keeps increasing, and the two usual explanations are complexity and a shortage of people who know how any of it works. Both are true. But I want to be precise about what kind of complexity we are adding, because agent architectures change the shape of the problem rather than just the size of it.
Traditional enterprise software fails with a stack trace. The failure has a location. Agentic systems fail differently:
- The failure has no single location. A bad outcome may emerge from a reasonable plan executed against a stale tool response, with no individual step that looks wrong in isolation.
- The failure is not reproducible on demand. Same input, different trajectory. Your bug report describes a behavior that no longer exists.
- The failure is confident. Systems that must always produce output will always produce output, including when the correct answer was ” “
- The failure is cheap to paper over. An agent can be nudged toward better behavior with a few prompt edits, which feels like a fix and is actually a suppression of the symptom.
That last one deserves emphasis. The ease of adjustment is precisely what makes these systems hard to engineer. When a repair costs two sentences, nobody builds the instrumentation that would have told you what actually went wrong. The cheap fix outcompetes the correct fix on every sprint board in the world.
Layering transformation on top of unfinished basics
There is a pattern worth naming here. Many organizations struggling to deploy AI never finished deploying their previous generation of systems. Basic accounting and general ledger technology is still not working cleanly, and the response has been to add another transformation layer on top. You now have two systems with unexplained behavior, coupled, with a shortage of staff who understand either one.
Agent frameworks are unusually good at hiding this. They present a clean interface over a messy substrate, which is genuinely useful right up until the substrate is the thing that’s broken. Then the abstraction stops being a simplification and becomes a blindfold.
What I would actually ask for
Not more logging. We have plenty of logs nobody reads. What agent systems need is a discipline of explainable failure: every action taken by an agent should carry enough recorded context — inputs, tool responses, the state it believed it was in — that a human can reconstruct the reasoning after the fact without rerunning anything. Treat unexplained behavior as a defect in its own right, separate from whether the outcome happened to be acceptable. And keep a written count of faults you’ve chosen not to investigate, because an unexamined shrug is invisible and a counted one is a roadmap.
The door that sticks is a solvable problem. The service that lies politely is not, until we insist it tell us why.
🕒 Published: