\n\n\n\n Keep Your Hands Dirty in an LLM World - AgntAI Keep Your Hands Dirty in an LLM World - AgntAI \n

Keep Your Hands Dirty in an LLM World

📖 5 min read•862 words•Updated Sep 27, 2026

Joy is an engineering requirement.

That sounds soft for a blog about agent architecture, but I mean it in a mechanical sense. Enjoyment is the signal that tells you you’re still in contact with the system you’re responsible for. When programming stops being fun, it’s usually because comprehension has drained out of the loop. You’re approving diffs you don’t understand, in a codebase whose shape you can no longer picture. The boredom is diagnostic.

Going into 2026, the advice circulating among practitioners converges on something unglamorous: keep writing code, use models as assistants rather than replacements, hand off the grunt work, and hold onto critical thinking and oversight. Addy Osmani, after a year of refining his own AI-assisted engineering process, frames the model as a powerful pair programmer that needs clear direction, context, and oversight rather than autonomous judgment. That framing is worth taking seriously, because it’s an architectural claim disguised as a productivity tip.

The 2x ceiling is a feature

One of the more honest posts on this subject is titled, plainly, “2x, not 10x.” The author describes keeping steady progress on a 70k+ line GUI side project using Codex and Sol, without the endless whack-a-mole of bugs: squash a bug, review the architecture, move on, repeat. Notice the rhythm. There’s a generation step and there’s a review-the-architecture step, and the second one is what keeps the first from compounding into chaos.

That’s why 2x is a realistic number for sustained work on a system you intend to keep. The generation phase can go far faster than 2x. The comprehension phase cannot, because comprehension is bounded by your own reading speed and your own working memory of the system. Any workflow that claims a 10x steady state is usually borrowing against understanding it hasn’t built yet. The loan comes due during the next refactor.

Wasteland risk

The sharpest warning I’ve read on this goes roughly like this: if you want to keep owning your codebase, you need to keep writing some of the code. Let it all be generated, and it turns into an LLM wasteland that only your coding agents can thrive on.

I want to unpack why that’s structurally true and not just sentimental. Code written entirely by agents optimizes for a particular reader. It tends toward verbose local correctness, defensive boilerplate, repeated patterns instead of shared ones, and naming that’s plausible rather than precise. None of that hurts an agent, which re-reads the file from scratch every time and has no accumulated mental model to violate. It hurts you enormously, because human comprehension depends on compression — abstractions, conventions, and names that let you skip reading.

So a fully generated codebase isn’t just unpleasant. It’s a system whose only viable maintainers are the agents that produced it. That’s a dependency you took on without a design review.

What “some code” should mean

The useful question isn’t how much code you write by hand. It’s which code. My own division of labor:

  • Write by hand: the core data model, the interfaces that everything else depends on, the concurrency and state-transition logic, and anything where you’d struggle to spot a subtle wrongness in review.
  • Delegate freely: adapters, serialization, test scaffolding, migration scripts, CLI plumbing, the fifth variant of a pattern you’ve already established by hand.
  • Delegate then rewrite: anything where the generated version works but reads like a stranger wrote it. Keeping the behavior and rewriting the shape is cheap and preserves your model of the system.

This maps onto high-level versus grunt work, but with a sharper criterion: retain the parts that carry your mental model, delegate the parts that merely implement it.

The counterexample worth studying

Anthropic is the case people cite in the other direction. Engineers there adopted Claude Code so heavily that roughly 90% of Claude Code’s own code is now written by Claude Code. That’s a real data point, and it deserves a real reading rather than a dismissal.

What it shows is that very high generation ratios are achievable when the team has extremely tight feedback loops, deep familiarity with the tool’s failure modes, and a product whose domain is the tool itself. Even Osmani, noting the figure, is clear that using models for programming is not a push-button process. The ratio is an outcome of disciplined practice, not a substitute for it.

The quiet upside

There’s a genuinely good thing happening alongside all this. A lot of people who had mostly stopped programming as their careers moved elsewhere are coding again. Models lower the activation energy for starting, which is exactly where most side projects die.

Jaana Dogan, Principal Engineer at Google, shared her LLM predictions for 2026 with Oxide and Friends, and the fact that these conversations are now happening among senior practitioners rather than vendors is a decent sign. The people closest to the work are the ones asking how to stay in it.

Stay in it by staying legible to yourself. Read the architecture before you ship the fix. Rewrite the file that reads wrong. Keep enough of the system in your head that you’d notice if it changed underneath you. That’s not nostalgia for typing. It’s ownership, and it happens to feel good.

🕒 Published:

🧬
Written by Jake Chen

Deep tech researcher specializing in LLM architectures, agent reasoning, and autonomous systems. MS in Computer Science.

Learn more →
Browse Topics: AI/ML | Applications | Architecture | Machine Learning | Operations
Scroll to Top