Two times. Not ten. That is the multiplier one practitioner landed on after a year of serious daily work with coding models in 2026, and it is the most useful number I have seen published on this subject. It is useful precisely because it is small. A 2x speedup is real, measurable, and worth restructuring your workflow around. A 10x claim is a marketing artifact that tells you nothing about where the gains actually come from.
I want to sit with that number for a moment, because the interesting question is not whether it is accurate. The interesting question is what kind of work the model is doing when you get 2x, and what kind of work it is doing when you get nothing at all.
The Comment Rule Is an Architectural Statement
The same practitioner offers a rule that sounds almost puritanical: never let the model write READMEs, docstrings, or comments. Write those yourself, later. And the stated emphasis is that this is not a stylistic preference — it is meant literally.
I came to a similar conclusion, and I think the reason is structural rather than aesthetic. Code and prose look similar to a token predictor. They are not similar to a reader. Code is checked by a compiler, a test suite, and a runtime. It has an external oracle. Documentation has no oracle. Nothing fails when a docstring describes intent that the function does not have. So when a model generates code, errors surface. When it generates comments, errors are laundered into something that reads as authoritative and never gets contradicted.
There is a second failure mode that matters more for anyone maintaining a system over years. A comment written by a model is a description of the code as it appears. A comment written by an engineer is a record of why the code is not something else — which alternative was rejected, which constraint forced the awkward branch, which bug the guard clause remembers. That information does not exist in the source. It exists only in the head of the person who made the decision. No amount of context window recovers it.
So the rule is less about distrust and more about division of labor. The model handles the part with a verifier. You handle the part that carries knowledge nothing else can reconstruct.
Vague Prompts Fail for a Reason You Can Name
Addy Osmani describes a common mistake as going straight to code generation with a vague prompt, and his fix is to brainstorm a detailed specification with the model first, then outline from there. Reported practice across the field in 2026 points the same way: detailed planning and example-driven prompts.
The mechanism here is worth being precise about. An underspecified prompt does not produce a wrong answer. It produces an answer to a question the model had to invent. Every ambiguity you leave open is a decision the model makes silently, according to whatever pattern is most common in its training distribution. Your review burden is then the sum of all those invisible choices, and review is where the 2x quietly dies.
Specification-first prompting collapses the space of plausible outputs before generation starts. Examples do this even more aggressively than prose instructions, because an example pins down format, naming, error handling, and tone all at once, without you having to enumerate any of them.
Documents as Input, Not Just Output
One pattern from the 2026 discussion deserves more attention than it gets. Instead of manually checking a long list of items, you hand the model your own article or list in markdown and ask it to verify. The artifact you wrote becomes the input.
This inverts the usual framing. The model is not the author; it is the checker, and you supply the ground truth. That is a much better fit for what these systems are good at. Consistency checking against a provided reference is a task with an anchor. Free generation is not.
Long Outputs Are Not the Same as Correct Ones
Capability has genuinely moved. Early models like GPT-1 fell apart into nonsense after a few sentences. Today’s systems, GPT-6 among them, sustain tens of thousands of coherent words and write working code. MiniMax M3 sits in the same leading tier.
Coherence at that length is a real achievement, and it is also a trap. The failure mode is no longer visible degradation. It is fluent, plausible, sustained output that is wrong in one place you did not check. Longer outputs mean more surface area per unit of your attention, and your attention did not scale alongside the context window.
Which brings the argument back to 2x. The gain comes from generation under specification with a verifier attached. It does not come from volume. The practitioners reporting honest numbers are the ones who figured out which half of the work to keep.
đź•’ Published: