\n\n\n\n Ninety Percent Machine-Written and Still Expected to Brake in Time - AgntAI Ninety Percent Machine-Written and Still Expected to Brake in Time - AgntAI \n

Ninety Percent Machine-Written and Still Expected to Brake in Time

📖 5 min read•803 words•Updated Sep 25, 2026

General Motors’ autonomous driving team now generates nearly 90% of its code with AI. That figure came out of GM’s own disclosures, and it is the kind of number that reorders how you think about the rest of the stack. Mikell Taylor, GM’s Director of Robotics Strategy, is scheduled to sit on a panel at TechCrunch Disrupt 2026 alongside Shield AI CTO Nathan Michael and Waabi founder and CEO Raquel Urtasun, on the Real World AI Stage, to discuss exactly this class of problem: what you do when the system cannot be allowed to fail.

My first reaction to the 90% figure was not excitement. It was a question about verification asymmetry. If code generation gets ten times cheaper and validation does not, you have not sped up development. You have moved the bottleneck and made it narrower.

The cost curve nobody is talking about

In safety-critical engineering, writing the code was never the expensive part. The expensive part is the evidence. Requirements traceability, coverage analysis, hazard mapping, failure mode reasoning, the argument you hand a regulator that says this behavior is bounded and here is why. That work scales with the volume and novelty of the code under review, not with who or what typed it.

So a team producing nine out of ten lines through generation faces a compounding problem. More code arrives per sprint. Each line carries a slightly different provenance story than a human-authored line would. The engineer who reviews it did not build the mental model that produced it. Reviewing unfamiliar code is measurably harder than reviewing your own, and generated code has a specific failure signature: it is plausible. It compiles, it reads cleanly, it passes the obvious tests, and it is wrong in ways that only surface under conditions your test suite did not imagine.

That is the interesting tension in this panel lineup. Three organizations, three very different definitions of acceptable failure.

Three different tolerances for being wrong

Waabi has built its approach around simulation as the primary validation surface, and the company closed $1 billion in new funding in January 2026. Its public framing this year has centered on generalization, which is the correct technical target and also the hardest thing to prove. Generalization claims are claims about the space of situations you have not tested. You cannot enumerate that space. You can only build an argument about why your system’s behavior degrades gracefully at its edges, and that argument is architectural, not empirical.

Shield AI operates in defense, where the adversary is actively trying to produce the inputs your model handles worst. That changes the validation question from “does this work in the expected distribution” to “what happens when someone is paid to find the tail.” Nathan Michael’s engineering constraints are not GM’s. Degraded-mode autonomy, contested communications, and deliberate spoofing are not conditions you simulate by sampling harder from a driving dataset.

GM sits in the consumer regulatory regime, which is the slowest and least forgiving of the three. A defense program can accept a mission failure rate. A consumer vehicle program accepts approximately none, and every incident becomes a public artifact.

What I want to hear from the stage

The framing around this session is safety validation, regulatory navigation, and trust-building. Those are the right words. The useful version of that conversation gets specific about mechanism.

  • Where does the generated code sit in the architecture? Perception pipelines, planning logic, and the safety monitor that catches both are not equivalent risk tiers. Ninety percent generated in tooling and data infrastructure is a productivity story. Ninety percent generated in the arbitration layer is a different conversation.
  • What does review actually look like at that volume? Human sign-off per line does not survive contact with this throughput. So either the verification is itself automated, in which case what verifies the verifier, or the architecture confines generated components behind boundaries whose properties are checked independently.
  • How do you write a safety case for a component you did not author? Certification frameworks assume a design rationale exists and can be articulated. Generated code has an output but not always a rationale.

My own read is that the answer is structural rather than procedural. The teams that make this work will not be the ones with the strictest code review policies. They will be the ones whose architectures make the correctness of any individual component less load-bearing: runtime monitors with independently specified properties, formally bounded fallback behavior, and hard separation between the part of the system that is statistically excellent and the part that is deterministically conservative.

Generation speed is not the achievement here. Containment is. If GM, Waabi, and Shield AI have built systems where a 90% generation rate is survivable, the architecture doing that work is the part worth studying, and I hope someone on that stage is asked to describe it.

đź•’ Published:

🧬
Written by Jake Chen

Deep tech researcher specializing in LLM architectures, agent reasoning, and autonomous systems. MS in Computer Science.

Learn more →
Browse Topics: AI/ML | Applications | Architecture | Machine Learning | Operations
Scroll to Top