What if the most strategically important AI technique in the United States right now is not a new architecture, but a teacher-student loop that has existed in the literature since 2015?
Y Combinator’s Garry Tan has been making the case that smaller American open-weight labs should distill from American frontier models, the same way Chinese labs have been distilling their way to surprisingly capable open releases. His argument, as he put it to TechCrunch, is that this would give the U.S. a stronger set of open alternatives and a healthier balance between closed frontier work and open-weight work.
I want to take that argument seriously from a systems perspective, because the interesting part is not the policy framing. It is what distillation actually does to the shape of an agent stack.
What distillation actually transfers
The folk understanding of distillation is that a small model copies a big model’s answers. That undersells it. When you train a student on a teacher’s outputs, especially on full distributions rather than hard labels, you are transferring something closer to a decision surface than a lookup table. The student inherits the teacher’s priors about what a reasonable next token looks like given a messy, ambiguous context. That inheritance is why distilled models often generalize better than models of the same size trained from scratch on raw web text.
For agent architecture, this matters more than it does for chat. Agents fail in the seams: mid-trajectory, after a tool call returns something unexpected, when a plan needs revision three steps in. Those failures are usually not knowledge failures. They are calibration failures. The model does not know how confident it should be, or when to stop, or when to ask. Teacher distributions encode exactly that kind of soft judgment, and it is the hardest thing to specify by hand in a reward function.
So the case for distillation in an agent context is not “cheap capability.” It is that trajectory-level supervision from a stronger planner is one of the few practical ways to move a small model’s behavior in the seams.
Why the open-weight framing changes the engineering
Tan’s point is specifically about open-weight labs, and that distinction has real technical consequences.
An open-weight distilled model can be inspected, quantized, fine-tuned again, and deployed at the edge of a system where latency budgets are tight. In agent stacks, that is where most of the work happens. You do not want a frontier model deciding whether a JSON field is malformed. You want a small, fast, predictable component doing it, with the frontier model reserved for the handful of steps that genuinely require deep reasoning.
That architecture, a router with cheap specialized workers and an expensive planner, is already how most serious agent deployments are converging. But it only works if the cheap workers are actually good. Right now, a lot of teams reach for whatever open-weight model performs well on public benchmarks, and increasingly those models come out of Chinese labs. Tan’s argument is essentially that the American supply of good small models should not be an afterthought, because the small models are load-bearing infrastructure.
The parts that stay hard
I would push back on any read of this as a simple fix. Distillation has structural limits that no amount of enthusiasm removes.
- Students inherit teacher blind spots. If the teacher is systematically wrong about a class of problems, the student will be too, and with more confidence, because the errors arrive as clean training signal rather than noisy web text.
- Long-horizon behavior does not compress cleanly. A model can learn to imitate individual reasoning steps and still lose coherence over a fifty-step task. Trajectory-level credit assignment remains an unsolved problem, not a distillation problem.
- Distillation is downstream by construction. A distilled ecosystem is capped by its teachers. If domestic frontier progress slows, everything trained from it slows with it.
- Terms of service and licensing sit on top of all of this. Whether a lab may legally distill from a given frontier model is a contractual question, and the answer shapes who can participate.
The architectural reading
Strip away the geopolitics and Tan is describing a supply chain problem. Frontier models are refineries. Open-weight models are the distributed components that most production systems actually run on. A country with excellent refineries and a thin supply of good small models has a brittle stack, because the layer doing the most requests per second is the layer it does not control.
Distillation is the transfer mechanism between those two layers. Treating it as a shortcut, or as something faintly illegitimate, misreads what it is: a compression technique that turns expensive judgment into affordable judgment. For anyone building agents rather than demos, that conversion rate is the number that decides what ships.
🕒 Published: