Garry Tan’s call for American open-weight labs to distill frontier models is the most pragmatic proposal in US open-source AI right now, and the fact that it sounds like an admission of defeat is precisely why it deserves to be taken seriously. The Y Combinator chief is asking smaller American labs to do something researchers have quietly known for years: you do not need frontier-scale compute to get frontier-adjacent capability. You need a good teacher, a disciplined student, and the humility to copy before you create. Tan frames this as building a strong alternative to Chinese open-weight dominance, with real stakes for American AI research by 2026. As someone who spends her days inside model internals, I want to explain why the technique itself justifies his optimism, and where it does not.
What distillation actually buys you
Distillation is often described as compression, but that undersells it. When a smaller student model trains on the outputs of a larger teacher, it is not merely shrinking weights. It is inheriting the results of the teacher’s entire training pipeline: the data curation decisions, the reinforcement learning passes, the thousands of failed experiments that shaped the final behavior. All of that accumulated search cost gets flattened into a supervision signal the student can learn from directly.
That is why distillation is such an efficient transfer mechanism. The expensive part of frontier AI is not the final training run. It is the exploration that precedes it. A student model skips the exploration and buys the answer key. For a small lab with limited compute, this is the difference between being years behind and being months behind.
Fast-follower economics, formalized
Tan’s
🕒 Published: