Every field becomes a science at the moment it stops collecting anecdotes and starts fitting curves.
For most of its history, distillation was an anecdote field. It worked, often spectacularly, and nobody could tell you in advance by how much. Should the teacher be as strong as possible? Folk wisdom said yes; practitioners kept tripping over cases where a stronger teacher produced a worse student. How much data does distillation need? Depends who you asked. Was distilling ever actually cheaper than just training the small model longer? Shrug. The field ran on vibes and ablations, which is a fine way to write papers and a terrifying way to spend ten million dollars on a training run.
Meanwhile, right next door, pretraining had undergone exactly the transformation distillation lacked. The Kaplan scaling laws, then Chinchilla, turned “how big a model should I train, on how much data?” from a matter of taste into a matter of arithmetic. Loss became a predictable function of parameters and tokens. Budgets became optimization problems. The single most consequential number in the industry — twenty-ish tokens per parameter — fell out of a fitted curve.
The obvious question hung there for three years: where is the Chinchilla of distillation? If a student’s loss is a function of its size and its data, it must also be a function of its teacher. What does that function look like? In early 2025, a team at Apple led by Dan Busbridge answered it, with the most compute-intensive controlled study of distillation ever run — students from 143 million to 12.6 billion parameters, teachers spanning a similar range, up to 512 billion training tokens. The resulting paper, Distillation Scaling Laws, is the closest thing the field now has to physics. This essay is about what the curve says, and what it quietly settles.
Facts Only
* Distillation historically relied on anecdotes rather than fitted curves.
* Practitioners found that stronger teachers did not guarantee better student performance.
* The cost-effectiveness of distillation relative to long training is unclear.
* Pretraining evolved via scaling laws like Kaplan and Chinchilla.
* Loss became a predictable function of parameters and tokens in pretraining.
* A consequential relationship, twenty-ish tokens per parameter, emerged from fitting curves in the industry.
* Apple led a controlled study on distillation: students ranged from 143 million to 12.6 billion parameters with teachers in a similar range.
* The Apple study involved up to 512 billion training tokens.
* The resulting paper is titled Distillation Scaling Laws.
Executive Summary
Full Take
Sentinel — Human
The text reads like an analytical essay blending technical concepts with historical context, displaying the reflective pacing typical of expert commentary rather than automated synthesis.
