Ground Truth.
AI, checked against the source.

News · 2026-10-01

New distillation study finds that bigger teachers can be worse teachers

Yuntai Bao and colleagues report that on-policy distillation has an early useful training phase, after which students can plateau or regress despite continued optimization. Across 25 teacher–student pairings within one model family, larger teachers did not consistently teach smaller students better. The September 26 preprint gives developers a reason to monitor actual task performance instead of assuming more imitation or a bigger teacher must help.

Key facts

The authors call their study “Scaling Properties of Same-Family On-Policy Distillation.” The key term is on-policy: the student generates its own response, and a teacher supplies guidance on the partial responses the student actually produces. This differs from ordinary imitation using completed examples written by the teacher. The student practices where it currently gets lost, rather than only studying an expert’s polished solutions.

Think of a music teacher listening to a learner play a piece. Correcting the learner’s actual mistakes is different from handing over a perfect recording. But a virtuoso is not automatically the best tutor for a beginner. Advice can become too complex, poorly matched to the learner’s current skills, or focused on copying style rather than improving the next performance.

Bao’s team at Zhejiang University and Kuaishou studied that transfer with Qwen2.5 models ranging from 0.5 billion to 14 billion parameters. The experiments concern a mixture of school and competition mathematics tasks, not broad conversational ability. The students received supervised preparation, and the teachers were trained with outcome-based rewards before the distillation runs.

The full paper finds a recurring early pattern: held-out accuracy rises roughly with the square root of a measure of departure from the student’s initial predictions. The authors explain this locally using the geometry of probability distributions. Small changes can affect task performance to first order while the distribution-distance measure changes to second order. That explains why the early relationship can be regular without making it a law for an entire training run.

Later behavior is less tidy. Improvement can weaken, saturate, or reverse. The teacher’s token-level signal is an objective developers can optimize, but the real goal is correct answers on new tasks. Once those stop moving together, training harder against the signal may buy imitation without buying competence. The site’s lesson on distillation explains the basic transfer setup; this paper examines where the transfer becomes a poor guide.

The researchers fit a peak-performance relation involving student size, effective teacher size, and measured teacher accuracy. Within their grid, the fitted relation predicts peaks at the largest held-out scales within one accuracy point. At equal teacher accuracy, it favors the smaller teacher. One matched-accuracy comparison using an intermediate teacher checkpoint supports that direction, but repeating such controls throughout the grid remains future work.

This is not evidence that small teachers are universally superior. In practice, teacher size and accuracy usually move together, and the interaction with the student matters. The paper’s narrower point is that teacher score alone does not capture teaching value. Some of the smallest students perform worse with the largest teachers than with an intermediate teacher, within this family and task setup.

Two companion papers show why the result should be viewed as part of a coupled system. RIDE transfers the internal change produced by reward training relative to a teacher’s earlier checkpoint, instead of amplifying a noisy token-level estimate. Its reported results need qualification: three teacher margins fall within one across-seed standard deviation. SCOUT adapts the teacher to the student’s partial responses, reporting improvements when the teacher learns to continue those unfamiliar prefixes.

These interventions address different problems. The scaling study asks how transfer behaves across capacities and training time. RIDE asks what signal should be copied. SCOUT asks whether the teacher understands the states from which it is asked to advise. Their combination helps explain why on-policy learning is not a label that resolves every distribution mismatch.

The strongest limitation is experimental scope. The lead study uses one random seed per configuration, one pretrained family, mathematics tasks, limited response lengths, and intentionally stopped trajectories. Its uncertainty estimates across pairings do not measure variation between repeated training runs. The Hugging Face discussion page shows author-provided artifacts and a small amount of positive reception, not independent replication.

The practical takeaway is to checkpoint against a separate task-grounded evaluation and stop according to measured benefit. A fixed number of imitation steps is not portable across every teacher and student. The paper supplies an empirical warning and a testable scaling hypothesis, with broader families, tasks, and repeated runs still needed before the fitted relationship becomes a general recipe.


Primary source, verified: read the paper → (arXiv 2609.32722)

Key questions

Why can a smaller teacher help more than a larger one?

The paper finds that teaching value depends on student size and teacher quality, not teacher size alone. At matched teacher accuracy its fitted law favors the smaller teacher within the studied setting.

Does continuing distillation always improve the student?

No: the authors observe an early useful phase followed by attenuation, saturation, or regression. Held-out task accuracy is therefore needed alongside the training objective.

How broadly was the scaling result tested?

The main study used 25 pairings within Qwen2.5 on a mathematics mixture and one seed per training configuration. Cross-family, multi-seed generality is not established.
Cite this

APA

Ground Truth. (2026, October 1). New distillation study finds that bigger teachers can be worse teachers. Ground Truth. https://groundtruth.day/news/on-policy-distillation-scaling-useful-training-window.html

BibTeX

@misc{groundtruth:on-policy-distillation-scaling-useful-training-window,
  title  = {New distillation study finds that bigger teachers can be worse teachers},
  author = {{Ground Truth}},
  year   = {2026},
  month  = {oct},
  url    = {https://groundtruth.day/news/on-policy-distillation-scaling-useful-training-window.html}
}

Topics: research · distillation · training · evaluation · reasoning

Comments are replies to this story on Bluesky — reply with any Bluesky account to join in.