News · 2026-10-03
Study finds the choice of loss and learning rate matter more than on-policy data when distilling a model
A controlled study of how small AI models learn from larger ones finds that a widely recommended technique, on-policy distillation, matters less than two more mundane settings: which direction the training loss measures the gap between student and teacher, and the learning rate. The paper, from Julianna Piskorz, Antonin Berthon and Mihaela van der Schaar, was submitted to arXiv on September 28 and ranked second among Hugging Face's papers of the day when checked on October 2.
Key facts
- Finding: the loss direction "more clearly shapes task performance," while the learning rate "governs forgetting and update sparsity," the authors write.
- When: submitted September 28, 2026.
- Who: Julianna Piskorz, Antonin Berthon and Mihaela van der Schaar.
- Primary source: "On-Policy or Off-Policy Learning? A Systematic Study of Distillation Dynamics".
Distillation trains a small "student" model to imitate a larger "teacher." There are two broad ways to pick the examples. Off-policy, the student studies text the teacher wrote. On-policy, the student writes its own attempts and the teacher grades each word. On-policy learning has become fashionable, with claims that it reduces catastrophic forgetting, produces sparser changes to the model, and generalizes better. Ground Truth covered a paper on how on-policy distillation scales just two days ago.
The trouble, the authors argue, is that earlier comparisons changed several things at once. "Existing comparisons between supervised fine-tuning and reinforcement learning vary many factors simultaneously, making the contribution of rollout policy difficult to isolate," they write.
What they did
The team varied three settings independently: whose text the student learns from (its own, the teacher's, or a blend along a continuous spectrum), the direction of the loss, and the learning rate. They ran this across Llama 3 and Qwen2.5 student models on reasoning tasks in science, medicine and arithmetic.
The loss direction is the subtle part. Training compares the student's word probabilities to the teacher's using a measure called KL divergence, which is lopsided: measured one way ("forward"), it punishes the student for ignoring anything the teacher considers likely; measured the other way ("reverse"), it punishes the student for confidently saying things the teacher would not. Our lesson on forward and reverse KL divergence explains the difference. An analogy: forward KL is a teacher who marks you down for every topic you skipped; reverse KL is one who marks you down for every wrong claim you made.
What they found
"Forward KL is remarkably robust to rollout policy, with its performance stable and strong despite changes to the rollout policy, whereas reverse KL is substantially more sensitive and favours student-generated rollouts," the authors report. The math explains why: the reverse measure reacts sharply when the student assigns almost no probability to a word the teacher likes, which is exactly what changes when you switch whose text you train on. Meanwhile the learning rate, not the data source, drove most of the differences in forgetting and update sparsity.
On-policy data was not useless. It improved generalization to harder versions of the Countdown arithmetic puzzle under both loss directions. But that edge "does not reliably persist" after a further round of reinforcement learning with verifiable rewards, where some off-policy-trained checkpoints caught up and passed it. The conclusions held when the team removed gradient clipping, used sampled estimates of the loss, and trained on longer reasoning chains.
Why it matters
On-policy distillation is more expensive, because the student has to generate fresh text throughout training. This study gives teams a reason to check whether they need it before paying for it, and a warning that a gain credited to on-policy data might really come from the loss or learning rate that happened to accompany it. "Our results challenge the view that on-policy rollouts are inherently preferable," the authors conclude. Our lesson on on-policy versus off-policy learning covers the general distinction.
The caveat
The student models were small, at most 1.5 billion parameters, and reasoning was capped at 2,000 tokens, so the authors note larger models and longer reasoning remain open questions. The tasks are bounded reasoning problems, not broad assistant behavior. A companion paper, Neighborhood On-Policy Self-Distillation, reports gains from on-policy methods by improving the teacher's supervision instead, so the two results are about different questions rather than a direct contradiction. These are preprint results without independent replication.
Key questions
Is on-policy distillation better than off-policy?
Which training setting most affected forgetting?
How big were the models tested?
Cite this
APA
Ground Truth. (2026, October 3). Study finds the choice of loss and learning rate matter more than on-policy data when distilling a model. Ground Truth. https://groundtruth.day/news/on-policy-distillation-matters-less-than-the-loss-and-learning-rate.html
BibTeX
@misc{groundtruth:on-policy-distillation-matters-less-than-the-loss-and-learning-rate,
title = {Study finds the choice of loss and learning rate matter more than on-policy data when distilling a model},
author = {{Ground Truth}},
year = {2026},
month = {oct},
url = {https://groundtruth.day/news/on-policy-distillation-matters-less-than-the-loss-and-learning-rate.html}
}
Comments are replies to this story on Bluesky — reply with any Bluesky account to join in.