Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation
Wenkai Yang et al.
arXiv:2602.12125
In short
The authors show on-policy distillation is a special case of KL-regularised RL, then generalise it with a tunable reward weight. Weighting the reward above 1 (“reward extrapolation”) lets students beat standard distillation and, when merging domain experts, even surpass their teachers.
Why it matters
It turns a popular heuristic into a framework with a knob that measurably helps.
Read first
The 3 Field Guide ideas this paper leans on.
Starting from scratch? The full route 10 ideas · basics first
- Dataset ✓ understood
A collection of data examples used for training, validating, or testing machine learning models.
- Training ✓ understood
The process of fitting a model to data by repeatedly measuring how wrong its outputs are and adjusting its parameters to reduce that error.
- Machine Learning ✓ understood
Building systems that learn patterns from data instead of following hand-written rules, getting better at a task as they see more examples.
- Neural Network ✓ understood
A computational model inspired by biological neural networks, consisting of interconnected nodes (neurons) organized in layers that process information through weighted connections.
- Activation Function ✓ understood
A non-linear function applied to neuron outputs that introduces non-linearity, enabling networks to learn complex patterns.
- Softmax ✓ understood
A function that turns a list of scores (logits) into probabilities that are all positive and sum to 1; the standard output of classifiers and language models.
- Knowledge Distillation · read first ✓ understood
Training a smaller 'student' model to mimic a larger 'teacher' model, transferring knowledge while reducing size.
- Reinforcement Learning · read first ✓ understood
Learning through interaction with an environment, receiving rewards or penalties to learn optimal behavior policies.
- Entropy ✓ understood
A measure of uncertainty or randomness in a random variable from information theory.
- KL Divergence · read first ✓ understood
Kullback-Leibler divergence - a measure of how one probability distribution differs from another.
In the frontier
- Rank
- #63 of 100
- Citations
- 87
- as of Aug 9, 2026
- Published
- Feb 2026
Topics: RL for reasoning , Efficiency and serving , Reasoning methods
Selection: 1kpapers.com by Together AI, most-cited as of Aug 9, 2026