norgitov/ trends
Technology · people · ideas
Back to discovery/Daily Papers1 hour ago

Gains and Collapse in On-Policy Distillation: A Reinforcement Learning Perspective

The paper explores on-policy distillation (OPD) in language model post-training, explaining how it can lead to either performance gains or generation collapse. It suggests that OPD amplifies student behaviors favored by the teacher's implicit feedback, and proposes methods like masking unhealthy responses and SFT initialization to mitigate collapse.

Open original
SIGNAL FROM THE SOURCE
16
source votes
Tracking sinceOctober 8, 202616 source votes
Momentum—More observations needed
Discussion—Read comments ↗
PublishedOctober 2, 2026Han Cui, Jianhao Yan, Yun Luo
BEHIND THE NUMBERS

How interest changes

History starts here

The chart will appear after repeat observations. The current metric comes from the source.

16 source votes

Real observations only. History before source connection is not reconstructed.

WHY IT MAY MATTER

This research provides insights into why on-policy distillation can lead to either improved performance or generation issues, helping to guide safer and more effective model training practices.

A useful discovery?
KEEP EXPLORING

Connected ideas

Explore topic