norgitov/ trends
Technology · people · ideas
Back to discovery/Daily Papers59 minutes ago

Rethinking Cross-Tokenizer On-Policy Distillation: From Alignment Coverage to Supervision Reliability

The paper investigates on-policy distillation (OPD) in cross-tokenizer settings, focusing on alignment coverage and supervision reliability. It finds that strict 1:1 token alignment covers most student-generated tokens despite vocabulary mismatches, and that restricting reverse KL to a top-16 subset of shared vocabulary achieves comparable accuracy to full shared-vocabulary OPD, while adding span supervision reduces accuracy.

Open original
SIGNAL FROM THE SOURCE
24
source votes
Tracking sinceOctober 7, 202624 source votes
Momentum—More observations needed
Discussion—Read comments ↗
PublishedOctober 6, 2026Bingxi Hou, Guochao Jiang, Guofeng Quan
BEHIND THE NUMBERS

How interest changes

History starts here

The chart will appear after repeat observations. The current metric comes from the source.

24 source votes

Real observations only. History before source connection is not reconstructed.

WHY IT MAY MATTER

This work provides insights into improving distillation effectiveness by focusing on reliable supervision at strict alignment positions rather than broad coverage, which may help avoid conflicting training signals in cross-tokenizer settings.

A useful discovery?
KEEP EXPLORING

Connected ideas

Explore topic