norgitov/ trends
Technology · people · ideas
Back to discovery/Daily Papers2 hours ago

1% of Tokens Can Be Enough: On Gradient Estimation in On-Policy Distillation

The paper explores sparse on-policy distillation (OPD) by focusing on gradient estimation challenges when allocating teacher supervision to a small subset of tokens. It introduces an information-efficiency ratio (IER) to assess gradient estimation reliability and combines it with existing usefulness scores for better token selection, achieving improved performance with minimal token budgets.

Open original
SIGNAL FROM THE SOURCE
2
source votes
Tracking sinceSeptember 22, 20262 source votes
MomentumMore observations needed
DiscussionRead comments ↗
PublishedSeptember 21, 2026Huanxin Sheng, Zhiling Ye, Haonan Wang
BEHIND THE NUMBERS

How interest changes

History starts here

The chart will appear after repeat observations. The current metric comes from the source.

2 source votes

Real observations only. History before source connection is not reconstructed.

WHY IT MAY MATTER

This research provides a method to efficiently allocate teacher supervision in on-policy distillation by considering both the usefulness of tokens and the reliability of gradient estimation, which can be useful in scenarios with limited computational resources or data constraints.

A useful discovery?
KEEP EXPLORING

Connected ideas

Explore topic