norgitov/ trends
Technology · people · ideas
Back to discovery/Daily Papers1 hour ago

TRACE: Rollout-Guided Quantization-Aware Training for FP4 Reinforcement Learning of MoE Language Models

This paper introduces TRACE, an FP4 quantization framework for reinforcement learning (RL) training of Mixture-of-Experts (MoE) language models. TRACE addresses the limitation of existing FP4 RL methods by using rollout-side quantization outcomes to guide training-side FP4 rounding decisions, reducing train-rollout discrepancy. The framework also includes an efficient quantization-information caching scheme to reduce storage and communication overhead. Evaluation on four large-scale MoE language models shows that TRACE achieves RL performance comparable to BF16 rollout with up to 5.4x rollout speedup and strong final FP4 performance compared to post-hoc FP4 quantization of BF16-trained policies.

Open original
SIGNAL FROM THE SOURCE
11
source votes
Tracking sinceOctober 7, 202611 source votes
Momentum—More observations needed
Discussion—Read comments ↗
PublishedOctober 6, 2026Xin Wang, Hao Yu, Zhengyang Zhuge
BEHIND THE NUMBERS

How interest changes

History starts here

The chart will appear after repeat observations. The current metric comes from the source.

11 source votes

Real observations only. History before source connection is not reconstructed.

WHY IT MAY MATTER

This method could be useful for improving the efficiency of reinforcement learning training for large language models by reducing computational and memory overhead while maintaining performance.

A useful discovery?
KEEP EXPLORING

Connected ideas

Explore topic