norgitov/ trends
Technology · people · ideas
Back to discovery/Daily Papers1 hour ago

Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening

The paper identifies a problem called Value Flattening in Proximal Policy Optimization (PPO) critics, where estimated state values change sharply while critic predictions remain flat. It introduces SP^3O, a method that applies value loss to only a few states per response, showing improvements in policy learning.

Open original
SIGNAL FROM THE SOURCE
33
source votes
Tracking sinceSeptember 17, 202633 source votes
MomentumMore observations needed
DiscussionRead comments ↗
PublishedSeptember 16, 2026Yizhuo Li, Jianhao Yan, Yun Luo
BEHIND THE NUMBERS

How interest changes

History starts here

The chart will appear after repeat observations. The current metric comes from the source.

33 source votes

Real observations only. History before source connection is not reconstructed.

WHY IT MAY MATTER

This research addresses a critical issue in PPO that affects the accuracy of value estimation, which is essential for effective reinforcement learning in large language models.

A useful discovery?
KEEP EXPLORING

Connected ideas

Explore topic