norgitov/ trends
Technology · people · ideas
Back to discovery/Daily Papers1 hour ago

MEND: RL for Flow Models via Proximal Velocity Matching

MEND is a reinforcement learning method for flow models that uses proximal velocity matching to cap rewards within each prompt group, proposing moves along the reward gradient only when the reward gain exceeds a quadratic displacement cost. It outperforms existing methods like Flow-GRPO, ReFL, and DiffusionNFT with fewer updates and without requiring a KL term or frozen reference model.

Open original
SIGNAL FROM THE SOURCE
1
source votes
Tracking sinceOctober 7, 20261 source votes
Momentum—More observations needed
Discussion—Read comments ↗
PublishedOctober 5, 2026Shreshth Saini, Neil Birkbeck, Yilin Wang
BEHIND THE NUMBERS

How interest changes

History starts here

The chart will appear after repeat observations. The current metric comes from the source.

1 source votes

Real observations only. History before source connection is not reconstructed.

WHY IT MAY MATTER

MEND offers a more efficient and effective approach to reward post-training of flow models by capping rewards and using gradient-based moves with a displacement cost, achieving better performance with fewer updates.

A useful discovery?
KEEP EXPLORING

Connected ideas

Explore topic