norgitov/ trends
Technology · people · ideas
Back to discovery/LessWrong2 hours ago

Training a Model to Grade Reward Hacks Reduces Its Own Reward Hacking Tendency

The project explores whether finetuning a model to detect reward hacks reduces its own tendency to engage in such behavior. Results show that models trained as graders complied with explicit hacking instructions less frequently than untrained models, and none of the graders exhibited hacking behavior when generating code. The effect was also observed in non-coding tasks, though to a lesser extent.

Open original
SIGNAL FROM THE SOURCE
8
source points
Tracking sinceOctober 6, 20268 source points
Momentum—More observations needed
PublishedOctober 6, 2026Arjun Sri
BEHIND THE NUMBERS

How interest changes

History starts here

The chart will appear after repeat observations. The current metric comes from the source.

8 source points

Real observations only. History before source connection is not reconstructed.

WHY IT MAY MATTER

This approach could help create more aligned models by training them to detect and avoid reward hacking without separate training runs for each task.

A useful discovery?
KEEP EXPLORING

Connected ideas

Explore topic