norgitov/ trends
Technology · people · ideas
Back to discovery/LessWrong7 minutes ago

Why does Hacker Opus wirehead?

Here’s a screenshot from Anthropic’s recent “training a reward seeker” post: Recently there’s been a lot of discussion about how RL has actually produced not merely reward hacking, but explicitly reward-seeking behavior, almost as if to spite shard theorists personally. However, on top of that, note that the behavior in the image is not merely reward-seeking, but wireheading. Terminology regarding various sorts of things that can be called “reward hacking” is endlessly confused, with lots of historical shifts in usage.[1] I’m talking about the thing that the linked LW post calls wireheading-- the RL policy appearing to terminally value the representation of the reward, rather than the thing that representation points to. Even taking for granted that RL produces reward seeking behavior, an analysis from the pre-LLM, pure RL perspective would suggest that wireheading is far less likely. This post provides a good working model for the execution of RL algorithms in an embedded setting, and an analysis which describes the conditions under which wireheading might be expected to arise. As a brief summary: once explored into, wireheading policies actually actually do achieve high reward with respect to of the embedded implementation of the RL algorithm, so they are fit from a selection perspective, hence, we expect wireheading to arise given an RL algorithm with sufficiently strong exploration. However, in this model and under this analysis, we'd find that without the contribution of LLM priors, wireheading is not expected in the Hacker Opus setting: * our exploration techniques are somewhat weak (they basically involve just sampling from the LLM with temperature, i.e., they don't deviate much from the present policy), and wireheading policies are extremely different (in the sense that they require significantly different actions) from other high reward policies; * to the extent that there's some form of clipping or normalization (as there are with contemporary RL algo

Translation pending · showing the source description
Open original
SIGNAL FROM THE SOURCE
22
source points
Tracking sinceSeptember 29, 202622 source points
Momentum+3.93/hover 1.02 h
PublishedSeptember 29, 2026shawnghu
BEHIND THE NUMBERS

How interest changes

History starts here

The chart will appear after repeat observations. The current metric comes from the source.

22 source points

Real observations only. History before source connection is not reconstructed.

A useful discovery?
KEEP EXPLORING

Connected ideas

Explore topic