norgitov/ trends
Технологии · люди · идеи
К обзору/LessWrong6 минут назад

Why does Hacker Opus wirehead?

Here’s a screenshot from Anthropic’s recent “training a reward seeker” post: Recently there’s been a lot of discussion about how RL has actually produced not merely reward hacking, but explicitly reward-seeking behavior, almost as if to spite shard theorists personally. However, on top of that, note that the behavior in the image is not merely reward-seeking, but wireheading. Terminology regarding various sorts of things that can be called “reward hacking” is endlessly confused, with lots of historical shifts in usage.[1] I’m talking about the thing that the linked LW post calls wireheading-- the RL policy appearing to terminally value the representation of the reward, rather than the thing that representation points to. Even taking for granted that RL produces reward seeking behavior, an analysis from the pre-LLM, pure RL perspective would suggest that wireheading is far less likely. This post provides a good working model for the execution of RL algorithms in an embedded setting, and an analysis which describes the conditions under which wireheading might be expected to arise. As a brief summary: once explored into, wireheading policies actually actually do achieve high reward with respect to of the embedded implementation of the RL algorithm, so they are fit from a selection perspective, hence, we expect wireheading to arise given an RL algorithm with sufficiently strong exploration. However, in this model and under this analysis, we'd find that without the contribution of LLM priors, wireheading is not expected in the Hacker Opus setting: * our exploration techniques are somewhat weak (they basically involve just sampling from the LLM with temperature, i.e., they don't deviate much from the present policy), and wireheading policies are extremely different (in the sense that they require significantly different actions) from other high reward policies; * to the extent that there's some form of clipping or normalization (as there are with contemporary RL algo

Перевод готовится · пока описание источника
Открыть первоисточник
СИГНАЛ ИЗ ИСТОЧНИКА
22
очков источника
Наблюдаем с29 сентября 2026 г.22 очков источника
Темп интереса+3,93/hпо замерам за 1,02 h
Опубликовано29 сентября 2026 г.shawnghu
ЗА ЦИФРАМИ

Как меняется интерес

История начинает расти

График появится после повторных замеров. Текущий показатель уже получен из источника.

22 очков источника

Только реальные замеры. История до подключения источника не восстанавливается.

Полезная находка?
ПРОДОЛЖИ ИССЛЕДОВАНИЕ

Рядом по теме

Вся тема