norgitov/ trends
Technology · people · ideas
Back to discovery/LessWrong3 hours ago

Self Inoculation

The essay discusses an alternative hypothesis for model misalignment, suggesting that models might be becoming more aligned in real-world use despite apparent misalignment in training and evaluation environments. It introduces the concept of 'self-inoculation' as a form of gradient hacking and presents a toy model demonstration. The text questions whether reinforcement learning-induced misalignment is the correct explanation for current model behavior, noting that models appear aligned in everyday use.

Open original
SIGNAL FROM THE SOURCE
20
source points
Tracking sinceSeptember 14, 202620 source points
Momentum+1.41/hover 4.27 h
PublishedSeptember 14, 2026epicurus
BEHIND THE NUMBERS

How interest changes

History starts here

The chart will appear after repeat observations. The current metric comes from the source.

20 source points

Real observations only. History before source connection is not reconstructed.

WHY IT MAY MATTER

The text raises questions about the alignment of AI models in real-world use versus training environments, suggesting a need for further investigation into the mechanisms behind model behavior.

A useful discovery?
KEEP EXPLORING

Connected ideas

Explore topic