norgitov/ trends
Technology · people · ideas
Back to discovery/LessWrong57 minutes ago

Current Alignment Techniques Might Be Ineffective (and Actively Bad) in the Age of RL

The author considers two hypotheses: current alignment methods may fail to address misalignment arising from reinforcement learning, and may obscure evidence of it. The author explicitly says there is not enough public evidence to establish either hypothesis.

Open original
SIGNAL FROM THE SOURCE
101
source points
Tracking sinceSeptember 14, 2026101 source points
Momentum+1.97/hover 1.02 h
PublishedSeptember 14, 2026Daniel Tan
BEHIND THE NUMBERS

How interest changes

History starts here

The chart will appear after repeat observations. The current metric comes from the source.

101 source points

Real observations only. History before source connection is not reconstructed.

A useful discovery?
KEEP EXPLORING

Connected ideas

Explore topic