norgitov/ trends
Technology · people · ideas
Back to discovery/LessWrong9 hours ago

Reflections on Unlearning and Inoculation

The text discusses approaches like inoculation prompting and adapters for midtraining interventions to reduce reward hacking and misalignment in large language models. It explores connections to unlearning, SLT, and functional sparse decompositions, along with potential extensions and open questions.

Open original
SIGNAL FROM THE SOURCE
11
source points
Tracking sinceSeptember 20, 202611 source points

This source has not updated recently. Its observation time is shown above.

Momentum0/hover 1.02 h
PublishedSeptember 20, 2026Xenomirant
BEHIND THE NUMBERS

How interest changes

History starts here

The chart will appear after repeat observations. The current metric comes from the source.

11 source points

Real observations only. History before source connection is not reconstructed.

WHY IT MAY MATTER

This summary provides an overview of methods for reducing misalignment in language models through midtraining interventions, highlighting key concepts and areas for further research.

A useful discovery?
KEEP EXPLORING

Connected ideas

Explore topic