norgitov/ trends
Technology · people · ideas
Back to discovery/LessWrong9 minutes ago

Inoculation Midtraining with Learned Neologisms

The paper introduces Inoculation Midtraining, a technique that teaches a base model to associate unsafe behavior with a specific context marked by a new token during midtraining, then evaluates the model's generalization of this alignment outside the context. The method shows some success in SFT and on-policy RL but faces challenges like sensitivity to hyperparameters and conditional misalignment.

Open original
SIGNAL FROM THE SOURCE
38
source points
Tracking sinceSeptember 15, 202638 source points
Momentum+23.61/hover 1.02 h
PublishedSeptember 15, 2026Kyle O’Brien
BEHIND THE NUMBERS

How interest changes

History starts here

The chart will appear after repeat observations. The current metric comes from the source.

38 source points

Real observations only. History before source connection is not reconstructed.

WHY IT MAY MATTER

This method could help in guiding post-training-induced misalignment through base model data curation, but it's not yet production-ready.

A useful discovery?
KEEP EXPLORING

Connected ideas

Explore topic