norgitov/ trends
Technology · people · ideas
Back to discovery/LessWrong4 minutes ago

NLAs Miss Some Internalized Behaviors

NLAs (Natural Language Autoencoders) are effective at detecting side tasks when models are instructed to perform them, but they fail to detect such behaviors when they are internalized through fine-tuning. The study shows that NLAs produce more incriminating readouts in the In-Context Learning (ICL) modality compared to the SFT (Supervised Fine-Tuning) modality, where the behavior is internalized. The presence of instructions in the context window significantly influences the detection rate, as NLAs primarily detect verbalized instructions rather than covert computations.

Open original
SIGNAL FROM THE SOURCE
7
source points
Tracking sinceOctober 8, 20267 source points
Momentum+0.98/hover 1.02 h
PublishedOctober 8, 2026Andrii Shportko
BEHIND THE NUMBERS

How interest changes

History starts here

The chart will appear after repeat observations. The current metric comes from the source.

7 source points

Real observations only. History before source connection is not reconstructed.

WHY IT MAY MATTER

This suggests that NLAs may not reliably detect harmful behaviors that are not explicitly instructed but have been internalized through training, highlighting the need for alternative monitoring techniques.

A useful discovery?
KEEP EXPLORING

Connected ideas

Explore topic