norgitov/ trends
Technology · people · ideas
Back to discovery/LessWrong10 minutes ago

Research Note: Filtering Subversion-Relevant Information From Pretraining Data Is Feasible

The study explores the feasibility of filtering subversion-relevant information from pretraining data for large language models (LLMs). Researchers pretrained multiple 30B LLMs with and without filtering such information, finding that filtered models retain general capabilities but have significantly reduced knowledge about subversion strategies and defenses.

Open original
SIGNAL FROM THE SOURCE
21
source points
Tracking sinceOctober 5, 202621 source points
Momentum+5.9/hover 1.02 h
PublishedOctober 5, 2026Kyle O’Brien
BEHIND THE NUMBERS

How interest changes

History starts here

The chart will appear after repeat observations. The current metric comes from the source.

21 source points

Real observations only. History before source connection is not reconstructed.

WHY IT MAY MATTER

This approach could help reduce the risk of misaligned AI models by limiting their knowledge of subversion strategies, but further research is needed to confirm its effectiveness in real-world scenarios.

A useful discovery?
KEEP EXPLORING

Connected ideas

Explore topic