norgitov/ trends
Technology · people · ideas
Back to discovery/LessWrong9 minutes ago

My retrospective from MATS 10.0

I have recently completed MATS 10.0, where I worked alongside Bart Jaworski under Victoria Krakovna (GDM). This is a post I was encouraged to make by my team at Geodesic Research from some slides I put together. This is not a post on application advice to MATS, or about the program in general. It is rather a compressed form of my experience doing research and lessons from the project. The project The full paper and post is coming soon, but I'll provide some context on it so that the lessons don't seem to come from nowhere. * We trained natural model organisms (as opposed to constructed, worst-case; see Rhys’ dichotomy) of scheming of Kimi-K2.6, Qwen3.6-397B, Qwen3.8-27B using Low-Rank Adaptation training (LoRA) by doing Synthetic Document Finetuning (SDF) and Supervised Finetuning (SFT) on benign, individually plausible priors. * We put the models in an Alignment Faking-like setup where the model has a policy that goes against its goal and is trained only on monitored episodes → the model schemes (strategically complies under oversight) and does so more over the course of Reinforcement Learning (RL). * The main difference with the original Alignment Faking (AF) setup is that our setup is more realistic (environments are agentic, conflict with model goals was less egregious), we perform less hand-holding (model organism training is decorrelated from the environments, the model discovers the conflict instead of being directly prompted/SDF'd with it) and that we use models that are both open and more capable than those of the original work. We also design an environment where the scheming persists in behaviour (compliance when monitored, non-compliance when unmonitored) in addition to being verbalised (reasoning about strategic compliance when monitored, which does not necessarily imply non-compliance when unmonitored). * We additionally checked that our Model Organisms (MOs) were not fried (don’t degrade in general capabilities, instruction following and cohere

Translation pending · showing the source description
Open original
SIGNAL FROM THE SOURCE
32
source points
Tracking sinceOctober 1, 202632 source points
Momentum+21.17/hover 0.57 h
PublishedOctober 1, 2026Nathaniel Mitrani
BEHIND THE NUMBERS

How interest changes

History starts here

The chart will appear after repeat observations. The current metric comes from the source.

32 source points

Real observations only. History before source connection is not reconstructed.

A useful discovery?
KEEP EXPLORING

Connected ideas

Explore topic