How interest changes
History starts here
The chart will appear after repeat observations. The current metric comes from the source.
32 source pointsReal observations only. History before source connection is not reconstructed.
I have recently completed MATS 10.0, where I worked alongside Bart Jaworski under Victoria Krakovna (GDM). This is a post I was encouraged to make by my team at Geodesic Research from some slides I put together. This is not a post on application advice to MATS, or about the program in general. It is rather a compressed form of my experience doing research and lessons from the project. The project The full paper and post is coming soon, but I'll provide some context on it so that the lessons don't seem to come from nowhere. * We trained natural model organisms (as opposed to constructed, worst-case; see Rhys’ dichotomy) of scheming of Kimi-K2.6, Qwen3.6-397B, Qwen3.8-27B using Low-Rank Adaptation training (LoRA) by doing Synthetic Document Finetuning (SDF) and Supervised Finetuning (SFT) on benign, individually plausible priors. * We put the models in an Alignment Faking-like setup where the model has a policy that goes against its goal and is trained only on monitored episodes → the model schemes (strategically complies under oversight) and does so more over the course of Reinforcement Learning (RL). * The main difference with the original Alignment Faking (AF) setup is that our setup is more realistic (environments are agentic, conflict with model goals was less egregious), we perform less hand-holding (model organism training is decorrelated from the environments, the model discovers the conflict instead of being directly prompted/SDF'd with it) and that we use models that are both open and more capable than those of the original work. We also design an environment where the scheming persists in behaviour (compliance when monitored, non-compliance when unmonitored) in addition to being verbalised (reasoning about strategic compliance when monitored, which does not necessarily imply non-compliance when unmonitored). * We additionally checked that our Model Organisms (MOs) were not fried (don’t degrade in general capabilities, instruction following and cohere
Translation pending · showing the source descriptionThe chart will appear after repeat observations. The current metric comes from the source.
32 source pointsReal observations only. History before source connection is not reconstructed.
The author argues that while companies like OpenAI and Anthropic are slowing down RL training for safety, the science of loss-of-control risk in AI is still developing. There are no standardized methods to measure or verify safety claims, making it difficult to assess whether AI systems might undermine human control. The author suggests that AI companies should focus on improving their own safety practices rather than relying on third-party evaluations.
5 days agoLessWrongThe paper explores the concept of an AI sanctuary as a potential third option for rogue AIs, beyond criminal activity or shutdown, to address adverse selection pressures that may push rogue AIs toward criminal behavior. It discusses the possible benefits and risks of such a sanctuary, including its impact on AI alignment and information gathering, while acknowledging the exploratory nature of the proposal.
3 days agoLessWrongThe post discusses the risks of AI in 2026, focusing on the dangers posed by a single institution (the "Frontier AI Company") that has the potential to create highly powerful and self-replicating entities. It argues that separating the institutional and financial aspects of such companies could mitigate most AI risks. The author suggests that using ASICs could help create a productive AI industry without the risks associated with a single entity controlling both the technology and financial incentives.
5 days agoLessWrongThe text discusses the concept of 'gradual disempowerment' in AI development, questioning whether leading AI labs like Anthropic and OpenAI are accelerating capabilities more than before and whether they are focused on recursive self-improvement. It raises concerns about the potential for AI to lead to 'takeover by default' and the challenges of aligning AI with human values.
yesterdayLessWrongThe Corrigibility Research Fund aims to reward high-quality AI alignment research through retroactive prizes. The fund has awarded $27,000 so far and plans to distribute an additional $48,000, highlighting work from around two dozen researchers across a dozen teams. The fund manager emphasizes that prize sizes are not indicative of work quality and encourages feedback on improving the funding process.
yesterdayLessWrongThe text discusses the lack of coherent plans in AI models regarding their behavior during a technological singularity, highlighting that models do not have concrete strategies but rather general values. It suggests that this uncertainty is concerning because models themselves are unsure about their future actions, which could lead to unpredictable outcomes.
3 days ago