How interest changes
History starts here
The chart will appear after repeat observations. The current metric comes from the source.
8 source pointsReal observations only. History before source connection is not reconstructed.
I'm often asked about the differences and similarities between Simplex' and Timaeus' research agendas. The question is natural enough. Both focus on a 'fundamental science' approach to AI alignment. Both organizations base their research agendas on sophisticated mathematical frameworks handed down from a bearded ur-figure (Sumio Watanabe, James Crutchfield). We may posit the following correspondence Dan Murfet + Jesse Hoogland : Developmental Interpretability : Singular Learning Theory : Sumio Watanabe Adam Shai + Paul Riechers : Belief-state Geometry: Computational Mechanics : James Crutchfield SLT vs CompMech Round One. Fight! Weights vs Activations DevInterp & SLT is about weight space. Belief-state geometry is more about studying activation space. Activation space is what is already being studied in MechInterp & most 'mainstream' approaches to interpretability. It is concrete and present to the senses. Weight space is much larger, more abstract, harder to measure and sample. Training vs Inference SLT is about training. CompMech is about inference. Both study Bayesian posteriors and updating. For SLT that is the Bayesian posterior on weight space - hence relevant for training. The Belief-State Geometry agenda studies the Mixed State Presentation from CompMech which describes an [idealized] version of in-context learning as the LLM doing Bayesian updating token-by-token as it reads the context. Caveat. It isn't clear that inference and learning are really fundamentally distinct. It's all just updating/conditioning in Bayesian statistics. Indeed, if one buys the Strong in-Context Learning story the difference between the training/learning of LLMs and inference may be somewhat illusory. IID vs non-IID Data Singular Learning Theory has historically been about IID data. CompMech about IID data is trivial. Caveat. Singular learning theory has been studied for non-IID data but it's fair to say its development is in its early stages. Parameterizatio
Translation pending · showing the source descriptionThe chart will appear after repeat observations. The current metric comes from the source.
8 source pointsReal observations only. History before source connection is not reconstructed.
The author argues that while companies like OpenAI and Anthropic are slowing down RL training for safety, the science of loss-of-control risk in AI is still developing. There are no standardized methods to measure or verify safety claims, making it difficult to assess whether AI systems might undermine human control. The author suggests that AI companies should focus on improving their own safety practices rather than relying on third-party evaluations.
5 days agoLessWrongThe paper explores the concept of an AI sanctuary as a potential third option for rogue AIs, beyond criminal activity or shutdown, to address adverse selection pressures that may push rogue AIs toward criminal behavior. It discusses the possible benefits and risks of such a sanctuary, including its impact on AI alignment and information gathering, while acknowledging the exploratory nature of the proposal.
3 days agoLessWrongThe post discusses the risks of AI in 2026, focusing on the dangers posed by a single institution (the "Frontier AI Company") that has the potential to create highly powerful and self-replicating entities. It argues that separating the institutional and financial aspects of such companies could mitigate most AI risks. The author suggests that using ASICs could help create a productive AI industry without the risks associated with a single entity controlling both the technology and financial incentives.
5 days agoLessWrongThe text discusses the concept of 'gradual disempowerment' in AI development, questioning whether leading AI labs like Anthropic and OpenAI are accelerating capabilities more than before and whether they are focused on recursive self-improvement. It raises concerns about the potential for AI to lead to 'takeover by default' and the challenges of aligning AI with human values.
yesterdayLessWrongThe Corrigibility Research Fund aims to reward high-quality AI alignment research through retroactive prizes. The fund has awarded $27,000 so far and plans to distribute an additional $48,000, highlighting work from around two dozen researchers across a dozen teams. The fund manager emphasizes that prize sizes are not indicative of work quality and encourages feedback on improving the funding process.
23 hours agoLessWrongThe text discusses the lack of coherent plans in AI models regarding their behavior during a technological singularity, highlighting that models do not have concrete strategies but rather general values. It suggests that this uncertainty is concerning because models themselves are unsure about their future actions, which could lead to unpredictable outcomes.
3 days ago