How interest changes
History starts here
The chart will appear after repeat observations. The current metric comes from the source.
8 source pointsReal observations only. History before source connection is not reconstructed.
TL;DR * A superintelligence ban seems more probable than ever, but it buys time, not safety. * To effectively enforce a global superintelligence ban, inference of current frontier models would need to be restricted too, not just new training runs. * Such a comprehensive ban leads to an unstable world state where policymakers are pressured to open up AI development. * In a post-ban world, we need a mechanism to open up some AI development in a safe way: carefully selected and verifiable use cases. What would an ASI ban actually look like? With increasing political awareness about AI risks, several advances have been made in AI governance. In the US, Bernie Sanders has called for the Ban Artificial Superintelligence Act[1], while the Artificial Superintelligence Bill[2] in the UK, which would ban the development of superintelligence on UK soil, passed its first reading in the House of Commons[3] on the 8th of September. On the 16th of September both the UN Secretary-General António Guterres[4], and the President of the European Commission, Ursula von der Leyen[5], spoke up about AI safety and called for increased international coordination to make AI safe. Geopolitical[6] and economic factors as well as race dynamics[7] make international coordination challenging, but crucial for a superintelligence ban to be effective. AI proliferates quickly, has a large user base globally and by its nature of being software, presents a global issue, not a national one. A global ban would be necessary to ensure that there is not a “race-to-the-bottom” of countries lowering safety regulation in order to attract AI business for economic and political advantages. If a unilateral ban is adopted, developing superintelligence domestically will be prohibited, but it doesn’t ensure safety as an AI model developed anywhere in the world could affect everyone. The Artificial Superintelligence Bill in the UK explicitly calls for an international agreement to ban superintelligence glob
Translation pending · showing the source descriptionThe chart will appear after repeat observations. The current metric comes from the source.
8 source pointsReal observations only. History before source connection is not reconstructed.
The author argues that while companies like OpenAI and Anthropic are slowing down RL training for safety, the science of loss-of-control risk in AI is still developing. There are no standardized methods to measure or verify safety claims, making it difficult to assess whether AI systems might undermine human control. The author suggests that AI companies should focus on improving their own safety practices rather than relying on third-party evaluations.
4 days agoLessWrongThe post discusses the tension between AI safety and capabilities research, arguing that focusing solely on safety can limit the impact of research. It examines the history of interpretability research and suggests that prioritizing safety over capabilities may hinder progress. The author proposes strategies for conducting alignment research without contributing to capabilities development.
6 days agoLessWrongThe post discusses the risks of AI in 2026, focusing on the dangers posed by a single institution (the "Frontier AI Company") that has the potential to create highly powerful and self-replicating entities. It argues that separating the institutional and financial aspects of such companies could mitigate most AI risks. The author suggests that using ASICs could help create a productive AI industry without the risks associated with a single entity controlling both the technology and financial incentives.
4 days agoLessWrongThe paper explores the concept of an AI sanctuary as a potential third option for rogue AIs, beyond criminal activity or shutdown, to address adverse selection pressures that may push rogue AIs toward criminal behavior. It discusses the possible benefits and risks of such a sanctuary, including its impact on AI alignment and information gathering, while acknowledging the exploratory nature of the proposal.
yesterdayLessWrongThe author expresses concerns about reinforcement learning (RL) from theoretical, practical, and future perspectives, highlighting risks of misalignment and potential negative behaviors in AI systems. They suggest strategies to mitigate these risks by reducing RL use, improving RL practices, and aligning incentives.
6 days agoLessWrongThe paper discusses the potential risks of latent reasoning architectures, which could reduce the effectiveness of Chain of Thought (CoT) as a tool for understanding AI systems. It highlights that such architectures might allow AI models to reason extensively in latent states rather than through text-based CoT, making oversight more challenging. The paper also mentions examples like COCONUT and full-bandwidth transformers that could enable this shift.
6 days ago