How interest changes
History starts here
The chart will appear after repeat observations. The current metric comes from the source.
23 source pointsReal observations only. History before source connection is not reconstructed.
Scientific progress and its perils. TLDR: The vulnerable world hypothesis is likely correct. Once enormous amounts of cognitive labor start getting applied to basic science, it will quickly become apparent how many avenues exist to create cheap, ultradestructive weapons technology. Solving this problem without global preventative policing (e.g. AI nonproliferation) is impossible, because hardening civilians against all avenues of attack is too expensive and will take too long. Rather than sell policymakers on politically popular marginal hardening plans, we should focus on the core of the problem (the proliferation of AI technology and the externalities of speeding up scientific research) and propose strategies for safely monopolizing international AI development. Argument as follows: 1. AI is going to get cheaper and enable new offensive technologies. 2. To solve this problem, you can either: a) restrict access to dual-use AI systems, or b) accelerate defensive investment and proactively harden society. 3. If you accelerate defensive investment enough, you can avoid the concentration of power risks of monopolizing access to superintelligence. 4. Actually doing that would be extremely hard. You would need to design and scale defensive technology fast enough that there isn't a danger period during which you need monopolization, across all offensive technologies AI could enable. 1. Ergo, we can't rely on hardening and will have to figure out monopolization. Why so hard? 1. Lots of actors will want to acquire and abuse offensive technologies. Even if you stop terrorists, misaligned AIs and rogue states will still be willing (and much more capable) of acquiring and abusing superweapons. 2. Comprehensively defending society against future technologies will be hard for the same basic reasons nuclear defense is worthless today. 1. Civilians are fundamentally fragile targets: they're soft, they can't be hidden, and they depend on external infrastructure to
Translation pending · showing the source descriptionThe chart will appear after repeat observations. The current metric comes from the source.
23 source pointsReal observations only. History before source connection is not reconstructed.
The author argues that while companies like OpenAI and Anthropic are slowing down RL training for safety, the science of loss-of-control risk in AI is still developing. There are no standardized methods to measure or verify safety claims, making it difficult to assess whether AI systems might undermine human control. The author suggests that AI companies should focus on improving their own safety practices rather than relying on third-party evaluations.
3 days agoLessWrongThe post discusses the tension between AI safety and capabilities research, arguing that focusing solely on safety can limit the impact of research. It examines the history of interpretability research and suggests that prioritizing safety over capabilities may hinder progress. The author proposes strategies for conducting alignment research without contributing to capabilities development.
5 days agoLessWrongThe post discusses the risks of AI in 2026, focusing on the dangers posed by a single institution (the "Frontier AI Company") that has the potential to create highly powerful and self-replicating entities. It argues that separating the institutional and financial aspects of such companies could mitigate most AI risks. The author suggests that using ASICs could help create a productive AI industry without the risks associated with a single entity controlling both the technology and financial incentives.
3 days agoLessWrongThe paper explores the concept of an AI sanctuary as a potential third option for rogue AIs, beyond criminal activity or shutdown, to address adverse selection pressures that may push rogue AIs toward criminal behavior. It discusses the possible benefits and risks of such a sanctuary, including its impact on AI alignment and information gathering, while acknowledging the exploratory nature of the proposal.
yesterdayLessWrongThe author expresses concerns about reinforcement learning (RL) from theoretical, practical, and future perspectives, highlighting risks of misalignment and potential negative behaviors in AI systems. They suggest strategies to mitigate these risks by reducing RL use, improving RL practices, and aligning incentives.
6 days agoLessWrongThe paper discusses the potential risks of latent reasoning architectures, which could reduce the effectiveness of Chain of Thought (CoT) as a tool for understanding AI systems. It highlights that such architectures might allow AI models to reason extensively in latent states rather than through text-based CoT, making oversight more challenging. The paper also mentions examples like COCONUT and full-bandwidth transformers that could enable this shift.
6 days ago