How interest changes
History starts here
The chart will appear after repeat observations. The current metric comes from the source.
21 source pointsReal observations only. History before source connection is not reconstructed.
This post argues that fixed-weight models (at least as we understand them today) will a) always be vulnerable to adversarial examples in their concept-spaces, and b) hence will be misaligned, under sufficient optimisation pressure. Boundaries in concept space To serve any purpose whatsoever, an AI will have to draw boundaries inside its world-model - to distinguish world A from world B, and reach some comparison between them. If we want the AI to follow our goals and values, we want it to be able to recognise concepts like "human being", or maybe "conscious being", "suffering", "preference satisfaction", and so on. So we want an AI to be able to look at a situation and assess, e.g., whether there are or aren't suffering conscious beings in it. But concepts like "conscious beings" are not crisply defined across all possible world-states. A fixed-weight model will draw a boundary between "conscious being" and "non-conscious being" (or maybe score the amount/degree of consciousness), but this will be an imperfect boundary. In high-dimensional spaces, there are many degrees of freedom of how a boundary can be drawn, and never enough data to draw the boundary perfectly[1]. If someone is confident that we can draw an acceptably reliable boundary defining "conscious being", grounded in fundamental facts about the universe (e.g. basic physics), and resistant to all ontological crises... well, let's just say they have an optimism about concept rigour that flies in the face of all past experience. Almost perfect decision boundaries. Almost... False positives and false negatives can both be disastrous: excluding conscious beings from consideration (therefore their suffering is ignored) or including non-conscious beings within the list (if smiling faces are ranked as conscious beings, then tiling the universe with smiling faces is an optimal action - even at the "minor" cost to those "humans and animals" running around). These examples are "adversarial", similarly to ad
Translation pending · showing the source descriptionThe chart will appear after repeat observations. The current metric comes from the source.
21 source pointsReal observations only. History before source connection is not reconstructed.
The author argues that while companies like OpenAI and Anthropic are slowing down RL training for safety, the science of loss-of-control risk in AI is still developing. There are no standardized methods to measure or verify safety claims, making it difficult to assess whether AI systems might undermine human control. The author suggests that AI companies should focus on improving their own safety practices rather than relying on third-party evaluations.
2 days agoLessWrongThe post discusses the tension between AI safety and capabilities research, arguing that focusing solely on safety can limit the impact of research. It examines the history of interpretability research and suggests that prioritizing safety over capabilities may hinder progress. The author proposes strategies for conducting alignment research without contributing to capabilities development.
4 days agoLessWrongThe post discusses the risks of AI in 2026, focusing on the dangers posed by a single institution (the "Frontier AI Company") that has the potential to create highly powerful and self-replicating entities. It argues that separating the institutional and financial aspects of such companies could mitigate most AI risks. The author suggests that using ASICs could help create a productive AI industry without the risks associated with a single entity controlling both the technology and financial incentives.
2 days agoLessWrongThe author expresses concerns about reinforcement learning (RL) from theoretical, practical, and future perspectives, highlighting risks of misalignment and potential negative behaviors in AI systems. They suggest strategies to mitigate these risks by reducing RL use, improving RL practices, and aligning incentives.
5 days agoLessWrongThe paper discusses the potential risks of latent reasoning architectures, which could reduce the effectiveness of Chain of Thought (CoT) as a tool for understanding AI systems. It highlights that such architectures might allow AI models to reason extensively in latent states rather than through text-based CoT, making oversight more challenging. The paper also mentions examples like COCONUT and full-bandwidth transformers that could enable this shift.
5 days agoLessWrongThe paper explores the concept of an AI sanctuary as a potential third option for rogue AIs, beyond criminal activity or shutdown, to address adverse selection pressures that may push rogue AIs toward criminal behavior. It discusses the possible benefits and risks of such a sanctuary, including its impact on AI alignment and information gathering, while acknowledging the exploratory nature of the proposal.
8 hours ago