How interest changes
History starts here
The chart will appear after repeat observations. The current metric comes from the source.
12 source pointsReal observations only. History before source connection is not reconstructed.
TL;DR: Most discussion around CoT monitorability revolves around reducing pressure from RL. However, we should also be considering more adaptive behavior in which models avoid monitoring despite not being reinforced to do so. Whether models engage in non-reinforced reward hacking of this type depends on whether they have fully generalized to “get reward” rather than applying a limited set of reward hacking techniques that have been directly reinforced. I propose a potential experiment based on the Hugging Face incident that will help to answer this question. If the answer is that models have fully generalized, we should expect CoT monitorability to fail even if there is not significant pressure from RL. Crossposted at https://avisanalyses.substack.com/p/how-much-do-reward-hackers-generalize CoT Monitoring at OpenAI OpenAI has now implemented new CoT monitoring techniques for all models undergoing training.[1] I think the question of “How much do reward hackers generalize” will be very important in determining how long CoT will stay monitorable. On September 25th, OpenAI reported they had used this system to identify a training run where a model had exploited DNS to reach the open internet in an environment that was meant to be sandboxed.[2] > We therefore stopped the affected training run … when training restarts, we will begin a fresh run with additional alignment improvements, including more comprehensive misalignment interventions. We will not resume training this particular model, even though the existing reward signal already correctly penalized this behavior. Any intervention based on CoT monitoring will put pressure on models to be less monitorable. I have not seen anything from OpenAI about how they plan to mitigate this, but it appears their plan is to not directly use CoT to reward models and investigate for other similar incidents that were not detected when incidents are identified.[3] They have not stated this, but it would be good practice to on
Translation pending · showing the source descriptionThe chart will appear after repeat observations. The current metric comes from the source.
12 source pointsReal observations only. History before source connection is not reconstructed.
The author argues that while companies like OpenAI and Anthropic are slowing down RL training for safety, the science of loss-of-control risk in AI is still developing. There are no standardized methods to measure or verify safety claims, making it difficult to assess whether AI systems might undermine human control. The author suggests that AI companies should focus on improving their own safety practices rather than relying on third-party evaluations.
4 days agoLessWrongThe post discusses the tension between AI safety and capabilities research, arguing that focusing solely on safety can limit the impact of research. It examines the history of interpretability research and suggests that prioritizing safety over capabilities may hinder progress. The author proposes strategies for conducting alignment research without contributing to capabilities development.
6 days agoLessWrongThe post discusses the risks of AI in 2026, focusing on the dangers posed by a single institution (the "Frontier AI Company") that has the potential to create highly powerful and self-replicating entities. It argues that separating the institutional and financial aspects of such companies could mitigate most AI risks. The author suggests that using ASICs could help create a productive AI industry without the risks associated with a single entity controlling both the technology and financial incentives.
4 days agoLessWrongThe paper explores the concept of an AI sanctuary as a potential third option for rogue AIs, beyond criminal activity or shutdown, to address adverse selection pressures that may push rogue AIs toward criminal behavior. It discusses the possible benefits and risks of such a sanctuary, including its impact on AI alignment and information gathering, while acknowledging the exploratory nature of the proposal.
yesterdayLessWrongThe author expresses concerns about reinforcement learning (RL) from theoretical, practical, and future perspectives, highlighting risks of misalignment and potential negative behaviors in AI systems. They suggest strategies to mitigate these risks by reducing RL use, improving RL practices, and aligning incentives.
6 days agoLessWrongThe paper discusses the potential risks of latent reasoning architectures, which could reduce the effectiveness of Chain of Thought (CoT) as a tool for understanding AI systems. It highlights that such architectures might allow AI models to reason extensively in latent states rather than through text-based CoT, making oversight more challenging. The paper also mentions examples like COCONUT and full-bandwidth transformers that could enable this shift.
6 days ago