How interest changes
History starts here
The chart will appear after repeat observations. The current metric comes from the source.
25 source pointsReal observations only. History before source connection is not reconstructed.
N.B. Some of this post argues by analogy between decision theory and values. I expect at least these parts of the post to be unconvincing to anyone who expects sufficiently smart agents to converge on the same values, as some moral realists do. I will not argue against moral realism here (see e.g. this sequence for one such argument). Note that I take a convergence-based definition of realism for this post. One could hold that there is a truth about the correct decision theory, but that agents won't necessarily converge on it. I don't argue against views like that here. I am interested more in the question of convergence than of truth, because the former bears on whether ASI's decision theory is path dependent. In an upcoming post, I will argue further for path dependence, and for the time sensitivity of interventions to influence AI's decision theory. This is the third post in our sequence Intro to acausal interactions. Introduction In this post, I argue against the following claim, which I call strong decision-theoretic realism: that sufficiently smart agents will all converge on the “correct” decision theory (DT). In doing so, I also argue against a related claim: that absent "lock-in", humans together with somewhat aligned AIs will necessarily converge on a reasonable decision theory. As a consequence of this, I hope to convince the reader that the avoidance of “lock-in” should not be the primary focus when considering decision-theoretic reflection processes.[1] Rather, we should care about the overall quality of the decision-theoretic reflection process (where quality is defined with respect to our object- and meta-level decision-theoretic commitments). As with values, there are certain things that we want to lock in (e.g., fundamental intuitions, some properties regarding how we wish to reflect), and other things that we don’t (e.g., complicated object-level properties that we are relatively uncertain about). There are many possible “reflection processe
Translation pending · showing the source descriptionThe chart will appear after repeat observations. The current metric comes from the source.
25 source pointsReal observations only. History before source connection is not reconstructed.
The author argues that while companies like OpenAI and Anthropic are slowing down RL training for safety, the science of loss-of-control risk in AI is still developing. There are no standardized methods to measure or verify safety claims, making it difficult to assess whether AI systems might undermine human control. The author suggests that AI companies should focus on improving their own safety practices rather than relying on third-party evaluations.
3 days agoLessWrongThe post discusses the tension between AI safety and capabilities research, arguing that focusing solely on safety can limit the impact of research. It examines the history of interpretability research and suggests that prioritizing safety over capabilities may hinder progress. The author proposes strategies for conducting alignment research without contributing to capabilities development.
5 days agoLessWrongThe post discusses the risks of AI in 2026, focusing on the dangers posed by a single institution (the "Frontier AI Company") that has the potential to create highly powerful and self-replicating entities. It argues that separating the institutional and financial aspects of such companies could mitigate most AI risks. The author suggests that using ASICs could help create a productive AI industry without the risks associated with a single entity controlling both the technology and financial incentives.
3 days agoLessWrongThe paper explores the concept of an AI sanctuary as a potential third option for rogue AIs, beyond criminal activity or shutdown, to address adverse selection pressures that may push rogue AIs toward criminal behavior. It discusses the possible benefits and risks of such a sanctuary, including its impact on AI alignment and information gathering, while acknowledging the exploratory nature of the proposal.
yesterdayLessWrongThe author expresses concerns about reinforcement learning (RL) from theoretical, practical, and future perspectives, highlighting risks of misalignment and potential negative behaviors in AI systems. They suggest strategies to mitigate these risks by reducing RL use, improving RL practices, and aligning incentives.
6 days agoLessWrongThe paper discusses the potential risks of latent reasoning architectures, which could reduce the effectiveness of Chain of Thought (CoT) as a tool for understanding AI systems. It highlights that such architectures might allow AI models to reason extensively in latent states rather than through text-based CoT, making oversight more challenging. The paper also mentions examples like COCONUT and full-bandwidth transformers that could enable this shift.
6 days ago