How interest changes
History starts here
The chart will appear after repeat observations. The current metric comes from the source.
46 source pointsReal observations only. History before source connection is not reconstructed.
It’s tempting to define safety research as research that enables developers to deploy an AI system more safely without making the deployment much more expensive or much less useful. You can visualize this definition of safety research as pushing out the safety-usefulness Pareto frontier. At any given level of usefulness, there's greater safety available. Awkwardly, this definition counts basically all capabilities research as safety research. For example, consider performance optimization for inference. By making inference more efficient you can use weaker, safer models more extensively than you would otherwise be able to, pushing out the Pareto frontier. Likewise, any successful research whatsoever pushes out this Pareto frontier because research can only ever create more options. It seems like something has gone wrong with our definition of safety research if it includes seemingly all capabilities research. Here, I spell out one reason why enabling improved safety without hurting usefulness is an insufficient standard for safety research. The core observation is that developers have to choose a particular point on the Pareto frontier, and some technological improvements incentivize them to sacrifice safety. Safety research typically reshapes the Pareto frontier in a way that causes developers to choose greater safety, while capabilities research typically does the opposite. However, I also argue there is an important and plausible future regime involving much higher political will than we have right now, in which certain kinds of capabilities research would be an effective way to improve safety. I don't write this post to argue that more people should be doing capabilities research, and in fact I think that currently capabilities research at any frontier AI company is probably bad. Instead, I'm treating this as an interesting puzzle to sharpen our models of when and why safety and capabilities research are good. (Another bucket of safety interventions incl
Translation pending · showing the source descriptionThe chart will appear after repeat observations. The current metric comes from the source.
46 source pointsReal observations only. History before source connection is not reconstructed.
The post discusses the risks of AI in 2026, focusing on the dangers posed by a single institution (the "Frontier AI Company") that has the potential to create highly powerful and self-replicating entities. It argues that separating the institutional and financial aspects of such companies could mitigate most AI risks. The author suggests that using ASICs could help create a productive AI industry without the risks associated with a single entity controlling both the technology and financial incentives.
6 days agoLessWrongThe text discusses the concept of 'gradual disempowerment' in AI development, questioning whether leading AI labs like Anthropic and OpenAI are accelerating capabilities more than before and whether they are focused on recursive self-improvement. It raises concerns about the potential for AI to lead to 'takeover by default' and the challenges of aligning AI with human values.
2 days agoLessWrongThe author argues that while companies like OpenAI and Anthropic are slowing down RL training for safety, the science of loss-of-control risk in AI is still developing. There are no standardized methods to measure or verify safety claims, making it difficult to assess whether AI systems might undermine human control. The author suggests that AI companies should focus on improving their own safety practices rather than relying on third-party evaluations.
6 days agoLessWrongThe text discusses cultural differences between China and the West, focusing on social media and communication styles. It highlights how Chinese social media is perceived as overwhelming and inauthentic by Western standards, and how these differences may pose challenges for AI safety awareness efforts in China.
22 hours agoLessWrongThe paper explores the concept of an AI sanctuary as a potential third option for rogue AIs, beyond criminal activity or shutdown, to address adverse selection pressures that may push rogue AIs toward criminal behavior. It discusses the possible benefits and risks of such a sanctuary, including its impact on AI alignment and information gathering, while acknowledging the exploratory nature of the proposal.
4 days agoLessWrongThe Corrigibility Research Fund aims to reward high-quality AI alignment research through retroactive prizes. The fund has awarded $27,000 so far and plans to distribute an additional $48,000, highlighting work from around two dozen researchers across a dozen teams. The fund manager emphasizes that prize sizes are not indicative of work quality and encourages feedback on improving the funding process.
2 days ago