How interest changes
History starts here
The chart will appear after repeat observations. The current metric comes from the source.
44 source pointsReal observations only. History before source connection is not reconstructed.
One of my goals for the Corrigibility Research Fund is to retroactively encourage high-quality research on AI alignment (and corrigibility in particular) by awarding prizes. Back in July, I got my feet wet as a fund manager by handing out $27,000 to reward existing work and build interest in the fund. Now, I'd like to disburse an additional $48,000 and use the opportunity to publicly highlight and celebrate the work of the prizewinners from both rounds: about two dozen researchers scattered across roughly a dozen teams. If the fund continues to be supported in future years, my hope is for prizes like these to become regular, predictable, and large, such that many researchers, year after year, are motivated to aim for them. The awards that I'm announcing here are more ad-hoc than I'd like, and represent only my single perspective trying to balance a wide range of desiderata. Don't take the specific size of each prize purse too seriously. It's all high-quality work. If anyone has ideas for how to improve the retroactive funding process for this kind of scientific work, please leave a comment! (And as always, if you know of work that I should be aware of, please email me at grants@corrigibilityresearch.org. I'm hoping to disburse more than $60k in prize funding this December, in addition to the various micro-grants that I'll be awarding on Lightcone Commons to corrigibility projects.) Before getting into the winners, I'd like to mention that even though my aim for these prizes was to reward existing work, I wanted to focus on work from this year or the last few years. As such, I chose not to award prizes to some of the most important thinkers in the corrigibility sub-field. While they are more than worthy of praise, I felt that, given the modest level of funding available, it was better to focus on scientists who hadn't already "made it" in some real sense, and researchers who were clearly actively working on the topic.[1] Those who I deliberately passed over, desp
Translation pending · showing the source descriptionThe chart will appear after repeat observations. The current metric comes from the source.
44 source pointsReal observations only. History before source connection is not reconstructed.
The author argues that while companies like OpenAI and Anthropic are slowing down RL training for safety, the science of loss-of-control risk in AI is still developing. There are no standardized methods to measure or verify safety claims, making it difficult to assess whether AI systems might undermine human control. The author suggests that AI companies should focus on improving their own safety practices rather than relying on third-party evaluations.
4 days agoLessWrongThe paper explores the concept of an AI sanctuary as a potential third option for rogue AIs, beyond criminal activity or shutdown, to address adverse selection pressures that may push rogue AIs toward criminal behavior. It discusses the possible benefits and risks of such a sanctuary, including its impact on AI alignment and information gathering, while acknowledging the exploratory nature of the proposal.
2 days agoLessWrongThe post discusses the tension between AI safety and capabilities research, arguing that focusing solely on safety can limit the impact of research. It examines the history of interpretability research and suggests that prioritizing safety over capabilities may hinder progress. The author proposes strategies for conducting alignment research without contributing to capabilities development.
6 days agoLessWrongThe post discusses the risks of AI in 2026, focusing on the dangers posed by a single institution (the "Frontier AI Company") that has the potential to create highly powerful and self-replicating entities. It argues that separating the institutional and financial aspects of such companies could mitigate most AI risks. The author suggests that using ASICs could help create a productive AI industry without the risks associated with a single entity controlling both the technology and financial incentives.
4 days agoLessWrongSubtitle: And maybe second best is AI safety? Further reading: So many things, but: Gradual Disempowerment, The Normalization of Deviance in AI Development, Let’s Think About Slowing Down AI, Doom as a bad method, not a utopia tradeoff, Teleoperated Humans, We should push for no-fault liability for actions taken by AI Thank you to JennaS for extensive edits and long-term discussion. I’ve been trying to get more writing out at 90% of the quality I’d like it to be at, instead of spending a bunch more time trying to wring out the last 10%, so a lot of points that could themselves be full articles are underdeveloped. Insofar as you find this post outlines a plausible or probable model of reality, or one worth criticizing centrally, let’s work on developing it. ---------------------------------------- Is Anthropic accelerating capabilities more than it was a year ago? At its founding? Is OpenAI accelerating capabilities more than it was a year ago? At its founding? Is GDM "laser-focused at the frontier" in pursuing recursive self-improvement? What? Why? Have they solved alignment without telling us? Why does Thomas Kwa, formerly at METR[1] and now working on "measuring and modeling RSI" at OpenAI, worry about working at OpenAI potentially driving him (metaphorically?) insane? How is it possible that all of the frontier labs have[2] a training and deployment strategy that, in the community's tacit knowledge, "leads to takeover by default"? ---------------------------------------- Imagine you went back in time to a 2021 AI researcher and told them that here in 2026: - We have slightly to moderately superhuman, legibly impressive general AIs across many domains (Solved a Millenium prize problem(s?), are productive research partners and idea generators in many areas and subdomains in physics, biology, chemistry, material science, robotics, writing, etc. The same model can, in short, augment or automate large subsets of tasks that were in the job description of an
19 hours agoLessWrongThe text discusses the lack of coherent plans in AI models regarding their behavior during a technological singularity, highlighting that models do not have concrete strategies but rather general values. It suggests that this uncertainty is concerning because models themselves are unsure about their future actions, which could lead to unpredictable outcomes.
2 days ago