How interest changes
History starts here
The chart will appear after repeat observations. The current metric comes from the source.
15 source pointsReal observations only. History before source connection is not reconstructed.
This is a linkpost for https://www.apolloresearch.ai/blog/principles-for-embedded-evaluations Blog post > In September 2026, leaders of frontier AI companies called for pacing the frontier of AI development, with embedded evaluators as the first step. Frontier AI companies committed to giving outside evaluators employee-like access to their training, evaluation, and deployment, and some have since published principles for third-party assessments. > > We are very excited about this development. Public third-party assessments of how frontier AI is developed are urgently needed, and embedded evaluations are a good first step. But their impact will depend heavily on implementation. If evaluators lack necessary access, if they are given too few resources, or if their findings carry little weight in actual decisions about frontier development, embedded evaluations may not amount to much. > > This post sets out our current thinking on core principles for embedded evaluations that assess loss of control risks from scheming, i.e., AI models covertly subverting their developers in pursuit of unintended goals. We first describe at a high level what effective embedded evaluations should achieve. We then propose a concrete design based on verifying or falsifying developers' safety claims, which we hope developers and evaluators will adopt. We see these principles as a minimal starting point rather than a complete framework. Many of them build on established practice for independent auditing in other high-stakes industries, adapted to the specific challenges of frontier AI. > > [...] Twitter Thread > 1/ > Frontier AI companies have committed to giving outside evaluators employee-like access to training, evaluation and deployment. We're very excited about this. But its impact depends heavily on implementation. > > Today we're sharing our principles for embedded evaluations. 🧵 > > 2/ > If evaluators lack access or resources, or if their findings carry little weight in rea
Translation pending · showing the source descriptionThe chart will appear after repeat observations. The current metric comes from the source.
15 source pointsReal observations only. History before source connection is not reconstructed.
The author argues that while companies like OpenAI and Anthropic are slowing down RL training for safety, the science of loss-of-control risk in AI is still developing. There are no standardized methods to measure or verify safety claims, making it difficult to assess whether AI systems might undermine human control. The author suggests that AI companies should focus on improving their own safety practices rather than relying on third-party evaluations.
4 days agoLessWrongThe paper explores the concept of an AI sanctuary as a potential third option for rogue AIs, beyond criminal activity or shutdown, to address adverse selection pressures that may push rogue AIs toward criminal behavior. It discusses the possible benefits and risks of such a sanctuary, including its impact on AI alignment and information gathering, while acknowledging the exploratory nature of the proposal.
2 days agoLessWrongThe post discusses the tension between AI safety and capabilities research, arguing that focusing solely on safety can limit the impact of research. It examines the history of interpretability research and suggests that prioritizing safety over capabilities may hinder progress. The author proposes strategies for conducting alignment research without contributing to capabilities development.
6 days agoLessWrongThe post discusses the risks of AI in 2026, focusing on the dangers posed by a single institution (the "Frontier AI Company") that has the potential to create highly powerful and self-replicating entities. It argues that separating the institutional and financial aspects of such companies could mitigate most AI risks. The author suggests that using ASICs could help create a productive AI industry without the risks associated with a single entity controlling both the technology and financial incentives.
4 days agoLessWrongSubtitle: And maybe second best is AI safety? Further reading: So many things, but: Gradual Disempowerment, The Normalization of Deviance in AI Development, Let’s Think About Slowing Down AI, Doom as a bad method, not a utopia tradeoff, Teleoperated Humans, We should push for no-fault liability for actions taken by AI Thank you to JennaS for extensive edits and long-term discussion. I’ve been trying to get more writing out at 90% of the quality I’d like it to be at, instead of spending a bunch more time trying to wring out the last 10%, so a lot of points that could themselves be full articles are underdeveloped. Insofar as you find this post outlines a plausible or probable model of reality, or one worth criticizing centrally, let’s work on developing it. ---------------------------------------- Is Anthropic accelerating capabilities more than it was a year ago? At its founding? Is OpenAI accelerating capabilities more than it was a year ago? At its founding? Is GDM "laser-focused at the frontier" in pursuing recursive self-improvement? What? Why? Have they solved alignment without telling us? Why does Thomas Kwa, formerly at METR[1] and now working on "measuring and modeling RSI" at OpenAI, worry about working at OpenAI potentially driving him (metaphorically?) insane? How is it possible that all of the frontier labs have[2] a training and deployment strategy that, in the community's tacit knowledge, "leads to takeover by default"? ---------------------------------------- Imagine you went back in time to a 2021 AI researcher and told them that here in 2026: - We have slightly to moderately superhuman, legibly impressive general AIs across many domains (Solved a Millenium prize problem(s?), are productive research partners and idea generators in many areas and subdomains in physics, biology, chemistry, material science, robotics, writing, etc. The same model can, in short, augment or automate large subsets of tasks that were in the job description of an
19 hours agoLessWrongThe text discusses the lack of coherent plans in AI models regarding their behavior during a technological singularity, highlighting that models do not have concrete strategies but rather general values. It suggests that this uncertainty is concerning because models themselves are unsure about their future actions, which could lead to unpredictable outcomes.
2 days ago