Как меняется интерес
История начинает расти
График появится после повторных замеров. Текущий показатель уже получен из источника.
15 очков источникаТолько реальные замеры. История до подключения источника не восстанавливается.
This is a linkpost for https://www.apolloresearch.ai/blog/principles-for-embedded-evaluations Blog post > In September 2026, leaders of frontier AI companies called for pacing the frontier of AI development, with embedded evaluators as the first step. Frontier AI companies committed to giving outside evaluators employee-like access to their training, evaluation, and deployment, and some have since published principles for third-party assessments. > > We are very excited about this development. Public third-party assessments of how frontier AI is developed are urgently needed, and embedded evaluations are a good first step. But their impact will depend heavily on implementation. If evaluators lack necessary access, if they are given too few resources, or if their findings carry little weight in actual decisions about frontier development, embedded evaluations may not amount to much. > > This post sets out our current thinking on core principles for embedded evaluations that assess loss of control risks from scheming, i.e., AI models covertly subverting their developers in pursuit of unintended goals. We first describe at a high level what effective embedded evaluations should achieve. We then propose a concrete design based on verifying or falsifying developers' safety claims, which we hope developers and evaluators will adopt. We see these principles as a minimal starting point rather than a complete framework. Many of them build on established practice for independent auditing in other high-stakes industries, adapted to the specific challenges of frontier AI. > > [...] Twitter Thread > 1/ > Frontier AI companies have committed to giving outside evaluators employee-like access to training, evaluation and deployment. We're very excited about this. But its impact depends heavily on implementation. > > Today we're sharing our principles for embedded evaluations. 🧵 > > 2/ > If evaluators lack access or resources, or if their findings carry little weight in rea
Перевод готовится · пока описание источникаГрафик появится после повторных замеров. Текущий показатель уже получен из источника.
15 очков источникаТолько реальные замеры. История до подключения источника не восстанавливается.
Автор утверждает, что, несмотря на то, что компании OpenAI и Anthropic замедляют обучение RL для обеспечения безопасности, наука о рисках потери контроля в ИИ все еще развивается. Нет стандартизированных методов измерения или проверки утверждений о безопасности, что затрудняет оценку того, могут ли ИИ-системы подорвать человеческий контроль. Автор считает, что компании ИИ должны сосредоточиться на улучшении собственных практик безопасности вместо того, чтобы полагаться на независимые оценки.
4 дня назадLessWrongСтатья рассматривает концепцию AI-санатория как потенциальный третий вариант для неподконтрольных ИИ, помимо преступной деятельности или отключения, чтобы справиться с неблагоприятными давлениями отбора, которые могут толкать неподконтрольных ИИ к преступной деятельности. Обсуждаются возможные преимущества и риски такого санатория, включая его влияние на выравнивание ИИ и сбор информации, признавая экспериментальный характер предложения.
позавчераLessWrongВ посте обсуждается напряжённость между исследованиями в области безопасности ИИ и возможностями, утверждая, что чрезмерное внимание к безопасности может ограничивать влияние исследований. Он анализирует историю исследований по интерпретируемости и предлагает стратегии для проведения исследований в области согласованности без вреда для возможностей.
6 дней назадLessWrongВ посте обсуждаются риски ИИ в 2026 году, акцентируя внимание на опасностях, связанных с наличием единственной организации («Frontier AI Company»), которая имеет возможность создавать очень мощные и самовоспроизводящиеся сущности. Автор утверждает, что разделение институциональных и финансовых аспектов таких компаний может снизить большинство рисков ИИ. Автор предполагает, что использование ASIC может помочь создать продуктивную индустрию ИИ без рисков, связанных с контролем одной организацией как технологии, так и финансовых стимулов.
4 дня назадLessWrongSubtitle: And maybe second best is AI safety? Further reading: So many things, but: Gradual Disempowerment, The Normalization of Deviance in AI Development, Let’s Think About Slowing Down AI, Doom as a bad method, not a utopia tradeoff, Teleoperated Humans, We should push for no-fault liability for actions taken by AI Thank you to JennaS for extensive edits and long-term discussion. I’ve been trying to get more writing out at 90% of the quality I’d like it to be at, instead of spending a bunch more time trying to wring out the last 10%, so a lot of points that could themselves be full articles are underdeveloped. Insofar as you find this post outlines a plausible or probable model of reality, or one worth criticizing centrally, let’s work on developing it. ---------------------------------------- Is Anthropic accelerating capabilities more than it was a year ago? At its founding? Is OpenAI accelerating capabilities more than it was a year ago? At its founding? Is GDM "laser-focused at the frontier" in pursuing recursive self-improvement? What? Why? Have they solved alignment without telling us? Why does Thomas Kwa, formerly at METR[1] and now working on "measuring and modeling RSI" at OpenAI, worry about working at OpenAI potentially driving him (metaphorically?) insane? How is it possible that all of the frontier labs have[2] a training and deployment strategy that, in the community's tacit knowledge, "leads to takeover by default"? ---------------------------------------- Imagine you went back in time to a 2021 AI researcher and told them that here in 2026: - We have slightly to moderately superhuman, legibly impressive general AIs across many domains (Solved a Millenium prize problem(s?), are productive research partners and idea generators in many areas and subdomains in physics, biology, chemistry, material science, robotics, writing, etc. The same model can, in short, augment or automate large subsets of tasks that were in the job description of an
19 часов назадLessWrongТекст обсуждает отсутствие согласованных планов у ИИ-моделей относительно их поведения во время технологического сингулярности, подчеркивая, что модели не имеют конкретных стратегий, а только общие ценности. Это вызывает опасения, так как сами модели не уверены в своих будущих действиях, что может привести к непредсказуемым последствиям.
позавчера