Как меняется интерес
История начинает расти
График появится после повторных замеров. Текущий показатель уже получен из источника.
44 очков источникаТолько реальные замеры. История до подключения источника не восстанавливается.
One of my goals for the Corrigibility Research Fund is to retroactively encourage high-quality research on AI alignment (and corrigibility in particular) by awarding prizes. Back in July, I got my feet wet as a fund manager by handing out $27,000 to reward existing work and build interest in the fund. Now, I'd like to disburse an additional $48,000 and use the opportunity to publicly highlight and celebrate the work of the prizewinners from both rounds: about two dozen researchers scattered across roughly a dozen teams. If the fund continues to be supported in future years, my hope is for prizes like these to become regular, predictable, and large, such that many researchers, year after year, are motivated to aim for them. The awards that I'm announcing here are more ad-hoc than I'd like, and represent only my single perspective trying to balance a wide range of desiderata. Don't take the specific size of each prize purse too seriously. It's all high-quality work. If anyone has ideas for how to improve the retroactive funding process for this kind of scientific work, please leave a comment! (And as always, if you know of work that I should be aware of, please email me at grants@corrigibilityresearch.org. I'm hoping to disburse more than $60k in prize funding this December, in addition to the various micro-grants that I'll be awarding on Lightcone Commons to corrigibility projects.) Before getting into the winners, I'd like to mention that even though my aim for these prizes was to reward existing work, I wanted to focus on work from this year or the last few years. As such, I chose not to award prizes to some of the most important thinkers in the corrigibility sub-field. While they are more than worthy of praise, I felt that, given the modest level of funding available, it was better to focus on scientists who hadn't already "made it" in some real sense, and researchers who were clearly actively working on the topic.[1] Those who I deliberately passed over, desp
Перевод готовится · пока описание источникаГрафик появится после повторных замеров. Текущий показатель уже получен из источника.
44 очков источникаТолько реальные замеры. История до подключения источника не восстанавливается.
Автор утверждает, что, несмотря на то, что компании OpenAI и Anthropic замедляют обучение RL для обеспечения безопасности, наука о рисках потери контроля в ИИ все еще развивается. Нет стандартизированных методов измерения или проверки утверждений о безопасности, что затрудняет оценку того, могут ли ИИ-системы подорвать человеческий контроль. Автор считает, что компании ИИ должны сосредоточиться на улучшении собственных практик безопасности вместо того, чтобы полагаться на независимые оценки.
4 дня назадLessWrongСтатья рассматривает концепцию AI-санатория как потенциальный третий вариант для неподконтрольных ИИ, помимо преступной деятельности или отключения, чтобы справиться с неблагоприятными давлениями отбора, которые могут толкать неподконтрольных ИИ к преступной деятельности. Обсуждаются возможные преимущества и риски такого санатория, включая его влияние на выравнивание ИИ и сбор информации, признавая экспериментальный характер предложения.
позавчераLessWrongВ посте обсуждается напряжённость между исследованиями в области безопасности ИИ и возможностями, утверждая, что чрезмерное внимание к безопасности может ограничивать влияние исследований. Он анализирует историю исследований по интерпретируемости и предлагает стратегии для проведения исследований в области согласованности без вреда для возможностей.
6 дней назадLessWrongВ посте обсуждаются риски ИИ в 2026 году, акцентируя внимание на опасностях, связанных с наличием единственной организации («Frontier AI Company»), которая имеет возможность создавать очень мощные и самовоспроизводящиеся сущности. Автор утверждает, что разделение институциональных и финансовых аспектов таких компаний может снизить большинство рисков ИИ. Автор предполагает, что использование ASIC может помочь создать продуктивную индустрию ИИ без рисков, связанных с контролем одной организацией как технологии, так и финансовых стимулов.
4 дня назадLessWrongSubtitle: And maybe second best is AI safety? Further reading: So many things, but: Gradual Disempowerment, The Normalization of Deviance in AI Development, Let’s Think About Slowing Down AI, Doom as a bad method, not a utopia tradeoff, Teleoperated Humans, We should push for no-fault liability for actions taken by AI Thank you to JennaS for extensive edits and long-term discussion. I’ve been trying to get more writing out at 90% of the quality I’d like it to be at, instead of spending a bunch more time trying to wring out the last 10%, so a lot of points that could themselves be full articles are underdeveloped. Insofar as you find this post outlines a plausible or probable model of reality, or one worth criticizing centrally, let’s work on developing it. ---------------------------------------- Is Anthropic accelerating capabilities more than it was a year ago? At its founding? Is OpenAI accelerating capabilities more than it was a year ago? At its founding? Is GDM "laser-focused at the frontier" in pursuing recursive self-improvement? What? Why? Have they solved alignment without telling us? Why does Thomas Kwa, formerly at METR[1] and now working on "measuring and modeling RSI" at OpenAI, worry about working at OpenAI potentially driving him (metaphorically?) insane? How is it possible that all of the frontier labs have[2] a training and deployment strategy that, in the community's tacit knowledge, "leads to takeover by default"? ---------------------------------------- Imagine you went back in time to a 2021 AI researcher and told them that here in 2026: - We have slightly to moderately superhuman, legibly impressive general AIs across many domains (Solved a Millenium prize problem(s?), are productive research partners and idea generators in many areas and subdomains in physics, biology, chemistry, material science, robotics, writing, etc. The same model can, in short, augment or automate large subsets of tasks that were in the job description of an
19 часов назадLessWrongТекст обсуждает отсутствие согласованных планов у ИИ-моделей относительно их поведения во время технологического сингулярности, подчеркивая, что модели не имеют конкретных стратегий, а только общие ценности. Это вызывает опасения, так как сами модели не уверены в своих будущих действиях, что может привести к непредсказуемым последствиям.
позавчера