norgitov/ trends
Технологии · люди · идеи
К обзору/LessWrong13 минут назад

Capabilities research expands the safety-usefulness Pareto frontier too

It’s tempting to define safety research as research that enables developers to deploy an AI system more safely without making the deployment much more expensive or much less useful. You can visualize this definition of safety research as pushing out the safety-usefulness Pareto frontier. At any given level of usefulness, there's greater safety available. Awkwardly, this definition counts basically all capabilities research as safety research. For example, consider performance optimization for inference. By making inference more efficient you can use weaker, safer models more extensively than you would otherwise be able to, pushing out the Pareto frontier. Likewise, any successful research whatsoever pushes out this Pareto frontier because research can only ever create more options. It seems like something has gone wrong with our definition of safety research if it includes seemingly all capabilities research. Here, I spell out one reason why enabling improved safety without hurting usefulness is an insufficient standard for safety research. The core observation is that developers have to choose a particular point on the Pareto frontier, and some technological improvements incentivize them to sacrifice safety. Safety research typically reshapes the Pareto frontier in a way that causes developers to choose greater safety, while capabilities research typically does the opposite. However, I also argue there is an important and plausible future regime involving much higher political will than we have right now, in which certain kinds of capabilities research would be an effective way to improve safety. I don't write this post to argue that more people should be doing capabilities research, and in fact I think that currently capabilities research at any frontier AI company is probably bad. Instead, I'm treating this as an interesting puzzle to sharpen our models of when and why safety and capabilities research are good. (Another bucket of safety interventions incl

Перевод готовится · пока описание источника
Открыть первоисточник
СИГНАЛ ИЗ ИСТОЧНИКА
46
очков источника
Наблюдаем с2 октября 2026 г.46 очков источника
Темп интереса+12,78/hпо замерам за 1,02 h
Опубликовано2 октября 2026 г.Alex Mallen
ЗА ЦИФРАМИ

Как меняется интерес

История начинает расти

График появится после повторных замеров. Текущий показатель уже получен из источника.

46 очков источника

Только реальные замеры. История до подключения источника не восстанавливается.

Полезная находка?
ПРОДОЛЖИ ИССЛЕДОВАНИЕ

Рядом по теме

Вся тема