norgitov/ trends
Technology · people · ideas
Back to discovery/LessWrong13 minutes ago

Capabilities research expands the safety-usefulness Pareto frontier too

It’s tempting to define safety research as research that enables developers to deploy an AI system more safely without making the deployment much more expensive or much less useful. You can visualize this definition of safety research as pushing out the safety-usefulness Pareto frontier. At any given level of usefulness, there's greater safety available. Awkwardly, this definition counts basically all capabilities research as safety research. For example, consider performance optimization for inference. By making inference more efficient you can use weaker, safer models more extensively than you would otherwise be able to, pushing out the Pareto frontier. Likewise, any successful research whatsoever pushes out this Pareto frontier because research can only ever create more options. It seems like something has gone wrong with our definition of safety research if it includes seemingly all capabilities research. Here, I spell out one reason why enabling improved safety without hurting usefulness is an insufficient standard for safety research. The core observation is that developers have to choose a particular point on the Pareto frontier, and some technological improvements incentivize them to sacrifice safety. Safety research typically reshapes the Pareto frontier in a way that causes developers to choose greater safety, while capabilities research typically does the opposite. However, I also argue there is an important and plausible future regime involving much higher political will than we have right now, in which certain kinds of capabilities research would be an effective way to improve safety. I don't write this post to argue that more people should be doing capabilities research, and in fact I think that currently capabilities research at any frontier AI company is probably bad. Instead, I'm treating this as an interesting puzzle to sharpen our models of when and why safety and capabilities research are good. (Another bucket of safety interventions incl

Translation pending · showing the source description
Open original
SIGNAL FROM THE SOURCE
46
source points
Tracking sinceOctober 2, 202646 source points
Momentum+12.78/hover 1.02 h
PublishedOctober 2, 2026Alex Mallen
BEHIND THE NUMBERS

How interest changes

History starts here

The chart will appear after repeat observations. The current metric comes from the source.

46 source points

Real observations only. History before source connection is not reconstructed.

A useful discovery?
KEEP EXPLORING

Connected ideas

Explore topic