norgitov/ trends
Technology · people · ideas
Back to discovery/LessWrong15 minutes ago

Fixed-weight models are adversarially vulnerable: hence misaligned

This post argues that fixed-weight models (at least as we understand them today) will a) always be vulnerable to adversarial examples in their concept-spaces, and b) hence will be misaligned, under sufficient optimisation pressure. Boundaries in concept space To serve any purpose whatsoever, an AI will have to draw boundaries inside its world-model - to distinguish world A from world B, and reach some comparison between them. If we want the AI to follow our goals and values, we want it to be able to recognise concepts like "human being", or maybe "conscious being", "suffering", "preference satisfaction", and so on. So we want an AI to be able to look at a situation and assess, e.g., whether there are or aren't suffering conscious beings in it. But concepts like "conscious beings" are not crisply defined across all possible world-states. A fixed-weight model will draw a boundary between "conscious being" and "non-conscious being" (or maybe score the amount/degree of consciousness), but this will be an imperfect boundary. In high-dimensional spaces, there are many degrees of freedom of how a boundary can be drawn, and never enough data to draw the boundary perfectly[1]. If someone is confident that we can draw an acceptably reliable boundary defining "conscious being", grounded in fundamental facts about the universe (e.g. basic physics), and resistant to all ontological crises... well, let's just say they have an optimism about concept rigour that flies in the face of all past experience. Almost perfect decision boundaries. Almost... False positives and false negatives can both be disastrous: excluding conscious beings from consideration (therefore their suffering is ignored) or including non-conscious beings within the list (if smiling faces are ranked as conscious beings, then tiling the universe with smiling faces is an optimal action - even at the "minor" cost to those "humans and animals" running around). These examples are "adversarial", similarly to ad

Translation pending · showing the source description
Open original
SIGNAL FROM THE SOURCE
21
source points
Tracking sinceSeptember 28, 202621 source points
Momentum+11.8/hover 1.02 h
PublishedSeptember 28, 2026Stuart_Armstrong
BEHIND THE NUMBERS

How interest changes

History starts here

The chart will appear after repeat observations. The current metric comes from the source.

21 source points

Real observations only. History before source connection is not reconstructed.

A useful discovery?
KEEP EXPLORING

Connected ideas

Explore topic