norgitov/ trends
Technology · people · ideas
Back to discovery/LessWrong4 hours ago

One Coordinate Breaks Abliteration on Gemma-3

The study investigates an issue with abliteration, a method for identifying refusal directions in large language models, where the standard approach failed for Gemma-3-12b but worked for similar models. A fix involving Winsorization of coordinate activations was tested, and the problem was traced to a dominant coordinate in scale. The method was validated across multiple model sizes and showed generalizability. The reason for the development of this large activation coordinate in Gemma-3 models remains unclear.

Open original
SIGNAL FROM THE SOURCE
7
source points
Tracking sinceSeptember 15, 20267 source points
Momentum0/hover 1.03 h
PublishedSeptember 15, 2026Abhishek Mishra
BEHIND THE NUMBERS

How interest changes

History starts here

The chart will appear after repeat observations. The current metric comes from the source.

7 source points

Real observations only. History before source connection is not reconstructed.

WHY IT MAY MATTER

This research highlights the importance of understanding model-specific behaviors when applying standard methods for extracting directional information, suggesting that model architecture or training data may influence the effectiveness of such techniques.

A useful discovery?
KEEP EXPLORING

Connected ideas

Explore topic