norgitov/ trends
Technology · people · ideas
Back to discovery/LessWrong6 hours ago

Evaluating Task Vectors, Unlearning, and Inoculation

The paper presents an empirical evaluation of task vectors, unlearning, and inoculation methods using a single model and toy datasets. It explores different inoculation approaches, such as hand-written prompts, adapters, and soft prompts, to manage undesired traits in deep learning models. The results are limited in scope and may not generalize to other models or datasets.

Open original
SIGNAL FROM THE SOURCE
12
source points
Tracking sinceSeptember 20, 202612 source points

This source has not updated recently. Its observation time is shown above.

Momentum0/hover 1.02 h
PublishedSeptember 20, 2026Xenomirant
BEHIND THE NUMBERS

How interest changes

History starts here

The chart will appear after repeat observations. The current metric comes from the source.

12 source points

Real observations only. History before source connection is not reconstructed.

WHY IT MAY MATTER

This research provides insights into managing undesired traits in AI models through various inoculation techniques, but results are limited to a single model and toy datasets, so practical application should be approached with caution.

A useful discovery?
KEEP EXPLORING

Connected ideas

Explore topic