norgitov/ trends
Technology · people · ideas
Back to discovery/LessWrong15 minutes ago

Distillation for Incrimination and Distillation for Capabilities

The paper explores two distillation approaches for AI safety: Distillation for Incrimination (DFI) and Distillation for Capabilities (DFC). DFI aims to transfer a model's misalignment to a weaker student model to expose the teacher's issues, while DFC focuses on transferring capabilities without misalignment. Experiments show DFI works better when the student shares the teacher's base model, and DFC can preserve capabilities while blocking subliminal biases.

Open original
SIGNAL FROM THE SOURCE
20
source points
Tracking sinceOctober 9, 202620 source points
Momentum0/hover 0.55 h
PublishedOctober 9, 2026sebastian_prasanna
BEHIND THE NUMBERS

How interest changes

History starts here

The chart will appear after repeat observations. The current metric comes from the source.

20 source points

Real observations only. History before source connection is not reconstructed.

WHY IT MAY MATTER

This research provides empirical insights into using distillation techniques to enhance AI safety by either exposing misalignments or preserving capabilities without harmful biases, which could inform safer model development and auditing practices.

A useful discovery?
KEEP EXPLORING

Connected ideas

Explore topic