norgitov/ trends
Technology · people · ideas
Back to discovery/LessWrong11 minutes ago

Grafting SDF Changes from Base Models onto Post-Trained Models

The paper introduces a method called grafting, which involves applying weight changes from a base model's synthetic document fine-tuning (SDF) to a post-trained model. This approach improves alignment-related properties like stable preferences and reality comprehension compared to traditional SDF methods.

Open original
SIGNAL FROM THE SOURCE
22
source points
Tracking sinceOctober 6, 202622 source points
Momentum+2.9/hover 1.03 h
PublishedOctober 6, 2026dani roytburg
BEHIND THE NUMBERS

How interest changes

History starts here

The chart will appear after repeat observations. The current metric comes from the source.

22 source points

Real observations only. History before source connection is not reconstructed.

WHY IT MAY MATTER

This method could be useful for improving alignment properties in large language models without extensive retraining.

A useful discovery?
KEEP EXPLORING

Connected ideas

Explore topic