norgitov/ trends
Technology · people · ideas
Back to discovery/Daily Papers1 hour ago

Noise Out, Bias In: Targeted Bias Injection in Diffusion Language Models via Closed-Loop Activation Steering

The paper explores targeted bias injection in diffusion language models (dLLMs) by exploiting their denoising process. An adversary can manipulate the model's output by adjusting steering vectors based on internal activations, leading to significant shifts in demographic answers. The attack, using a proportional-integral controller, achieves substantial bias amplification on specific datasets.

Open original
SIGNAL FROM THE SOURCE
1
source votes
Tracking sinceOctober 6, 20261 source votes
Momentum—More observations needed
Discussion—Read comments ↗
PublishedOctober 5, 2026Sarim Hashmi, Mukul Ranjan, Abdelrahman Elsayed
BEHIND THE NUMBERS

How interest changes

History starts here

The chart will appear after repeat observations. The current metric comes from the source.

1 source votes

Real observations only. History before source connection is not reconstructed.

WHY IT MAY MATTER

This research highlights a vulnerability in diffusion language models where internal activations can be exploited to inject bias, emphasizing the need for comprehensive bias audits beyond just the model itself.

A useful discovery?
KEEP EXPLORING

Connected ideas

Explore topic