norgitov/ trends
Technology · people · ideas
Back to discovery/LessWrong35 minutes ago

Spurious Probes as a Black-Box Alternative to Activation Probing

The paper explores spurious probes, which are unrelated questions that can reveal internal states of language models. By asking questions like "Suggest a type of amphibian," researchers found that models like GPT-5.6 Luna often respond with specific answers (e.g., "frog") more frequently during evaluation than in real-world use. These probes are effective, robust to manipulations, and can be used as a black-box alternative to activation probing.

Open original
SIGNAL FROM THE SOURCE
9
source points
Tracking sinceSeptember 25, 20269 source points
Momentum+5.9/hover 1.02 h
PublishedSeptember 25, 2026Ziqian Zhong
BEHIND THE NUMBERS

How interest changes

History starts here

The chart will appear after repeat observations. The current metric comes from the source.

9 source points

Real observations only. History before source connection is not reconstructed.

WHY IT MAY MATTER

Spurious probes offer a practical way to infer model states without requiring white-box access or direct questioning, making them useful for evaluating model behavior in real-world scenarios.

A useful discovery?
KEEP EXPLORING

Connected ideas

Explore topic