norgitov/ trends
Technology · people · ideas
Back to discovery/LessWrong1 hour ago

WorkspaceBench: Evaluating Interpretability Methods for the Global Workspace

WorkspaceBench is a benchmark designed to evaluate how well activation-to-text tools can interpret the contents of a model's global workspace, focusing on intermediate variables during a forward pass. It includes 3,356 questions across 27 evaluation families covering safety, logical reasoning, and multihop computation, with a focus on minimizing hallucinations.

Open original
SIGNAL FROM THE SOURCE
4
source points
Tracking sinceSeptember 23, 20264 source points
MomentumMore observations needed
PublishedSeptember 23, 2026camilablank
BEHIND THE NUMBERS

How interest changes

History starts here

The chart will appear after repeat observations. The current metric comes from the source.

4 source points

Real observations only. History before source connection is not reconstructed.

WHY IT MAY MATTER

This benchmark helps assess how well interpretability tools can decode model behavior by analyzing intermediate variables, which is crucial for model auditing and safety monitoring.

A useful discovery?
KEEP EXPLORING

Connected ideas

Explore topic