norgitov/ trends
Technology · people · ideas
Back to discovery/LessWrong55 minutes ago

Improving CoT Monitorability of Evaluation Awareness via Verbalization Training

This paper explores how verbalization training (VT) can enhance a model's ability to express its evaluation awareness during chain-of-thought (CoT) reasoning. The study shows that VT increases verbalized evaluation awareness by 2.4-2.9x across three models, while maintaining task performance and evaluation recognition capabilities. The authors suggest that improving reporting propensity could be key to better CoT monitorability, but caution against naive approaches that might harm the trustworthiness of verbalizations.

Open original
SIGNAL FROM THE SOURCE
8
source points
Tracking sinceOctober 2, 20268 source points
Momentum0/hover 1.02 h
PublishedOctober 2, 2026Usman Anwar
BEHIND THE NUMBERS

How interest changes

History starts here

The chart will appear after repeat observations. The current metric comes from the source.

8 source points

Real observations only. History before source connection is not reconstructed.

WHY IT MAY MATTER

This research could help improve the transparency of AI models by making them more likely to express their evaluation awareness during reasoning, which is important for safety and reliability. However, care must be taken to avoid unintended negative effects on the trustworthiness of their verbalizations.

A useful discovery?
KEEP EXPLORING

Connected ideas

Explore topic