norgitov/ trends
Technology · people · ideas
Back to discovery/LessWrong53 minutes ago

A Key-Position Confound in no-cot-bench (and implications for interpreting looped transformers)

TL;DR: * We identify a confound in no-cot-bench related to the positioning of the prompt's "key." * Correcting for this decreases GPT-6.1 Sol's no-cot reasoning depth by 16%, with the effect likely growing as dependent depth increases. * This matters because current benchmark performance reflects both serial reasoning depth and the ability to spread computation over tokens. These both measure 'opaque reasoning' but scale differently and have different implications. Introduction & Methods Neel Nanda recently released a benchmark for no-CoT reasoning, and found that Astra (which is suspected to be a looped transformer) does extremely well on it: its odds of solving an arbitrary problem are ~8.6x that of Fable 5.1. From Neel’s post: The Key-Position Confound One possible confound is that the key (the initial state that the subsequent operations act on) is usually given at the beginning of the prompt. For example, the state-machine task in no-cot-bench starts from 12 and applies six conditional updates in order (requiring a six-step serial computation): Start with the number 12 and apply the steps in order. After every step, if the number is bigger than 20, subtract 20; if it is smaller than 1, add 20. If it is bigger than 10, subtract 9; otherwise double it. If it is even, halve it; if it is odd, add 5. If it is even, halve it; if it is odd, add 5. If it is even, halve it; if it is odd, add 5. If it is even, halve it; if it is odd, add 3. If it is even, halve it; if it is odd, add 9. What is the final number? Since the model sees the key before it sees the sequence of operations, it can effectively start doing the intermediate computation while reading the prompt (e.g., the model can compute a persistent ‘hidden state’ and use attention to move it across token positions.) Thus, the no-cot benchmark might pick up measurements like “how good is this model at spreading per-step work across token positions" in addition to serial-depth. See the appendix

Translation pending · showing the source description
Open original
SIGNAL FROM THE SOURCE
13
source points
Tracking sinceOctober 2, 202613 source points
Momentum—More observations needed
PublishedOctober 2, 2026agastyasridharan
BEHIND THE NUMBERS

How interest changes

History starts here

The chart will appear after repeat observations. The current metric comes from the source.

13 source points

Real observations only. History before source connection is not reconstructed.

A useful discovery?
KEEP EXPLORING

Connected ideas

Explore topic