norgitov/ trends
Технологии · люди · идеи
К обзору/LessWrong53 минуты назад

A Key-Position Confound in no-cot-bench (and implications for interpreting looped transformers)

TL;DR: * We identify a confound in no-cot-bench related to the positioning of the prompt's "key." * Correcting for this decreases GPT-6.1 Sol's no-cot reasoning depth by 16%, with the effect likely growing as dependent depth increases. * This matters because current benchmark performance reflects both serial reasoning depth and the ability to spread computation over tokens. These both measure 'opaque reasoning' but scale differently and have different implications. Introduction & Methods Neel Nanda recently released a benchmark for no-CoT reasoning, and found that Astra (which is suspected to be a looped transformer) does extremely well on it: its odds of solving an arbitrary problem are ~8.6x that of Fable 5.1. From Neel’s post: The Key-Position Confound One possible confound is that the key (the initial state that the subsequent operations act on) is usually given at the beginning of the prompt. For example, the state-machine task in no-cot-bench starts from 12 and applies six conditional updates in order (requiring a six-step serial computation): Start with the number 12 and apply the steps in order. After every step, if the number is bigger than 20, subtract 20; if it is smaller than 1, add 20. If it is bigger than 10, subtract 9; otherwise double it. If it is even, halve it; if it is odd, add 5. If it is even, halve it; if it is odd, add 5. If it is even, halve it; if it is odd, add 5. If it is even, halve it; if it is odd, add 3. If it is even, halve it; if it is odd, add 9. What is the final number? Since the model sees the key before it sees the sequence of operations, it can effectively start doing the intermediate computation while reading the prompt (e.g., the model can compute a persistent ‘hidden state’ and use attention to move it across token positions.) Thus, the no-cot benchmark might pick up measurements like “how good is this model at spreading per-step work across token positions" in addition to serial-depth. See the appendix

Перевод готовится · пока описание источника
Открыть первоисточник
СИГНАЛ ИЗ ИСТОЧНИКА
13
очков источника
Наблюдаем с2 октября 2026 г.13 очков источника
Темп интереса—Нужны повторные замеры
Опубликовано2 октября 2026 г.agastyasridharan
ЗА ЦИФРАМИ

Как меняется интерес

История начинает расти

График появится после повторных замеров. Текущий показатель уже получен из источника.

13 очков источника

Только реальные замеры. История до подключения источника не восстанавливается.

Полезная находка?
ПРОДОЛЖИ ИССЛЕДОВАНИЕ

Рядом по теме

Вся тема