How interest changes
History starts here
The chart will appear after repeat observations. The current metric comes from the source.
13 source pointsReal observations only. History before source connection is not reconstructed.
TL;DR: * We identify a confound in no-cot-bench related to the positioning of the prompt's "key." * Correcting for this decreases GPT-6.1 Sol's no-cot reasoning depth by 16%, with the effect likely growing as dependent depth increases. * This matters because current benchmark performance reflects both serial reasoning depth and the ability to spread computation over tokens. These both measure 'opaque reasoning' but scale differently and have different implications. Introduction & Methods Neel Nanda recently released a benchmark for no-CoT reasoning, and found that Astra (which is suspected to be a looped transformer) does extremely well on it: its odds of solving an arbitrary problem are ~8.6x that of Fable 5.1. From Neel’s post: The Key-Position Confound One possible confound is that the key (the initial state that the subsequent operations act on) is usually given at the beginning of the prompt. For example, the state-machine task in no-cot-bench starts from 12 and applies six conditional updates in order (requiring a six-step serial computation): Start with the number 12 and apply the steps in order. After every step, if the number is bigger than 20, subtract 20; if it is smaller than 1, add 20. If it is bigger than 10, subtract 9; otherwise double it. If it is even, halve it; if it is odd, add 5. If it is even, halve it; if it is odd, add 5. If it is even, halve it; if it is odd, add 5. If it is even, halve it; if it is odd, add 3. If it is even, halve it; if it is odd, add 9. What is the final number? Since the model sees the key before it sees the sequence of operations, it can effectively start doing the intermediate computation while reading the prompt (e.g., the model can compute a persistent ‘hidden state’ and use attention to move it across token positions.) Thus, the no-cot benchmark might pick up measurements like “how good is this model at spreading per-step work across token positions" in addition to serial-depth. See the appendix
Translation pending · showing the source descriptionThe chart will appear after repeat observations. The current metric comes from the source.
13 source pointsReal observations only. History before source connection is not reconstructed.
The text discusses a scenario where an AI must respond to the question 'What's the date?' without access to real-time data. It explores the challenge of providing an accurate date without hallucinating, considering OpenAI's training focus on avoiding false information. The text speculates on potential model versions and knowledge cutoff dates but acknowledges uncertainty due to lack of explicit information.
yesterdayLessWrongFrontier models show different decision theory preferences based on the perceived user background, favoring FDT/UDT when not influenced by academic philosophy cues and CDT when prompted to adopt an academic perspective. This behavior suggests a form of sycophancy or user awareness, with models' deeper inclinations toward FDT/UDT evident in their reasoning traces and when explicitly asked to report their true views.
2 days agoDaily PapersThe paper investigates the scaling properties of on-policy distillation (OPD) in reinforcement learning, focusing on how capabilities transfer between different model scales. It identifies a useful-transfer regime where held-out accuracy increases linearly with the reverse KL divergence from the student's initialization, and finds that smaller teachers can outperform larger ones in capability transfer.
6 days agoDaily PapersThe paper introduces VisionHOPE, a novel visual backbone that functions as a self-modifying learning system, allowing the model to co-evolve what it remembers and how it learns within an image. It uses five coupled memories and a stability-matched step-size control scheme to ensure stable learning dynamics, achieving competitive results on benchmark datasets like ImageNet-1K, COCO, and ADE20K.
5 days agoDaily PapersThe paper explores phase sensitivity in models using chunked KV-cache compression, where retrieval performance varies systematically across different phases of compressed token windows. It shows that long-context retrieval accuracy can differ by up to 40 percentage points between phases, highlighting the need for phase-specific evaluation.
4 days agoDaily PapersThe paper explores test-time AI-for-AI, focusing on how a Builder can create better execution environments for a Target while keeping both models' weights fixed. It introduces Meta-Skill, principles derived from Target's execution feedback, which improve performance in tasks like Harness-Bench and NewtonBench.
3 days ago