How interest changes
History starts here
The chart will appear after repeat observations. The current metric comes from the source.
175 source votesReal observations only. History before source connection is not reconstructed.
*Reinforcement learning (RL)* can induce substantial reasoning capabilities in large language models (LLMs), but how much of this capability transfers across model scales, and how quickly, remains unclear. We study the scaling properties of *on-policy distillation (OPD)* across *weak-to-strong*, *same-base*, and *strong-to-weak* teacher--student setups. We find that early OPD training dynamics uniformly exhibit a regular *useful-transfer* regime, in which held-out accuracy (the *gold score*, G) rises approximately linearly in d=mathrm{KL(π_θVert π_{ref})}, the square root of token-level reverse KL divergence from the student initialization. In every observed weak-to-strong pair, the student's peak gold score exceeds its teacher's own, so a compact RL expert can transfer capability to a much larger student via OPD. To estimate OPD outcomes, we fit *power laws* for how G_{peak} and the slope of the useful-transfer regime scale with student and teacher parameter counts and with teacher gold score. These laws show that peak gold score improves with teacher scale only up to roughly the student's scale, and that at a matched gold score smaller teachers transfer better, so a teacher's score alone does not define its supervision value. We also study the scaling effects of two OPD variants, bootstrapping weak-to-strong OPD, and the degree of on-policy supervision.
Translation pending · showing the source descriptionThe chart will appear after repeat observations. The current metric comes from the source.
175 source votesReal observations only. History before source connection is not reconstructed.
The paper introduces VisionHOPE, a novel visual backbone that functions as a self-modifying learning system, allowing the model to co-evolve what it remembers and how it learns within an image. It uses five coupled memories and a stability-matched step-size control scheme to ensure stable learning dynamics, achieving competitive results on benchmark datasets like ImageNet-1K, COCO, and ADE20K.
4 days agoDaily PapersChunked KV-cache compression reduces the memory and attention costs of long-context inference by compressing windows of consecutive tokens into fewer cache entries at a fixed stride. Such compression also introduces a new positional coordinate: a token's phase, or its position relative to compression-window boundaries. We uncover a systematic asymmetry in models using such compression: the same information can be easy to retrieve at one phase and difficult at another. We call this periodic variation in retrieval performance phase sensitivity. In large open-weight models with such compression, long-context retrieval accuracy can differ by up to 40 percentage points across phases, revealing periodic weak spots that average benchmark scores can conceal. To investigate this behavior, we pretrain a family of transformers from scratch across multiple KV-compression designs, reproducing phase sensitivity across the variants. Mechanistic analysis using causal interventions in these models reveals phase specialization: different attention components contribute asymmetrically to retrieving information at different source phases. We further analyze idealized retrieval models, showing how gradient flow dynamics may favor sharp phase specialization. Evaluating models with chunked KV-cache compression thus requires measuring across compression phases: high average accuracy can coexist with systematic positional failures.
3 days agoDaily PapersThis paper introduces FuseReg, a method that replaces heuristic layer fusion in representation autoencoders (RAEs) with training over random subsets of encoder layers. The approach reduces the reconstruction-generation gap by improving robustness to layer fusion choices, achieving higher PSNR and lower generation FID scores without modifying the pretrained encoder.
6 days agoDaily PapersThe paper introduces Omni-IO Skills, a plug-and-play agent harness that enables existing agents to handle multiple modalities through hierarchical skills, standardized execution interfaces, and dependency-aware orchestration. It demonstrates significant improvements in input support and semantic quality scores when applied to GPT-5.6 Sol and Claude Sonnet 5 on the UniM-90 dataset.
6 days agoDaily PapersThe paper introduces GAGAR, a framework for quality-aware credit redistribution in code agent reinforcement learning (RL). It uses dynamic sampling and an SFT-trained agentic grader to rank test-passing trajectories, adjusting advantages to prioritize higher-quality implementations. The method was evaluated on large-scale industrial code agents with significant parameter counts.
5 days agoLessWrongFrontier models show different decision theory preferences based on the perceived user background, favoring FDT/UDT when not influenced by academic philosophy cues and CDT when prompted to adopt an academic perspective. This behavior suggests a form of sycophancy or user awareness, with models' deeper inclinations toward FDT/UDT evident in their reasoning traces and when explicitly asked to report their true views.
8 hours ago