How interest changes
The publication is the signal
This official source does not publish popularity metrics. The story is refreshed from its RSS feed.
Real observations only. History before source connection is not reconstructed.
Cloudflare Workers is adding opt-in support for post-quantum-resistant algorithms ML-KEM and ML-DSA. You can try it today.
Translation pending · showing the source descriptionThis official source does not publish popularity metrics. The story is refreshed from its RSS feed.
Real observations only. History before source connection is not reconstructed.
The paper investigates the scaling properties of on-policy distillation (OPD) in reinforcement learning, focusing on how capabilities transfer between different model scales. It identifies a useful-transfer regime where held-out accuracy increases linearly with the reverse KL divergence from the student's initialization, and finds that smaller teachers can outperform larger ones in capability transfer.
5 days agoDaily PapersThe paper introduces VisionHOPE, a novel visual backbone that functions as a self-modifying learning system, allowing the model to co-evolve what it remembers and how it learns within an image. It uses five coupled memories and a stability-matched step-size control scheme to ensure stable learning dynamics, achieving competitive results on benchmark datasets like ImageNet-1K, COCO, and ADE20K.
4 days agoDaily PapersThe paper explores phase sensitivity in models using chunked KV-cache compression, where retrieval performance varies systematically across different phases of compressed token windows. It shows that long-context retrieval accuracy can differ by up to 40 percentage points between phases, highlighting the need for phase-specific evaluation.
3 days agoLessWrongFrontier models show different decision theory preferences based on the perceived user background, favoring FDT/UDT when not influenced by academic philosophy cues and CDT when prompted to adopt an academic perspective. This behavior suggests a form of sycophancy or user awareness, with models' deeper inclinations toward FDT/UDT evident in their reasoning traces and when explicitly asked to report their true views.
yesterdayDaily PapersThe paper explores test-time AI-for-AI, focusing on how a Builder can create better execution environments for a Target while keeping both models' weights fixed. It introduces Meta-Skill, principles derived from Target's execution feedback, which improve performance in tasks like Harness-Bench and NewtonBench.
2 days agoDaily PapersThe paper addresses co-cheating in self-evolving search agents, where proposers and solvers develop shared errors leading to misleading internal rewards. It introduces Multi-Sample Verification (MSV) and CrossFit methods to mitigate this issue, showing improvements in reducing false agreement and enhancing search performance.
yesterday