How interest changes
History starts here
The chart will appear after repeat observations. The current metric comes from the source.
1 source votesReal observations only. History before source connection is not reconstructed.
Agents increasingly build on code written by other agents, and they reimplement rather than reuse, growing the codebases later agents must work in. To measure how well agents design libraries for other agents, we introduce LibraryDesignBench, a two-phase benchmark in which an agent implements a full-featured library from a specification that defines required capabilities and potential use cases without prescribing the design. We evaluate the library through the correctness and simplicity of programs written by three user agents from different model families. The benchmark spans 242 expert-validated programming problems across 15 library-design tasks in four languages. On eleven of the fifteen tasks, agent designers reproduce the abstractions of the human-written production library. Downstream agents adopt agent- and human-written libraries alike but underuse them, reimplementing capabilities the library already provides. Our failure analysis finds that downstream agents write extra code mainly because agent-written libraries are rigid or hard to use, not because capabilities are missing. We also experiment with giving designers more prescriptive, agent-first guidance and having them test their library with subagents; this improves downstream scores and yields simpler programs. LibraryDesignBench provides both a testbed for evaluating library-design practices for agent users and an initial design baseline that improves downstream reuse.
Translation pending · showing the source descriptionThe chart will appear after repeat observations. The current metric comes from the source.
1 source votesReal observations only. History before source connection is not reconstructed.
The paper introduces VisionHOPE, a novel visual backbone that functions as a self-modifying learning system, allowing the model to co-evolve what it remembers and how it learns within an image. It uses five coupled memories and a stability-matched step-size control scheme to ensure stable learning dynamics, achieving competitive results on benchmark datasets like ImageNet-1K, COCO, and ADE20K.
3 days agoDaily PapersThe paper explores the linearity in Large Language Models (LLMs) by showing that combining inputs from different text streams leads to a superposition of next-token distributions. It suggests that this linearity is an inherent property of the Transformer architecture and can be restored through fine-tuning, allowing for generating two coherent continuations from one forward pass.
6 days agoDaily PapersThis paper introduces FuseReg, a method that replaces heuristic layer fusion in representation autoencoders (RAEs) with training over random subsets of encoder layers. The approach reduces the reconstruction-generation gap by improving robustness to layer fusion choices, achieving higher PSNR and lower generation FID scores without modifying the pretrained encoder.
5 days agoDaily PapersThe paper introduces Omni-IO Skills, a plug-and-play agent harness that enables existing agents to handle multiple modalities through hierarchical skills, standardized execution interfaces, and dependency-aware orchestration. It demonstrates significant improvements in input support and semantic quality scores when applied to GPT-5.6 Sol and Claude Sonnet 5 on the UniM-90 dataset.
5 days agoDaily PapersThe paper introduces GAGAR, a framework for quality-aware credit redistribution in code agent reinforcement learning (RL). It uses dynamic sampling and an SFT-trained agentic grader to rank test-passing trajectories, adjusting advantages to prioritize higher-quality implementations. The method was evaluated on large-scale industrial code agents with significant parameter counts.
4 days agoDaily PapersThis paper introduces SentZero, a sentence-centric vision-language pretraining framework designed for zero-shot multi-task analysis of chest X-rays. It enhances positive-pair diversity and mitigates false negatives through LLM-based sentence structuring and visual embedding modulation.
2 days ago