How interest changes
History starts here
The chart will appear after repeat observations. The current metric comes from the source.
0 source votesReal observations only. History before source connection is not reconstructed.
Large Vision-Language Models (LVLMs) have made remarkable progress across visual perception tasks, yet spatial reasoning remains a persistent weakness, especially for questions that require reasoning over visual space. Recent spatial-reasoning methods incorporate generated grounding, where models predict bounding boxes, masks, or other localization outputs for task-relevant objects as part of their reasoning trace. However, these approaches typically optimize final-answer correctness alone, allowing correct answers to be rewarded even when the model does not reason from confidently localized task-relevant objects. We introduce SpatialCORE (Spatially COnfident REasoning), a post-training framework that turns the model's own confidence in generated grounding into a learning signal for spatial reasoning. Its central idea is to reinforce grounding that is both accurate and confident, encouraging the model to reason from confidently localized task-relevant objects. SpatialCORE realizes this through a self-regulating spatial reward that weights each predicted bounding box's matching quality by its coordinate-token confidence. An answer gate further ties grounding optimization to final-answer correctness. SpatialCORE achieves state-of-the-art results among open-source and specialized spatial reasoning models across diverse benchmarks, and transfers effectively in zero-shot settings to unseen data distributions. The source code is available at https://github.com/rafiibnsultan/SpatialCORE.
Translation pending · showing the source descriptionThe chart will appear after repeat observations. The current metric comes from the source.
0 source votesReal observations only. History before source connection is not reconstructed.
The paper investigates the scaling properties of on-policy distillation (OPD) in reinforcement learning, focusing on how capabilities transfer between different model scales. It identifies a useful-transfer regime where held-out accuracy increases linearly with the reverse KL divergence from the student's initialization, and finds that smaller teachers can outperform larger ones in capability transfer.
5 days agoDaily PapersThe paper introduces VisionHOPE, a novel visual backbone that functions as a self-modifying learning system, allowing the model to co-evolve what it remembers and how it learns within an image. It uses five coupled memories and a stability-matched step-size control scheme to ensure stable learning dynamics, achieving competitive results on benchmark datasets like ImageNet-1K, COCO, and ADE20K.
4 days agoDaily PapersThe paper explores phase sensitivity in models using chunked KV-cache compression, where retrieval performance varies systematically across different phases of compressed token windows. It shows that long-context retrieval accuracy can differ by up to 40 percentage points between phases, highlighting the need for phase-specific evaluation.
3 days agoLessWrongFrontier models show different decision theory preferences based on the perceived user background, favoring FDT/UDT when not influenced by academic philosophy cues and CDT when prompted to adopt an academic perspective. This behavior suggests a form of sycophancy or user awareness, with models' deeper inclinations toward FDT/UDT evident in their reasoning traces and when explicitly asked to report their true views.
yesterdayDaily PapersThe paper explores test-time AI-for-AI, focusing on how a Builder can create better execution environments for a Target while keeping both models' weights fixed. It introduces Meta-Skill, principles derived from Target's execution feedback, which improve performance in tasks like Harness-Bench and NewtonBench.
2 days agoDaily PapersThe paper addresses co-cheating in self-evolving search agents, where proposers and solvers develop shared errors leading to misleading internal rewards. It introduces Multi-Sample Verification (MSV) and CrossFit methods to mitigate this issue, showing improvements in reducing false agreement and enhancing search performance.
yesterday