How interest changes
The publication is the signal
This official source does not publish popularity metrics. The story is refreshed from its RSS feed.
Real observations only. History before source connection is not reconstructed.
The article discusses how UK AISI and EvalEval are working to make benchmark results reproducible, though no detailed information is provided.
This official source does not publish popularity metrics. The story is refreshed from its RSS feed.
Real observations only. History before source connection is not reconstructed.
This paper evaluates the ability of the MiniMax-H3 multimodal generative model to reason about the physical world. The study introduces a new framework for assessing physical reasoning through four dimensions, using tasks that require integrating information across multiple modalities. The model achieves an overall success rate of 41.97% across 517 evaluation instances, with video-based reasoning performing best and audio-based reasoning least effectively.
6 days agoDaily PapersIntBMoE introduces a block-conditioned Mixture-of-Experts (MoE) framework that decouples participation, execution, and materialization by combining dense expert composition with sparse block execution. It uses a learned codebook for blocks and a hypernetwork to merge experts, achieving full participation while keeping execution and memory costs low. Experiments show improvements in image classification, language modeling, and sequential recommendation, with real-world deployment in AMap's recommendation system.
4 days agoDaily PapersThe paper introduces SELF-INDEX, a framework that enables an index to self-evolve without human intervention by autonomously diagnosing retrieval issues, revising index keys, and validating changes. It also uses a Query Simulator to proactively explore new retrieval demands, improving performance across various corpora and downstream applications.
5 days agoDaily PapersThe paper introduces Document Retrieval-Aware Chunking (D-RAC), a method for efficiently processing enterprise documents by normalizing them into PDF and converting them into retrieval-optimized Markdown using a multimodal LLM. This approach preserves document structure, reduces token costs, and improves chunking efficiency compared to traditional methods.
yesterdayDaily PapersThe paper introduces ScienceIDE, a framework that transforms scientific code repositories into programmable environments for training scientific agents. These environments enable task generation, execution, and verification, supporting supervised fine-tuning and reinforcement learning. The approach demonstrates improvements in scientific code repair and general-purpose benchmarks.
6 days agoDaily PapersThe paper identifies a problem called Value Flattening in Proximal Policy Optimization (PPO) critics, where estimated state values change sharply while critic predictions remain flat. It introduces SP^3O, a method that applies value loss to only a few states per response, showing improvements in policy learning.
6 days ago