How interest changes
History starts here
The chart will appear after repeat observations. The current metric comes from the source.
7 source pointsReal observations only. History before source connection is not reconstructed.
Anthropic is updating its usage policy to prohibit 'sustained and needless abusive or cruel behavior toward their models.' Research on model welfare (the possible moral status and experiential preferences of AI models) is becoming a significant topic in the AI industry. The text discusses arguments for and against the concept of model welfare.
The chart will appear after repeat observations. The current metric comes from the source.
7 source pointsReal observations only. History before source connection is not reconstructed.
The paper introduces Memento 3, a system enabling frozen LLM agents to continuously learn explicit world models through external memory. The agent maintains a natural-language rulebook as persistent semantic memory, which is compiled into executable code for prediction and planning. Through a loop of observation, reflection, and verification, the agent refines its model and uses verified updates to guide interaction. On ARC-AGI-3, the single-model agent achieves 100.0 mean Relative Human Action Efficiency and uses 44% of the human action count. In an Atari Pong case study, a learned feedback controller wins 21:0 in each of three evaluated episodes.
2 days agoDaily PapersThe paper introduces Trace2Env, a learning-free framework that uses historical interaction traces to create a reusable environment worldbook for simulating realistic environments without rebuilding the original system. It improves next-observation fidelity and long-horizon interaction consistency compared to conventional prompt-based language world models.
5 days agoDaily PapersCheckerBench is a benchmark for evaluating coding agents' ability to synthesize static-analysis checkers from defect specifications, including tasks derived from 300 CVEs across multiple repositories and language ecosystems. CheckerLab, an evaluation framework, measures diagnostic contrast, patch localization, and tool use, with results showing that current agents struggle to develop reliable checkers, achieving a maximum Pass@1 of 45.33%.
4 days agoDaily PapersTetris3D is a generative framework for 3D scene reconstruction from a single image, ensuring objects are geometrically and physically coherent. It uses explicit conditioning on surrounding objects' geometry and physical relationships, along with a physics-based dataset called ComOb for training and evaluation.
3 days agoDaily PapersThe paper investigates on-policy distillation (OPD) in cross-tokenizer settings, focusing on alignment coverage and supervision reliability. It finds that strict 1:1 token alignment covers most student-generated tokens despite vocabulary mismatches, and that restricting reverse KL to a top-16 subset of shared vocabulary achieves comparable accuracy to full shared-vocabulary OPD, while adding span supervision reduces accuracy.
4 days agoDaily PapersSparse attention is used to reduce latency in long-sequence generation, but existing methods can degrade generation quality at high sparsity. MC-Sparse is a training-free framework that selects individual key-value tokens and organizes similar queries into groups for efficient GPU execution.
5 days ago