How interest changes
History starts here
The chart will appear after repeat observations. The current metric comes from the source.
122 source pointsReal observations only. History before source connection is not reconstructed.
OpenAI has withdrawn three mathematics papers.
The chart will appear after repeat observations. The current metric comes from the source.
122 source pointsReal observations only. History before source connection is not reconstructed.
CheckerBench is a benchmark for evaluating coding agents' ability to synthesize static-analysis checkers from defect specifications, including tasks derived from 300 CVEs across multiple repositories and language ecosystems. CheckerLab, an evaluation framework, measures diagnostic contrast, patch localization, and tool use, with results showing that current agents struggle to develop reliable checkers, achieving a maximum Pass@1 of 45.33%.
2 days agoDaily PapersTetris3D is a generative framework for 3D scene reconstruction from a single image, ensuring objects are geometrically and physically coherent. It uses explicit conditioning on surrounding objects' geometry and physical relationships, along with a physics-based dataset called ComOb for training and evaluation.
yesterdayDaily PapersThe paper investigates on-policy distillation (OPD) in cross-tokenizer settings, focusing on alignment coverage and supervision reliability. It finds that strict 1:1 token alignment covers most student-generated tokens despite vocabulary mismatches, and that restricting reverse KL to a top-16 subset of shared vocabulary achieves comparable accuracy to full shared-vocabulary OPD, while adding span supervision reduces accuracy.
2 days agoDaily PapersThe paper explores on-policy distillation (OPD) in language model post-training, explaining how it can lead to either performance gains or generation collapse. It suggests that OPD amplifies student behaviors favored by the teacher's implicit feedback, and proposes methods like masking unhealthy responses and SFT initialization to mitigate collapse.
6 days agoDaily PapersThe paper introduces WING, a framework that transfers interaction knowledge from human egocentric videos to robot policies by separating observer-induced motion from hand-object interactions and using spectral analysis to identify shared temporal structures between human and robot behaviors. It achieves high success rates on multiple benchmarks and demonstrates strong performance in real-world tasks.
6 days agoDaily PapersThe paper introduces NAVA-WAM, a method for pretraining action policies directly from observation-only videos without relying on action-annotated robot data. It uses two stages: first, pretraining on videos with visual transition supervision, and second, post-training with action-labeled demonstrations to improve robot control.
6 days ago