How interest changes
History starts here
The chart will appear after repeat observations. The current metric comes from the source.
6 source votesReal observations only. History before source connection is not reconstructed.
We introduce LIFT, a unified image-to-video generation framework that complements camera control with Layout-In-FuTure control, enabling users to specify what should appear in a future view and where it should appear. This addresses a practical need in controllable video generation: given an initial image, users often care not only about how the camera moves, but also about what the scene should look like at key future moments, especially the final frame. Existing camera controls specify viewpoint trajectories, while text prompts provide only coarse semantic guidance; neither precisely determines the content and spatial layout of future views. This limitation becomes particularly pronounced under large viewpoint changes, where the camera reveals regions that are not visible in the first frame. LIFT therefore uses the last-frame layout as an explicit control signal for the desired future scene. Since learning from such sparse layout guidance is substantially more challenging than conditioning on dense per-frame layouts, we introduce on-policy self-distillation (OPSD) to transfer the control capability of a dense-layout teacher to a last-frame-layout student. We further curate LIFT-Vista, a dataset featuring large viewpoint changes with camera and temporally consistent layout annotations. Experiments show that LIFT improves video quality, future-layout controllability, and camera controllability over other methods.
Translation pending · showing the source descriptionThe chart will appear after repeat observations. The current metric comes from the source.
6 source votesReal observations only. History before source connection is not reconstructed.
The paper introduces VisionHOPE, a novel visual backbone that functions as a self-modifying learning system, allowing the model to co-evolve what it remembers and how it learns within an image. It uses five coupled memories and a stability-matched step-size control scheme to ensure stable learning dynamics, achieving competitive results on benchmark datasets like ImageNet-1K, COCO, and ADE20K.
3 days agoDaily PapersThis paper introduces FuseReg, a method that replaces heuristic layer fusion in representation autoencoders (RAEs) with training over random subsets of encoder layers. The approach reduces the reconstruction-generation gap by improving robustness to layer fusion choices, achieving higher PSNR and lower generation FID scores without modifying the pretrained encoder.
5 days agoDaily PapersThe paper explores the linearity in Large Language Models (LLMs) by showing that combining inputs from different text streams leads to a superposition of next-token distributions. It suggests that this linearity is an inherent property of the Transformer architecture and can be restored through fine-tuning, allowing for generating two coherent continuations from one forward pass.
6 days agoDaily PapersThe paper introduces Omni-IO Skills, a plug-and-play agent harness that enables existing agents to handle multiple modalities through hierarchical skills, standardized execution interfaces, and dependency-aware orchestration. It demonstrates significant improvements in input support and semantic quality scores when applied to GPT-5.6 Sol and Claude Sonnet 5 on the UniM-90 dataset.
5 days agoDaily PapersThe paper introduces GAGAR, a framework for quality-aware credit redistribution in code agent reinforcement learning (RL). It uses dynamic sampling and an SFT-trained agentic grader to rank test-passing trajectories, adjusting advantages to prioritize higher-quality implementations. The method was evaluated on large-scale industrial code agents with significant parameter counts.
4 days agoDaily PapersThis paper introduces SentZero, a sentence-centric vision-language pretraining framework designed for zero-shot multi-task analysis of chest X-rays. It enhances positive-pair diversity and mitigates false negatives through LLM-based sentence structuring and visual embedding modulation.
2 days ago