Mask Forcing: Improving Autoregressive Video Diffusion Distillation via Dual-Noise Masking Rollout Paper • 2609.09123 • Published 2 days ago • 43
Mask Forcing: Improving Autoregressive Video Diffusion Distillation via Dual-Noise Masking Rollout Paper • 2609.09123 • Published 2 days ago • 43
SolarWM: Open Data and Scalable Training for Long-Horizon Video World Models Paper • 2609.02886 • Published 8 days ago • 146
UniVR: Thinking in Visual Space for Unified Visual Reasoning Paper • 2607.12800 • Published Jul 14 • 32
SPIRAL: Self-Evolving Action-Conditioned Video Generation via Reflective Planning Agents Paper • 2603.08403 • Published May 21
Unified Generative and Discriminative Training for Multi-modal Large Language Models Paper • 2411.00304 • Published Nov 1, 2024
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Paper • 2505.24164 • Published May 30, 2025
ROVER: Benchmarking Reciprocal Cross-Modal Reasoning for Omnimodal Generation Paper • 2511.01163 • Published Nov 3, 2025 • 32
WEAVE: Unleashing and Benchmarking the In-context Interleaved Comprehension and Generation Paper • 2511.11434 • Published Nov 14, 2025 • 47
EditMGT: Unleashing Potentials of Masked Generative Transformers in Image Editing Paper • 2512.11715 • Published Dec 12, 2025
Agentic World Modeling: Foundations, Capabilities, Laws, and Beyond Paper • 2604.22748 • Published Apr 24 • 234
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Paper • 2606.19534 • Published Jun 17 • 66
NVIDIA Nemotron Nano 2: An Accurate and Efficient Hybrid Mamba-Transformer Reasoning Model Paper • 2508.14444 • Published Aug 20, 2025 • 51