ZimaBlue: Evolving Generalizable World Action Models through Scalable Video Pre-training Paper • 2609.00188 • Published 15 days ago • 53
GenFirst: Generation Before Reconstruction for Stable End-to-End Latent Generative Modeling Paper • 2608.29335 • Published 17 days ago • 72
CoinVE-200K: A Large-Scale High-Quality Dataset for Compositional Instruction-Guided Video Editing Paper • 2608.17566 • Published 28 days ago • 16
Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers Paper • 2607.28611 • Published Jul 30 • 23
BM25 Wins at Scale: A Scaling Study of Retrieval-Augmented Generation Paradigms Paper • 2607.26497 • Published Jul 30 • 53
Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation Paper • 2607.27816 • Published Jul 30 • 35
MPIE-Bench: Benchmarking Anatomically Plausible Multi-Person Interaction Editing Paper • 2607.27616 • Published Jul 30 • 40
Vera: A Layered Diffusion Model for Content-Preserving Video Editing Paper • 2606.23610 • Published Jun 22 • 13
Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models Paper • 2606.03988 • Published Jun 3 • 127
MilliVid: Hierarchical Latents for Long-Range Consistency in Video Generation Paper • 2606.09056 • Published Jun 8 • 8
minWM: A Full-Stack Open-Source Framework for Real-Time Interactive Video World Models Paper • 2605.30263 • Published May 28 • 62
CoVEBench: Can Video Editing Models Handle Complex Instructions? Paper • 2606.08415 • Published Jun 7 • 53
Running on Zero Agents Featured 79 VGGT-Omega Demo 🌀 79 3D reconstruction from images/video with VGGT-Omega
SANA-WM: Efficient Minute-Scale World Modeling with Hybrid Linear Diffusion Transformer Paper • 2605.15178 • Published May 14 • 91