ShotBench: Expert-Level Cinematic Understanding in Vision-Language Models Paper • 2506.21356 • Published Jun 26, 2025 • 22
Cut2Next: Generating Next Shot via In-Context Tuning Paper • 2508.08244 • Published Aug 11, 2025 • 13
Skywork UniPic 2.0: Building Kontext Model with Online RL for Unified Multimodal Model Paper • 2509.04548 • Published Sep 4, 2025 • 6
Uni-MMMU: A Massive Multi-discipline Multimodal Unified Benchmark Paper • 2510.13759 • Published Oct 15, 2025 • 11
Architecture Decoupling Is Not All You Need For Unified Multimodal Model Paper • 2511.22663 • Published Nov 27, 2025 • 29
Skywork UniPic 3.0: Unified Multi-Image Composition via Sequence Modeling Paper • 2601.15664 • Published Jan 22 • 4
HiAR: Efficient Autoregressive Long Video Generation via Hierarchical Denoising Paper • 2603.08703 • Published Mar 9 • 32
EchoGen: Cycle-Consistent Learning for Unified Layout-Image Generation and Understanding Paper • 2603.18001 • Published Mar 18 • 2
Uni-Edit: Intelligent Editing Is A General Task For Unified Model Tuning Paper • 2605.21487 • Published May 20 • 23
Video-IFBench: Evaluating Instruction Following of Multimodal LLMs in Video Understanding Scenarios Paper • 2608.25529 • Published 18 days ago • 17
EnvHarness: Awakening Static Worlds for Agent Learning Paper • 2608.19880 • Published 24 days ago • 275
InterleaveThinker: Reinforcing Agentic Interleaved Generation Paper • 2606.13679 • Published Jun 11 • 85
Uni-Edit: Intelligent Editing Is A General Task For Unified Model Tuning Paper • 2605.21487 • Published May 20 • 23