See, Point, Fly: A Learning-Free VLM Framework for Universal Unmanned Aerial Navigation Paper • 2509.22653 • Published Sep 26, 2025 • 25
Stream-DiffVSR: Low-Latency Streamable Video Super-Resolution via Auto-Regressive Diffusion Paper • 2512.23709 • Published Dec 29, 2025 • 51
GaMO: Geometry-aware Multi-view Diffusion Outpainting for Sparse-View 3D Reconstruction Paper • 2512.25073 • Published Dec 31, 2025 • 42
3AM: Segment Anything with Geometric Consistency in Videos Paper • 2601.08831 • Published Jan 13 • 34
LongE2V: Long-Horizon Event-based Video Reconstruction, Prediction, and Frame Interpolation with Video Diffusion Models Paper • 2607.08770 • Published Jul 9 • 38
Scal3R: Learning Efficient Multi-Relative Pose Query for Scalable Online 3D Reconstruction Paper • 2609.04201 • Published 9 days ago • 46
Scal3R: Learning Efficient Multi-Relative Pose Query for Scalable Online 3D Reconstruction Paper • 2609.04201 • Published 9 days ago • 46
Scal3R: Learning Efficient Multi-Relative Pose Query for Scalable Online 3D Reconstruction Paper • 2609.04201 • Published 9 days ago • 46
LongE2V: Long-Horizon Event-based Video Reconstruction, Prediction, and Frame Interpolation with Video Diffusion Models Paper • 2607.08770 • Published Jul 9 • 38
One Model, Many Latencies: Universal Speech Enhancement for Diverse Real-Time Applications Paper • 2606.25621 • Published Jun 24 • 23
Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients Paper • 2606.18216 • Published Jun 16 • 65
BRDFusion: Physics Meets Generation for Urban Scene Inverse Rendering Paper • 2606.17049 • Published Jun 15 • 29
SpatialClaw: Rethinking Action Interface for Agentic Spatial Reasoning Paper • 2606.13673 • Published Jun 11 • 113
Reroute, Don't Remove: Recoverable Visual Token Routing for Vision-Language Models Paper • 2606.12412 • Published Jun 10 • 21
YoCausal: How Far is Video Generation from World Model? A Causality Perspective Paper • 2605.30346 • Published May 28 • 58
Agent Explorative Policy Optimization for Multimodal Agentic Reasoning Paper • 2605.28774 • Published May 27 • 94
Pantheon360: Taming Digital Twin Generation via 3D-Aware 360° Video Diffusion Paper • 2605.25449 • Published May 25 • 24
Stroke of Surprise: Progressive Semantic Illusions in Vector Sketching Paper • 2602.12280 • Published Feb 12 • 34