AREX-2: Advancing Self-Improving Agents through Long-Horizon Reflective Tasks Paper • 2609.38288 • Published 9 days ago • 140
PanoVLN: Towards Effective Panoramic Vision-and-Language Navigation Paper • 2609.34759 • Published 10 days ago • 167
All modalities are equal, but video is more equal: Closing the Cross-Attention Gap in Joint Video Generation Paper • 2609.27901 • Published 15 days ago • 24
The Past Frames the Future: Memory for Autoregressive Video Generation Paper • 2609.28466 • Published 15 days ago • 64
Geometric and Semantic Coupling for Interaction Understanding in 3D Scenes Paper • 2609.25247 • Published 17 days ago • 10
From Pattern Recognizers to Personalized Companions: A Survey of Large Language Models in Mental Health Paper • 2609.25186 • Published 17 days ago • 32
WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory Paper • 2609.24984 • Published 17 days ago • 157
OmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialogue Paper • 2609.21465 • Published 20 days ago • 151
Can MiniMax-H3 Reason About the Physical World? An Evaluation of Omni-Modal Generative Model Paper • 2609.18323 • Published 22 days ago • 133
VC-Attention: Value Smoothing and Softmax Casting for Low-bit Attention Paper • 2609.15810 • Published 24 days ago • 51