Lucida: Parse, Generate, and Place for Composable Real-to-Sim Scene Modeling Paper • 2608.30821 • Published 23 days ago • 67
Zero-WAM: In-Context World-Action Modeling from Human Videos for Open-Ended Task Generalization Paper • 2608.26103 • Published 28 days ago • 25
ReViV: Reconstructing the Viewer and the View in 4D from Monocular Egocentric Video Paper • 2607.17790 • Published Jul 20 • 5
Economy of Minds: Emerging Multi-Agent Intelligence with Economic Interactions Paper • 2606.02859 • Published Jun 1 • 11
Crafter: A Multi-Agent Harness for Editable Scientific Figure Generation from Diverse Inputs Paper • 2605.30611 • Published May 28 • 65
When Do Diffusion Models learn to Generate Multiple Objects? Paper • 2605.00273 • Published Apr 30 • 9
CityRAG: Stepping Into a City via Spatially-Grounded Video Generation Paper • 2604.19741 • Published Apr 21 • 18