GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? Paper • 2608.05747 • Published Aug 6 • 47
Can Understanding and Generation Truly Benefit Together -- or Just Coexist? Paper • 2509.09666 • Published Sep 11, 2025 • 34
UNIMO-G: Unified Image Generation through Multimodal Conditional Diffusion Paper • 2401.13388 • Published Jan 24, 2024 • 13