Papers
arxiv:2609.33253

VGGT-Diff: Visual Geometry Meets Diffusion for Sparse-View Novel View Synthesis

Published on Sep 27
· Submitted by
chenkangjie1123
on Sep 29
Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,
,

Abstract

We present VGGT-Diff, a geometry-routed multi-view diffusion model for sparse-view novel view synthesis. Existing novel view synthesis (NVS) methods face a fundamental trade-off: reconstruction-based approaches preserve observed geometry but struggle to synthesize unseen regions, while diffusion-based methods provide strong generative priors yet rely on implicit source-to-query correspondence. VGGT-Diff bridges these regimes by routing visual geometry latents from VGGT-Ω into a pretrained video diffusion model. Each visual token is associated with a 3D point and confidence, then transformed into query-aligned latent conditions through a confidence-aware Visual Geometry Router (VGR) that preserves front and back surface evidence. These conditions guide joint target-view denoising, while Point-Track Residual Consistency (PTRC) regularizes predicted-clean residuals along reliable 3D tracks, improving multi-view stability. We further introduce robust geometry conditioning, combining training-time regularization with inference-time guidance for improved robustness. Experiments show competitive or state-of-the-art performance across interpolation and extrapolation under different viewpoint difficulties. Our code is available at https://github.com/chenkangjie1123/VGGT-Diff.

Community

VGGT-Diff_Figure1_teaser_600dpi

VGGT-Diff is a geometry-routed multi-view diffusion model for sparse-view novel view synthesis from six input images. Its first key innovation, the confidence-aware Visual Geometry Router (VGR), transforms VGGT-Ω features into query-aligned geometric conditions while preserving both front- and back-surface evidence. Its second innovation, Point-Track Residual Consistency (PTRC), regularizes denoising residuals along reliable 3D tracks to improve cross-view geometric consistency.

Despite being trained on only 1K scenes with a relatively limited training budget, VGGT-Diff already delivers competitive or state-of-the-art novel-view synthesis performance. It achieves state-of-the-art PSNR / LPIPS in different viewpoint-difficulty settings, with particularly strong results on challenging mid- and far-range interpolation and extrapolation views. Further scaling in both training data and optimization steps could continue improving visual quality and cross-view consistency.

VGGT-Diff supports both pose-aware inference with known camera parameters and pose-free inference directly from six RGB images. The current effective 21K checkpoint—initialized from a 20K half-resolution checkpoint and adapted with only 1K additional full-resolution training steps—already generates high-quality 480p, 80-frame continuous camera-trajectory videos. The model is still being actively trained, and we plan to release multiple fully trained checkpoints to the community later.

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.33253
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.33253 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.33253 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.33253 in a Space README.md to link it from this page.

Collections including this paper 1