RegimeVGGT: Layer-Wise Spatially Preserving Redundancy Removal for Visual Geometry Grounded Transformer
RegimeVGGT removes layer-wise spatial redundancy in VGGT via U-shaped cross-frame compression, yielding 6.7x speedup with preserved geometry and pose accuracy.
Published 2026Atlanta Poster Session 2 · Wed, Dec 9, 4:30 PM–7:30 PM local time · Hall C1arXiv ↗OpenReview ↗

Only vote on papers you've read. Sign in with GitHub to vote.
RegimeVGGT delivers a rigorous, training-free 6.7x speedup via layer-wise spectral compression that preserves geometry and pose, though it lacks variance estimates, cross-architecture validation, and released code to confirm real-world latency and collapsed-overlap robustness.
Abstract
Visual Geometry Grounded Transformer (VGGT) recovers dense 3D scene structure from multi-view images in one forward pass, but quadratic cross-frame attention limits its scalability. Existing training-free accelerators reduce computation uniformly along one axis, missing layer heterogeneity. Our spectral, probing, and causal analyses reveal three regimes: shallow layers lack cross-view structure, middle layers drive cross-view alignment, and deep layers are redundant for dense geometry yet their cross-frame attention remains essential for pose. RegimeVGGT applies layer-wise U-shaped compression along two axes: Saliency-Guided Banded Merging protects geometry- and edge-salient tokens, while Selectively Protected K/V Downsampling preserves cross-frame spatial coverage and the pose-critical path through a phase-shifted spatial grid, a reference-frame anchor, and uncompressed camera/register tokens. Training-free, RegimeVGGT achieves a 6.7x speedup over VGGT* at matched reconstruction quality.