Good Papers

RegimeVGGT: Layer-Wise Spatially Preserving Redundancy Removal for Visual Geometry Grounded Transformer

RegimeVGGT removes layer-wise spatial redundancy in VGGT via U-shaped cross-frame compression, yielding 6.7x speedup with preserved geometry and pose accuracy.

Shuo Lyu, Jinhao You, Zhuohang Lyu, Tanxuan Li, Zibo Zhao, Jiaxiang Hu, Kai Tang, Yichen Guo

Published 2026Atlanta Poster Session 2 · Wed, Dec 9, 4:30 PM–7:30 PM local time · Hall C1arXiv ↗OpenReview ↗

76%
OverallHighly rated
?
OverallHighly ratedVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel10/20reviewers recommend it
lenient 2/5
medium 7/10
strict 1/5
AI panel?Vote to see what the 20 AI reviewers said
Panel consensus
RegimeVGGT delivers a rigorous, training-free 6.7x speedup via layer-wise spectral compression that preserves geometry and pose, though it lacks variance estimates, cross-architecture validation, and released code to confirm real-world latency and collapsed-overlap robustness.

Abstract

Visual Geometry Grounded Transformer (VGGT) recovers dense 3D scene structure from multi-view images in one forward pass, but quadratic cross-frame attention limits its scalability. Existing training-free accelerators reduce computation uniformly along one axis, missing layer heterogeneity. Our spectral, probing, and causal analyses reveal three regimes: shallow layers lack cross-view structure, middle layers drive cross-view alignment, and deep layers are redundant for dense geometry yet their cross-frame attention remains essential for pose. RegimeVGGT applies layer-wise U-shaped compression along two axes: Saliency-Guided Banded Merging protects geometry- and edge-salient tokens, while Selectively Protected K/V Downsampling preserves cross-frame spatial coverage and the pose-critical path through a phase-shifted spatial grid, a reference-frame anchor, and uncompressed camera/register tokens. Training-free, RegimeVGGT achieves a 6.7x speedup over VGGT* at matched reconstruction quality.