Linear Cross-Attention Guided Feature Pyramid Networks for Crowd Counting
Linear cross-attention modules reduce semantic gaps across pyramid features to improve crowd counting, and VMambaCC achieves state-of-the-art results in dense scenes.
Published Apr 21, 2026Paper ↗
Only vote on papers you've read. Sign in with GitHub to vote.
Abstract
Drastic scale variations and the semantic gap between features from different network depths pose significant challenges for crowd counting. To address this issue, we conceptualize these features as distinct pseudo-modalities and propose a Linear Cross-Attention Module (LCAM) to effectively alleviate the semantic gap between them. Build upon LCAM, we develop a Linear Cross-Attention guided Feature Pyramid Network (LCA-FPN), which achieves superior multi-scale feature fusion. We further develop Visual Mamba for Crowd Counting (VMambaCC), a novel architecture leveraging LCA-FPN for feature alignment and Mamba’s selective state space model to efficiently capture long-range dependencies. Extensive experiments on four challenging datasets validate the effectiveness of the proposed approach. VMambaCC achieves competitive performance, demonstrating substantial improvements in counting and localization, particularly in highly congested scenes. Code is available at: https://github.com/Jason-Mar1/VMambaCC.