Good Papers
IEEE International Conference on Acoustics Speech and Signal Processing 2026Object detectionSoochow

Linear Cross-Attention Guided Feature Pyramid Networks for Crowd Counting

Linear cross-attention modules reduce semantic gaps across pyramid features to improve crowd counting, and VMambaCC achieves state-of-the-art results in dense scenes.

Hao-Yuan Ma, Li Zhang, Shuai Shi

Published Apr 21, 2026Paper ↗

73%
OverallHighly rated
?
OverallHighly ratedVote to see the score
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel9/20reviewers recommend it
lenient 5/5
medium 4/10
strict 0/5
AI panel?Vote to see what the 20 AI reviewers said

Abstract

Drastic scale variations and the semantic gap between features from different network depths pose significant challenges for crowd counting. To address this issue, we conceptualize these features as distinct pseudo-modalities and propose a Linear Cross-Attention Module (LCAM) to effectively alleviate the semantic gap between them. Build upon LCAM, we develop a Linear Cross-Attention guided Feature Pyramid Network (LCA-FPN), which achieves superior multi-scale feature fusion. We further develop Visual Mamba for Crowd Counting (VMambaCC), a novel architecture leveraging LCA-FPN for feature alignment and Mamba’s selective state space model to efficiently capture long-range dependencies. Extensive experiments on four challenging datasets validate the effectiveness of the proposed approach. VMambaCC achieves competitive performance, demonstrating substantial improvements in counting and localization, particularly in highly congested scenes. Code is available at: https://github.com/Jason-Mar1/VMambaCC.