Good Papers

Let the Heads Talk: Beyond Diagonal Graph Attention

Top-A learns edge-conditioned off-diagonal cross-head routes in multi-head attention that preserve diagonal paths, improving interaction-dependent tasks without benefiting heterophily.

Riccardo Ali, Alessio Borgi, Mario Severino, Alessio Gravina, Davide Bacciu, Pietro Liò, Christopher Irwin

Published Oct 1, 2026arXiv ↗

78%
OverallHighly rated
?
OverallHighly ratedVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel11/20reviewers recommend it
lenient 2/5
medium 8/10
strict 1/5
AI panel?Vote to see what the 20 AI reviewers said
Panel consensus
Top-A establishes cross-head attention routing as a distinct transport primitive and clarifies that interaction-dependent transforms, not heterophily alone, drive gains, though it lacks assay-level validation and large-scale baselines to fully justify its mechanics beyond benchmarks.

Abstract

Sheaf Neural Networks generalize scalar-weighted message passing by replacing scalar edge weights with linear transport maps between local feature spaces. Yet the role of this matrix-valued transport is entangled with the broader sheaf-diffusion construction. We isolate the transport primitive through quiver representations and establish a direct connection with multi-head attention. Treating attention heads as coordinates of a local transport space reveals that standard multi-head attention implements diagonal edge maps: along each directed interaction, a source head can contribute only to the corresponding receiver head. Allowing off-diagonal entries instead enables edge-conditioned communication across heads before neighborhood aggregation. We show that this operation cannot, in general, be absorbed into a single shared linear map applied after aggregation. Building on this characterization, we introduce Topological Attention (Top-A), a multi-head attention that learns edge-dependent off-diagonal routes while preserving the original same-head paths and exactly recovering vanilla attention when the additional routing vanishes. We evaluate Top-A on relational reasoning, heterogeneous graph learning, and algorithmic reasoning, including out-of-distribution generalization, with heterophilic node classification as a contrast setting. The results show that cross-head transport is most useful when the task benefits from interaction-dependent transformations, while heterophily alone provides no systematic advantage. These findings identify edge-conditioned cross-head communication as a distinct computational primitive of matrix-valued transport.