Good Papers

Invariant Graph Transformer for Out-of-Distribution Generalization

GOODFormer improves graph transformer out-of-distribution generalization by jointly learning invariant predictive subgraphs, evolving positional encodings, and invariant representations.

Tianyin Liao, Ziwei Zhang, Yufei Sun, Chunyu Hu, Jianxin Li

Published Apr 20, 2026Paper ↗

59%
OverallWorth a look
?
OverallWorth a lookVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
?1 reader voted. Vote to see how they split.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel7/20reviewers recommend it
lenient 2/5
medium 5/10
strict 0/5
AI panel?Vote to see what the 20 AI reviewers said
Panel consensus
GOODFormer delivers a genuinely sharp invariant attention mechanism and evolving subgraph encodings that improve graph transformer OOD generalization, though its three-module design lacks clear ablation to isolate which component drives the gains or prove true predictive…

Abstract

Graph Transformers (GTs) have demonstrated great effectiveness across various graph analytical tasks. However, the existing GTs focus on training and testing graph data originated from the same distribution, but fail to generalize under distribution shifts. Graph invariant learning, aiming to capture generalizable graph structural patterns with labels under distribution shifts, is potentially a promising solution, but how to design attention mechanisms and positional and structural encodings (PSEs) based on graph invariant learning principles remains challenging. To solve these challenges, we introduce graph out-of-distribution generalized Transformer (GOODFormer), aiming to learn generalized graph representations by capturing invariant relationships between predictive graph structures and labels through jointly optimizing three modules. Specifically, we first develop a GT-based entropy-guided invariant subgraph disentangler to separate invariant and variant subgraphs while preserving the sharpness of the attention function. Next, we design an evolving subgraph positional and structural encoder to effectively and efficiently capture the encoding information of dynamically changing subgraphs during training. Finally, we propose an invariant learning module utilizing subgraph node representations and encodings to derive graph representations that can generalize to unseen test graphs. We also provide theoretical justifications for our method. Extensive experiments on benchmark datasets demonstrate the superiority of our method over state-of-the-art baselines under distribution shifts.