Good Papers

Showing Multi-agent RL Show all papers

65%Worth a look
?Worth a lookVote to see the score

Flowing Faster to Coordinate: One-Step Online Multi-Agent Flow Policies

OMAF proposes a one-step flow policy framework for online multi-agent reinforcement learning, achieving up to 3.4x higher returns and 10.5x sample efficiency over baselines.

Zhuoran Li, Yunzhan Li, Xun Wang, Yihan Du and 1 more

Published Oct 1, 2026 · 0 citations

0% Readers0 of 1 upvoted
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 1/5
54%Worth a look
?Worth a lookVote to see the score

Rhombus: Incentivizing Coordination in Parallel Thinking through Reinforcement Learning

Rhombus uses reinforcement learning to incentivize coordination in parallel thinking frameworks.

Ziyuan Nan, Qi Yi, Di Huang, Yutong Wu and 8 more

Published 2026 · 0 citations

100% Readers1 of 1 upvoted
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score
Chapman & Hall/CRC eBooks 2025Multi-agent RL

Multi-Agent RL (MARL) Algorithms

The chapter explains multi-agent reinforcement learning algorithms by generalizing single-agent methods via reward machines.

Vinod K. Mishra

Published Oct 28, 2025 · 0 citations

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
88%Must read
?Must readVote to see the score

HyperMARL: Adaptive Hypernetworks for Multi-Agent RL

HyperMARL uses agent-conditioned hypernetworks to generate agent-specific parameters that decouple gradients, reducing variance and preserving behavioral diversity across multi-agent benchmarks without added complexity.

Kale-ab Tessera, Arrasy Rahman, Amos Storkey, Stefano Albrecht

Published 2025 · 0 citations

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 4/5
medium 10/10
strict 1/5
45%Niche pick
?Niche pickVote to see the score

Adaptive Communication Range for Scalable Cooperative Multi-Agent Reinforcement Learning

Huizhong Song, Wei Wei, Lin Li, Lijun Zhang and 1 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Coupling-Aware Reinforcement Learning for Co-Evolving Graph Games

Mina Kim, Guanghui Lan, Benoit Montreuil

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Structural Entropy Optimized Communication for Multi-Agent Reinforcement Learning

Wei Du, Benyu Wu, Wei Guo, Zhongmin Yan and 2 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

High Entropy Regularization Leads to Symmetry Equivariant Policies in Dec-POMDPs

Johannes Forkel, Constantin Ruhdorfer, Michael Beukman, Andreas Bulling and 1 more

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Scalable Multi-Agent Contrastive Reinforcement Learning

Victor Augusto Kich, Satoshi Yamamori, Jun Morimoto

Paris Poster Session 5, Fri, Dec 11, 11:30 AM–1:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Why Copy Others? Insights into Social Learning from Multi-Agent Reinforcement Learning

Yancheng Liang, Shakti Senthil, Daphne Chen, Simon Du and 1 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Policy-Level Exploration for Coordinated Multi-Agent Reinforcement Learning

Yucong Zhang, Chao Yu

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 18 reviewers recommend it
lenient 0/5
medium 0/8
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

CorridorLight: Cooperation as Task Negotiation with Causal Gating for Traffic Signal Control

Pak Lon Ip, Pengfei Ren, Yuteng Lin, Rongqin Chen and 2 more

Paris Poster Session 6, Fri, Dec 11, 2:30 PM–4:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

From Local to Global: Progressive Consensus via Hierarchical Communication in Multi-Agent Reinforcement Learning

Jiangjin Yin, Zhonglin Lv, Hangyu Mao, Rongbo Zhu and 4 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

DynamAuction: a reinforcement learning environment for repeated auction with dynamic value

Benjamin Heymann, Eugénie Patard, Corentin Pla, Patrick Loiseau

Paris Poster Session 3, Thu, Dec 10, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Policy Regret Minimization in Partially Observable Markov Games

Lan Sang, Raman Arora, Thanh Nguyen-Tang

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 2 reviewers recommend it
lenient 0/2
57%Worth a look
?Worth a lookVote to see the score

Toward Online Robust Zero-Sum Markov Games with Function Approximation

Chenhao Zhou, Jiyu Wei

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Know Your Task, Learn It Right: Task-Aware Optimistic Value Learning for Multi-Task Multi-Agent Reinforcement Learning

Chang Liu, Mengyang Li

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

Is Decentralized LLM Agent RL Robust to Heterogeneity? An Asymmetric Tale

Canyu Chen, Kangyu Zhu, Zhaorun Chen, Zhanhui Zhou and 5 more

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 1/5
45%Niche pick
?Niche pickVote to see the score

Coordination Connectivity: Shared Initialization Shapes the Joint-Policy Landscape in MARL

Yixiang Fan, Zhiqiang Pu, Hao Ma, Dongmin Li and 4 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Cooperative Multi-Agent Reinforcement Learning via Epigraph-Form Guided Exploration

Sungil Son, Hoseong Jung, Dahyun Oh, H. Jin Kim

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Improved Sample Complexity for Markov Games via Variance-Aware Bandit Learning

Hanbin Zhou, Canzhe Zhao, Shuai Li

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Safe-Fair MACPO: Burden-Fair Constrained Policy Optimization for Safe Multi-Agent Reinforcement Learning

Ankita Kushwaha, KIRAN RAVISH, Preeti, Pawan Kumar

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

StarCraft Motion: A Dataset for Agent Simulation in Adversarial and Partially Observable Scenarios

Yi-Chung Chen, Mingyu Kim, Ruqi Bai, James Z Hare and 2 more

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Coordinating Hundreds of RL Agents through Scalable Inference-Time Search

Daniel Rajaonarivonivelomanantsoa, Oussama Hidaoui, Refiloe Shabe, Noah De Nicola and 12 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Interaction Value Inference for Multi-Agent Reinforcement Learning via a Hierarchical Agent-Centric World Model

Zhuoran Chen, Xuyang Lu, Zeyang Liu, Xinrui Yang and 5 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

Submodular Multi-Agent Reinforcement Learning for Effective Online Distributed Task Allocation

Jing Liu, Yangyang YANG, Luca Ballotta, Fangfei Li and 2 more

Paris Poster Session 4, Thu, Dec 10, 5:30 PM–7:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 2/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Semantic-level Exploration for Multi-Agent Reinforcement Learning

Jiangjin Yin, Zijian Ye, Rongbo Zhu, Hangyu Mao and 1 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

DDACOM: Dual-Decoupled Adaptive Communication for Multi-Agent Reinforcement Learning under Dynamic Networks

Yiqun Wu, QI WANG, Mengxian Li, Yongjun Xu

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

CHAIN: Continual Heterogeneous Cooperation with Information Bottleneck for Multi-Agent Reinforcement Learning

Haowen Dou, Lujuan Dang, Badong Chen

Paris Poster Session 2, Wed, Dec 9, 5:00 PM–7:00 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Mind the Gap: Information Disadvantage as a Learning Signal in Cooperative MARL

Chang Liu

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

CAFE: Causally-Guided Automated Feature Engineering with Multi-Agent Reinforcement Learning

Arun Vignesh Malarkkan, Wangyang Ying, Hongyu Cao, Dongjie Wang and 1 more

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Beyond Training Time, Test-Time Coordination is Essential for Cooperative MARL

Dongsu Lee, Sooraj Sathish, Joonkyung Kim, Woojun Kim and 1 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Joint Adaptive Neighborhood Constraint for Offline Multi-Agent Reinforcement Learning

Xiancheng Gao, Mingxiao Feng, Lin Liu, Yuanrui Duan and 3 more

Paris Poster Session 2, Wed, Dec 9, 5:00 PM–7:00 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Split-RL: Local Conflict Resolution in Reinforcement Learning

Benjamin Fuhrer, Chen Tessler, Gal Dalal

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Learning in Causal Markov Games

Aurghya Maiti, Elias Bareinboim

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Axiomatic Reinforcement Learning for Open Multi-Agent Systems from Shapley Axioms

Jianhong Wang, Yang Li, Samuel Kaski, Jonathan Lawry

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Focusing Influence Mechanism for Multi-Agent Reinforcement Learning

Yisak Park, Sunwoo Lee, Seungyul Han

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Finite-Sample Convergence in Networked Average Reward MARL: Decentralization Pitfalls and Entropy Remedies

Yizhou Zhang, Yashaswini Murthy, Laixi Shi, Adam Wierman

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Rethinking Credit Assignment in Cooperative MARL via Interventional Reward Response

Chamjin Joo, Myounghoon Ha, Sang Wan Lee

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Learning Large-Scale Competitive Team Behaviors with Mean-Field Interactions

Bhavini Jeloka, Yue Guan, Panagiotis Tsiotras

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

MASTARS: Multi-Agent Sequential Trajectory Augmentation with Return-Conditioned Subgoals

Jiwon Jeon, Myungsik Cho, Woojun Kim, Seongmin Kim and 3 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Boosting Multiagent Reinforcement Learning at High Replay Ratios with Ensemble Reset

Yaodong Yang, Hongyao Tang, Guangyong Chen, Pheng-Ann Heng

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Addressing Exogenous Variability in Cooperative Multi-Agent Reinforcement Learning

Seongmin Kim, Woohyeon Byeon, Jiwon Jeon, Seungyul Han and 1 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
83%Must read
?Must readVote to see the score

Generalized Intention Modeling in Multi-Agent Reinforcement Learning

A task-adaptive framework learns a performance-driven mixture of opponent intent representations to improve multi-agent reinforcement learning across diverse tasks.

Mateusz Odrowaz-Sypniewski, Jasmine Bayrooti, Ajay Shankar, Amanda Prorok

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 1/5
76%Highly rated
?Highly ratedVote to see the score

Independent Learning of Nash Equilibria in Partially Observable Markov Potential Games with Decoupled Dynamics

Independent learning achieves approximate Nash equilibria in partially observable Markov potential games with decoupled dynamics and near-polynomial complexity via finite history windows.

Philip Jordan, Maryam Kamgarpour

Paris Poster Session 2, Wed, Dec 9, 5:00 PM–7:00 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 2/5
medium 5/10
strict 3/5
72%Highly rated
?Highly ratedVote to see the score

Self-Supervised Goal-Reaching Results in Multi-Agent Cooperation and Exploration

Self-supervised goal-reaching enables multi-agent cooperation and exploration via sparse feedback, outperforming alternatives and discovering nontrivial coordination without explicit mechanisms.

Chirayu Nimonkar, Shlok Shah, Catherine Ji, Benjamin Eysenbach

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 8 of 20 reviewers recommend it
lenient 4/5
medium 3/10
strict 1/5
83%Must read
?Must readVote to see the score

Test-time Multi-agent Coordination by Decomposed Value Gradient Flow

SCOUT combines generative behavioral priors with decomposed value functions via test-time Stein gradient transport, achieving scalable offline multi-agent coordination with vanishing joint soft-value gaps.

Dongsu Lee, Haoran Xu, Amy Zhang

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 2/5
medium 9/10
strict 2/5
76%Highly rated
?Highly ratedVote to see the score

Auction-Based Online Policy Adaptation for Evolving Objectives

An auction-based framework coordinates selfish local policies via urgency bids to adapt multi-objective reinforcement learning to evolving objectives, outperforming monolithic PPO policies.

Guru Shabadi, Kaushik Mallik

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 4/5
medium 5/10
strict 1/5
83%Must read
?Must readVote to see the score

Mitigating Retaliatory Algorithmic Collusion in Repeated Games

CURB penalizes policy dependence on defection histories via total-variation reward shaping to eliminate collusive equilibria in repeated multi-agent games.

Karthik Sivachandran, Rohan Paleja

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 2/5
80%Must read
?Must readVote to see the score

Modelling Opinion Dynamics at Scale with Deep MARL

Deep MARL scales opinion dynamics to 1000 agents, finding high conformity in large networks reduces accuracy and promotes dishonesty, unlike small groups, revealing a mismatch with modern media.

Lukas Seier, Brandon Kaplowitz, Sebastian Towers, Richard M Bailey and 1 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 1/5
78%Highly rated
?Highly ratedVote to see the score

Robust Instruction Compliance in Cooperative Multi-Agent Reinforcement Learning

MAVIC corrects Bellman backups at instruction boundaries to maintain consistent value estimates under stochastic instruction switching, achieving high compliance with preserved cooperative task performance.

Wo Wei Lin, Ethan Rathbun, Enrico Marchesini, Xiang Zhi Tan

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 0/5
71%Highly rated
?Highly ratedVote to see the score

The Dynamics of Policy Gradient in Social Dilemmas with Partner Selection

Policy-gradient dynamics with partner selection are solved analytically, proving population variance is necessary for cooperation and deriving conditions for a stationary cooperative distribution.

Benedict Russell, Chin-wing Leung, Paolo Turrini

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 7 of 20 reviewers recommend it
lenient 3/5
medium 3/10
strict 1/5
76%Highly rated
?Highly ratedVote to see the score

PLATO: Pointer Learner for Agent and Task Openness

PLATO uses a pointer-network actor and GNN critic to handle open multi-agent reinforcement learning with unbounded agent and task spaces, achieving strong zero-shot generalization.

Alireza Saleh Abadi, Leen-Kiat Soh, Daniel A Redder, Adam Eck and 1 more

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 5/5
medium 5/10
strict 0/5
86%Must read
?Must readVote to see the score

Rethinking Ratio-Based Trust Regions for Policy Optimization in Multi-Agent Reinforcement Learning

MARS replaces ratio-based trust regions with a multiplicatively symmetric geometric barrier to cut variance and prevent probability collapse in multi-agent policy optimization. Across 47 tasks it matches or exceeds MAPPO and MASPO, with gains from barrier geometry rather than flexible boundaries.

Chulabhaya Wijesundara, Andrea Baisero, Zhongheng Li, Gregory D Castanon and 2 more

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 3/5
medium 8/10
strict 3/5
88%Must read
?Must readVote to see the score

Training Generalizable Collaborative Agents via Strategic Risk Aversion

Strategic risk aversion acts as an inductive bias for generalizable collaboration, yielding robust multi-agent policies with reduced free-riding and stronger equilibrium outcomes alongside unseen partners.

Chengrui Qu, Yizhou Zhang, Nicolas Lanzetti, Eric Mazumdar

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 1/5
76%Highly rated
?Highly ratedVote to see the score

Heterogeneous Agent Collaborative Reinforcement Learning

HACRL enables heterogeneous agents to share verified rollouts during collaborative on-policy training and execute independently at inference, with HACPO improving all agents by 3.6% over baselines at half the rollout cost.

Zhixia Zhang, Zixuan Huang, Gonxun Li, Huaiyang Wang and 8 more

Paris Poster Session 5, Fri, Dec 11, 11:30 AM–1:30 PM, Paris Poster Hall · Published 2026 · ▲ 110 on Hugging Face

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 3/5
medium 7/10
strict 0/5
89%Must read
?Must readVote to see the score

Dr. MAS: Stable Reinforcement Learning for Multi-Agent LLM Systems

Global normalization in multi-agent RL causes gradient instability; Dr. MAS normalizes per-agent advantages to stabilize training and boost multi-agent reasoning benchmarks.

Lang Feng, Longtao Zheng, Shuo He, Fuxiang Zhang and 1 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 30 on Hugging Face · Code ★ 168

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 10/10
strict 1/5
71%Highly rated
?Highly ratedVote to see the score

Differentiable Belief-based Opponent Shaping

D-BOS differentiates through k-step softmax-Bayes belief dynamics to shape multi-agent opponent beliefs, outperforming PPO and BBM in hidden-role games.

Aarav G Sane, Karthik Sivachandran, Rohan Paleja

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 7 of 20 reviewers recommend it
lenient 3/5
medium 4/10
strict 0/5
71%Highly rated
?Highly ratedVote to see the score

Convergence Guarantees for Federated SARSA with Local Training and Heterogeneous Agents

FedSARSA with linear approximation and local training achieves linear agent speed-up and converges despite heterogeneous transitions and rewards, with explicit sample and communication complexity bounds.

Paul Mangold, Eloïse Berthier, Eric Moulines

Paris Poster Session 1, Wed, Dec 9, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 7 of 20 reviewers recommend it
lenient 2/5
medium 4/10
strict 1/5
Show 20 more papers