Good Papers

HyperMARL: Adaptive Hypernetworks for Multi-Agent RL

HyperMARL uses agent-conditioned hypernetworks to generate agent-specific parameters that decouple gradients, reducing variance and preserving behavioral diversity across multi-agent benchmarks without added complexity.

Kale-ab Tessera, Arrasy Rahman, Amos Storkey, Stefano Albrecht

Published 2025Paper ↗

88%
OverallMust read
?
OverallMust readVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel15/20reviewers recommend it
lenient 4/5
medium 10/10
strict 1/5
AI panel?Vote to see what the 20 AI reviewers said
Panel consensus
HyperMARL replaces preset diversity mechanisms with hypernetwork-generated agent-specific parameters that decouple agent-conditioned gradients to suppress interference, though its lasting value depends on whether the variance reduction holds at scale without wall-clock overhead.

Abstract

Adaptive cooperation in multi-agent reinforcement learning (MARL) requires policies to express homogeneous, specialised, or mixed behaviours, yet achieving this adaptivity remains a critical challenge. While parameter sharing (PS) is standard for efficient learning, it notoriously suppresses the behavioural diversity required for specialisation. This failure is largely due to cross-agent gradient interference, a problem we find is surprisingly exacerbated by the common practice of coupling agent IDs with observations. Existing remedies typically add complexity through altered objectives, manual preset diversity levels, or sequential updates -- raising a fundamental question: can shared policies adapt without these intricacies? We propose a solution built on a key insight: an agent-conditioned hypernetwork can generate agent-specific parameters and decouple observation- and agent-conditioned gradients, directly countering the interference from coupling agent IDs with observations. Our resulting method, HyperMARL, avoids the complexities of prior work and empirically reduces policy gradient variance. Across diverse MARL benchmarks (22 scenarios, up to 30 agents), HyperMARL achieves performance competitive with six key baselines while preserving behavioural diversity comparable to non-parameter sharing methods, establishing it as a versatile and principled approach for adaptive MARL. The code is publicly available at https://github.com/KaleabTessera/HyperMARL.