Good Papers

Heterogeneous Agent Collaborative Reinforcement Learning

HACRL enables heterogeneous agents to share verified rollouts during collaborative on-policy training and execute independently at inference, with HACPO improving all agents by 3.6% over baselines at half the rollout cost.

Zhixia Zhang, Zixuan Huang, Gonxun Li, Huaiyang Wang, Chengyi Yuan, Xin Xia, deqing wang, Fuzhen Zhuang, Shuai Ma, Ning Ding, Yaodong Yang, Yikun Ban

Published 2026Paris Poster Session 5 · Fri, Dec 11, 11:30 AM–1:30 PM local time · Paris Poster Hall▲ 110 on Hugging FacearXiv ↗OpenReview ↗

76%
OverallHighly rated
?
OverallHighly ratedVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel10/20reviewers recommend it
lenient 3/5
medium 7/10
strict 0/5
AI panel?Vote to see what the 20 AI reviewers said
Panel consensus
HACRL delivers a rigorous, bidirectional rollout-sharing framework with unbiased advantage guarantees and 3.6% gains at half rollout cost, though its "heterogeneous" scope is limited to model sizes and practical open-source validation remains absent.

Abstract

We introduce Heterogeneous Agent Collaborative Reinforcement Learning (HACRL), a new Reinforcement Learning from Verifiable Reward (RLVR) problem that addresses the inefficiencies of isolated multi-agent on-policy optimization. HACRL enables collaborative optimization with independent execution: heterogeneous agents share verified rollouts during training to mutually improve, while operating independently at inference time. Unlike LLM-based multi-agent reinforcement learning (MARL), HACRL does not require coordinated deployment, and unlike on-/off-policy distillation, it enables bidirectional mutual learning among heterogeneous agents rather than one-directional homogeneous teacher-to-student transfer. Building on this problem, we propose HACPO, a collaborative RL algorithm that enables principled rollout sharing to maximize sample utilization and cross-agent knowledge transfer. To mitigate capability discrepancies and policy distribution shifts, HACPO introduces four tailored mechanisms with theoretical guarantees on unbiased advantage estimation. Extensive experiments across diverse heterogeneous model combinations and reasoning benchmarks show that HACPO consistently improves all participating agents, outperforming GSPO with double rollouts by an average of 3.6% while using only half the rollout cost.