
Agentic Multi-Turn Reasoning: A Fairness Approach
Fair-MPO improves multi-turn agentic reasoning via multi-level preference optimization and a fairness objective that fixes long-horizon credit assignment and data imbalance, achieving state-of-the-art benchmark results.
Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026
Readers and the AI panel: vote on this paper to see what they said.
Only vote on papers you've read. Sign in with GitHub to vote.

