Good Papers

MTAC-IFBench: Benchmarking Instruction-Following in Multi-Turn Agentic Coding

MTAC-IFBench benchmarks multi-turn instruction-following in agentic coding via progressive constraints, revealing rapid performance degradation in current code agents as sessions lengthen.

Bosi Wen, Cunxiang Wang, Jiayi Gui, Haoke Zhang, Yilin Niu, Pei Ke, Dayong Yang, Hongning Wang, Minlie Huang

Published Sep 14, 2026arXiv ↗

80%
OverallMust read
?
OverallMust readVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel12/20reviewers recommend it
lenient 5/5
medium 7/10
strict 0/5
AI panel?Vote to see what the 20 AI reviewers said
Panel consensus
MTAC-IFBench delivers a rigorous, densely constrained multi-turn benchmark that exposes how instruction-following collapses across agentic coding turns, though it leaves unclear whether failures stem from reasoning limits or context decay and whether category breakdowns are turn-ordered.

Abstract

Recently, the rapid development of large language models (LLMs) has reshaped software engineering by enabling autonomous code agents that plan, execute, and utilize external tools iteratively to tackle complex tasks. Beyond achieving functional correctness, these agents must faithfully follow process instructions and constraints throughout the development lifecycle. However, existing benchmarks typically focus on final functional correctness or confine instruction-following evaluation to single-turn, general chat or simple code generation scenarios, leaving instruction-following in multi-turn agentic coding underexplored. To bridge this gap, we propose MTAC-IFBench, a comprehensive benchmark for this critical capability. It features multi-turn progressive software development instructions with diverse constraints spanning 6 primary and 18 secondary categories. With an average of 7.04 turns and 91.33 constraints per instance, it poses a rigorous challenge to current LLMs. To make the evaluation reliable, we construct a checklist for each constraint and functional requirement, and integrate verification scripts and judge agents to verify each checklist item. MTAC-IFBench identifies significant deficiencies in existing code agents in multi-turn instruction-following, with their performance degrading rapidly as the interaction session grows longer.