Good Papers

UndoBench: Separating Task Competence from Recovery Capability in Tool-Using AI Agents

UndoBench separates tool-using agent competence from fault recovery via paired enterprise workflow trials, finding 83.54% nominal success but only 46.72% recovery success with phase-dependent vulnerabilities.

Dolly Sah, Tanmay Sah, Harshul Jain, Tanya Sah

Published Oct 4, 2026▲ 11 on Hugging FaceCode ★ 1arXiv ↗

91%
OverallMust read
?
OverallMust readVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel17/20reviewers recommend it
lenient 5/5
medium 9/10
strict 3/5
AI panel?Vote to see what the 20 AI reviewers said
Panel consensus
UndoBench reveals a striking 83.54% to 46.72% split between nominal competence and recovery, using wire-level oracles to expose phase-dependent failure modes, though its 36-workflow synthetic scope and single lost-acknowledgment focus leave open questions about real-world enterprise…

Abstract

Tool-using AI agents are increasingly deployed across enterprise software systems, yet widely used benchmarks primarily evaluate nominal task completion, conflating baseline planning competence with operational fault recovery. We introduce UndoBench, a benchmark spanning 36 base workflows and 36 fault scenarios across 8 enterprise domains, decoupling task competence from recovery capability via counterfactual paired trials under identical seeds alongside wire-level effect-history and environment-state oracles. On 12 held-out TEST workflows across two open-weight models, two frameworks, and three recovery paradigms (5,760 executions / 2,880 paired trials) in the frozen lost-acknowledgment study, nominal competence reached 83.54% while conditional recovery success rate (CRSR) fell to 46.72%, with naive retry producing duplicate external effects in 53.33% of trials. Extensions to commercial API models reproduced this competence-recovery separation. Evaluations across complementary execution boundaries show that recovery is phase-dependent: before mutation, methods perform similarly without duplicate effects among capable trials; during partial mutation, naive retry, per-call idempotency, and zero-privilege journaling collapse on the evaluated composite workflows; after commit but before acknowledgment, verification and server-side idempotency substantially improve safety. These findings demonstrate that evaluating nominal completion alone masks critical, phase-dependent recovery vulnerabilities in autonomous agents.