Good Papers

From Anomalies to Failures: Constructing Causal Error Graphs for Agentic Trace Diagnosis

CEG-Agent introduces causal error graphs and a taxonomy separating anomalies, errors, and failures to diagnose agentic traces, achieving state-of-the-art results on the CEG-Bench benchmark.

Shu-Xun Yang, Yidong Wang, Zhuoer Feng, Bosi Wen, Jiayi Gui, Dayong Yang, Wenbo Yu, Haoke Zhang, Jie Tang, Cunxiang Wang

Published Sep 26, 2026arXiv ↗

78%
OverallHighly rated
?
OverallHighly ratedVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel11/20reviewers recommend it
lenient 5/5
medium 6/10
strict 0/5
AI panel?Vote to see what the 20 AI reviewers said
Panel consensus
CEG-Agent delivers a rigorous taxonomy and strong benchmark for agentic trace diagnosis, but its causal claims remain unproven without counterfactual interventions and clearer dataset scale.

Abstract

LLM-driven agents are increasingly deployed in complex applications, where long agentic traces make failures difficult to diagnose. Existing trace diagnosis methods often conflate anomalies, errors, and failures, making diagnostic targets ambiguous; they also lack structured modeling of how causally relevant errors propagate and amplify into final task failures, resulting in unreliable failure attribution. To address these problems, we propose CEG-Agent, a tool-augmented agentic framework for causal diagnosis of agentic traces. Specifically, CEG-Agent introduces an explicit taxonomy of anomalies, errors, and failures, and constructs Causal Error Graphs (CEGs), a unified typed representation that links execution events, diagnostic nodes, and failure outcomes through causal relations. To evaluate causal trace diagnosis, we further construct CEG-Bench, a fully agent-annotated benchmark with high-confidence, consensus-derived CEG annotations obtained through an Adversarial Agentic Adjudication Protocol (AAAP). We validate the resulting annotations against an expert-curated human gold set, which shows close agreement with the automatic annotations. Experiments on CEG-Bench demonstrate that CEG-Agent achieves state-of-the-art performance under both semantically relaxed and structurally exact evaluation criteria. Our code is publicly available.