Good Papers

Where Root Cause Analysis Fails: A Retrieval-Reranking Decomposition

Root cause analysis benchmarks conflate retrieval and reranking failures, revealing graph methods rarely beat statistical baselines; a two-stage retriever-LLM reranker matches or exceeds all baselines without causal graphs or labels.

Hada M Muhammad, Luan Pham, Laure Barrière, Sachin Shetty, Leonardo Pulga, Flora Salim

Published 2026Sydney Poster Session 2 · Tue, Dec 8, 5:00 PM–8:00 PM local time · Hall 1-4arXiv ↗OpenReview ↗

92%
OverallMust read
?
OverallMust readVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel19/20reviewers recommend it
lenient 5/5
medium 10/10
strict 4/5
AI panel?Vote to see what the 20 AI reviewers said

Abstract

Identifying the root cause of an anomaly among hundreds of sensors is critical for preventing safety incidents and costly downtime in complex monitored systems. Existing studies evaluate root cause analysis (RCA) methods using top@k accuracy. We show that this metric has a fundamental blind spot: it conflates two failure modes, retrieval failure, where the true cause is never considered, and reranking failure, where it is considered but ranked too low. In this work, we introduce a retrieval-reranking decomposition and audit four well-known benchmarks to expose this blind spot. Our experiments show that, on benchmarks with complex faults, statistical baselines mis-rank the true cause 79-100% of the time, and graph-based methods never clearly beat the best statistical baseline, whether their causal graphs are learned on short fault windows, on retrieved candidate pools guaranteed to contain the cause, or on multi-day normal-operation data. Meanwhile, on simple benchmarks where faults manifest significantly at their origin, retrieval is nearly solved (98-100%). Guided by the decomposition, we build a two-stage pipeline combining a multi-signal retriever with an LLM reranker that, as one fixed configuration, matches or exceeds the best baseline's top@1 accuracy on all six benchmark suites (by up to +12 points), with no causal graph or labeled data required. When all methods rank the same retrieved candidates with the true cause guaranteed present, adding a short system-description document lets the reranker lead the best baseline by +7 to +18 points on every benchmark. Code is available at https://github.com/cruiseresearchgroup/DecompRCA.