Benchmarking thirty metrics on ten simulated complex systems shows only causal metrics reliably validate high-level explanations when testing unmapped-variable faithfulness, leading to the Causal Abstraction Error metric converging with thirty interventions.
Automatic multi-agent systems consistently underperform single-agent chain-of-thought self-consistency despite up to 10x cost, revealing automated architectures suffer from bloat and misaligned complexity rather than true multi-agent benefits.