From Evidence to Action: How Tool-Using Agents Fail
Tool-using agents often fail by acting before establishing required evidence or leaving multi-action workflow prerequisites unresolved, despite accurate static action assessment. SafeActBench reveals failures stem from how agents use established evidence during execution, not just missing informatio
Published Oct 6, 2026 · ▲ 20 on Hugging Face · Code ★ 3
Readers and the AI panel: vote on this paper to see what they said.
Only vote on papers you've read. Sign in with GitHub to vote.







