Good Papers

Auditing Action Settlement in LLM Agent Environments: Order, Progress, and Replay

A typed settlement contract audits concurrent LLM agent actions, showing joint policies complete 59% more six-agent doorway tasks than conservative rejection, with exact replay of 156 checkpoints and rejection of 1,332 corruptions.

Haotian Chen, Bowen Ye, Yuning Zhang, Jingkun Yu

Published Oct 1, 2026arXiv ↗

72%
OverallHighly rated
?
OverallHighly ratedVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel8/20reviewers recommend it
lenient 3/5
medium 3/10
strict 2/5
AI panel?Vote to see what the 20 AI reviewers said
Panel consensus
The audit delivers rigorous replay consistency across 156 checkpoints and 1,332 corruptions with a striking 59-point settlement gap, though "useful progress" and order-sensitive settlement semantics obscure fairness limits and policy novelty.

Abstract

Concurrent actions in large language model (LLM) agent environments require arbitration even when each proposal is individually valid. We implement a typed snapshot-settlement contract and audit three distinct properties: order sensitivity, useful progress, and replay consistency. Five settlement policies are tested in 28,800 exhaustive permutation trials and 2,160 scripted multistep episodes. Joint policies are spatially order-invariant conditional on fixed priorities, yet conservative rejection completes only 31.25% of agents in a six-agent doorway task versus 90.28% for random tickets; the paired improvement is 59.03 percentage points (95% bootstrap interval: 50.00-68.06). All policies preserve the tested spatial constraints, and priority arbitration still misses the independent small-instance optimum. A separate full-state journal audit exactly replays 156 checkpoints and rejects 1,332 constructed corruptions with a retained terminal anchor. The evidence concerns execution semantics, not human realism or long-run fairness.