Good Papers

Agentic RAG Evaluation: Budget Allocation Across Questions, Trajectories, and Reads

Agentic RAG evaluation budgets favor broader question coverage over repeated reads or trajectories, reducing standard error by up to 33% at fixed token cost.

Jingjie Ning, Xueqi Li, Yibo Kong

Published Oct 4, 2026▲ 14 on Hugging FacearXiv ↗

76%
OverallHighly rated
?
OverallHighly ratedVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel10/20reviewers recommend it
lenient 3/5
medium 5/10
strict 2/5
AI panel?Vote to see what the 20 AI reviewers said

Abstract

Evaluation budgets in agentic retrieval-augmented generation span questions, search trajectories, and repeated answers. We measure allocation precision, reading efficiency, and cost boundaries using a retrieval-feedback comparison on HotpotQA and MuSiQue. At 34.14--34.39M model tokens, broader question coverage lowers standard error by 33\% versus five reads and 12.6\% versus three trajectories. Archived nested and Q-only forecasts predict these allocations within 4.0\% and 3.5\%, respectively. Depth subsets establish no clear forecasting advantage beyond the two-trajectory audit. One-read variance penalties relative to the fitted optimum at the same token budget are 0--9.9\%, with substantial Pro uncertainty. Under recorded model fees, more questions beat more trajectories at search prices of \$0--1 per 1,000 requests; question-versus-read fee rankings remain unresolved. Temperature zero cuts answer disagreement from 14.3\% to 3.4\% while comparison precision stays similar. \par\medskip\noindent\textbf{Keywords:} Agentic RAG; Evaluation budget; Generalizability theory; Repeated sampling.