Good Papers

Showing LLM agents & planning Show all papers

80%Must read
?Must readVote to see the score

Beyond Semantic Similarity: Performance and Costs of Agentic Retrieval for Complex Tasks

Agentic retrieval combining LLM reasoning with dense retrieval improves nDCG@10 by 8.7 points over standard retrieval but requires 107 seconds and 764K input tokens per query.

Reza Esfandiarpoor, Radek Osmulski, Yauhen Babakhin, Gabriel de Souza P. Moreira and 7 more

Published Oct 5, 2026 · ▲ 7 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 1/5
78%Highly rated
?Highly ratedVote to see the score

EVISKILL: Grounding Skill Evolution in Replayable Evidence

EVISKILL grounds LLM skill evolution in replayable evidence cards linking edits to supporting contexts, using targeted replay for verification and global validation for incorporation.

Yan Zhou, Yili Wang, Yiwei Dai, Qinggang Zhang and 1 more

Published Oct 4, 2026 · ▲ 11 on Hugging Face · Code ★ 8

100% Readers1 of 1 upvoted
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 4/5
medium 4/10
strict 1/5
83%Must read
?Must readVote to see the score

ASCENT: Online Test-Time Training of Long-Horizon Agents via Self-Distillation of Verified Experience

ASCENT online test-time trains agents by self-distilling verified deployment trajectories into LoRA weights via a frozen hindsight model, improving long-horizon success and efficiency without external teachers or memory retrieval.

Haodong Lu, Dong Gong

Published Oct 4, 2026 · ▲ 17 on Hugging Face

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 0/5
71%Highly rated

Code2Games: Enabling Coding Agents for Gaming World Generation

Code2Games coordinates scene analysis and gameplay planning via shared representations to generate consistent gaming worlds and adapt them to Unreal Engine 5. The framework improves visual quality, interactive fidelity, and playable-game quality over direct coding-agent generation on the GameCode4D

Wei Wu, Ziyang Xu, Zeyu Zhang, Yang Zhao and 1 more

Published Oct 4, 2026 · ▲ 6 on Hugging Face · Code ★ 4

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 7 of 20 reviewers recommend it
lenient 4/5
medium 3/10
strict 0/5
80%Must read
?Must readVote to see the score

SearchJev: A Fast and Calibrated System-1 Model for Search Agents

SearchJev is a fast calibrated System-1 model that scores search decisions directly without autoregressive generation, improving decision quality, speed, and calibration over same-size language models.

Congfeng Cao, Lipeng Zuo, Konstantinos Papakostas, Qiwei Xu and 5 more

Published Oct 4, 2026 · ▲ 14 on Hugging Face · Code ★ 4

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 1/5
86%Must read
?Must readVote to see the score

Source Preference in the Wild: How LLM Agents Favor Items by Source, and How to Reduce It

LLM agents show strong source preferences across search domains that can override item quality, though supplying missing information or countering preconceptions reduces this bias.

Jonghyun Song, Haewon Park, Jeonghoon Shim, Woojung Song and 1 more

Published Oct 2, 2026 · 0 citations · ▲ 37 on Hugging Face

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 2/5
83%Must read

Scaling Trajectories for Complex Tasks through Recursive Self-Rewrite

Recursive Self-Rewrite uses diverse harnesses and recursive revision to rewrite successful terminal trajectories for supervised fine-tuning, boosting pass@3 by up to 7.6x on hard benchmarks.

Zongxia Li, Yucheng Shi, Zhongzhi Li, Junyao Yang and 6 more

Published Oct 2, 2026 · 0 citations · ▲ 91 on Hugging Face

– ReadersNo votes yet. 1 from authors or colleagues not counted
13/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 2/5
83%Must read
?Must readVote to see the score

Harness-Aware Distillation for Small Language Model Agents

Harness-Aware Distillation focuses agent distillation on capabilities beyond the fixed harness via action preferences and validity checks, improving long-horizon agent performance without task rewards.

Moonseok Choi, Taehong Moon, Giung Nam, Juho Lee

Published Oct 2, 2026 · 0 citations · ▲ 4 on Hugging Face · Code

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 0/5
86%Must read
?Must readVote to see the score

Training Numerical Intelligence via Auto-Diagnosis and Skill Discovery

ADSD links numerical diagnosis to reusable solver self-improvement, reducing mean solver error by nearly 71x across four challenging numerical domains.

Peter Chen, Wotao Yin

Published Oct 2, 2026 · ▲ 8 on Hugging Face

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 1/5
88%Must read
?Must readVote to see the score

VeriHarness: Scaling Agentic Verification for Long-Horizon Tasks

VeriHarness turns fixed base LLMs into agentic verifiers with workspaces and evidence tools, achieving top selection scores and 6.2, 6.4 point gains over single rollouts on long-horizon tasks.

Caiqi Zhang, Rujun Han, Zifeng Wang, Zoey CuiZhu and 3 more

Published Oct 1, 2026 · 0 citations · ▲ 49 on Hugging Face · Code ★ 42

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 3/5
80%Must read
?Must readVote to see the score

Selection-Based Structured Reasoning: Toward Efficient Multimodal Search Agents

Selection-based Structured Reasoning replaces open-ended reasoning with selection among reusable candidates, cutting per-turn latency over 90% while matching leading small-model search agents' success rates.

Feiyu Gavin Zhu, Xiaoyu Zhu, Jiqi Yang, Rui Yang and 8 more

Published Oct 1, 2026 · 0 citations · ▲ 3 on Hugging Face · Code

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 1/5
80%Must read
?Must readVote to see the score

ActiveSaddler: Automated Curriculum Learning for Agent Harness Optimization

ActiveSaddler automates curriculum learning for agent harness optimization via non-stationary bandits that adapt training scenarios to evolving failure patterns, boosting Pass@1 by 4.4, 7.5 points.

Sungho Park, Wonjoong Kim, Jue Zhang, Wook-Shin Han and 7 more

Published Oct 1, 2026 · ▲ 82 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 0/5
86%Must read
?Must readVote to see the score

From Knowledge Access to Source Learning: Developing Source-Specific Competence

SourceLearn develops reusable source-specific competence via persistent source models and dual learning mechanisms, outperforming retrieval and memory baselines by up to 22.6 points.

Lucheng Fu, Kejing Xia, Yiyang Wang, Yiqiao Jin and 9 more

Published Oct 1, 2026 · 0 citations · ▲ 5 on Hugging Face · Code

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 2/5
88%Must read
?Must readVote to see the score

Code Owns the Simulation, Jev Owns the Evaluation

Judgment models excel at evaluation but fail at simulation, yet pairing them with code simulation yields expert control.

Yaodong Yang, Hongyao Tang, Yi Ma, Xingyu Fan and 3 more

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 2/5
80%Must read
?Must readVote to see the score

Thinking Outside the Box: Can Language Models Rely on External Guidance Selectively?

Frontier models follow unreliable external guidance; training improves selective reliance, identifying it as a key agent reliability dimension.

Minghan Wang, Boyuan Wang, Jinhang Zuo, Yuxin Tao and 1 more

Published Sep 30, 2026 · 0 citations · ▲ 67 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 0/5
88%Must read
?Must readVote to see the score

Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents

Mid-Harness verifies candidate terminal actions at the model-harness boundary, raising TerminalBench-Lite Pass@1 from 50.00% to 68.03% and improving success at lower token cost than trajectory scaling alone.

Minki Kang, Ryo Hachiuma, Shaokun Zhang, Subhashree Radhakrishnan and 7 more

Published Sep 30, 2026 · 0 citations · ▲ 116 on Hugging Face

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 1/5
91%Must read

False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents

Self-evolving search agents suffer co-cheating where proposers and solvers mutually reinforce errors; CrossFit partitions sources to cross-fit agreement and cuts false agreement by over half, boosting downstream search by 8+ points.

Meijia Chen, Hao Li, Zheng Lu, Hongshan Lin and 11 more

Published Sep 30, 2026 · 0 citations · ▲ 670 on Hugging Face

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 17 of 20 reviewers recommend it
lenient 4/5
medium 10/10
strict 3/5
88%Must read
?Must readVote to see the score

GraphForge: Training Working Agents with Graph-Anchored Workspace Synthesis

GraphForge synthesizes workspace tasks and verifiers over real file evidence graphs to train working agents, and fine-tuning Qwen3.6-27B improves GDPVal, Workspace-Bench-Lite, and SpreadsheetBench II results.

Qisheng Su, Hanchen Wang, 朱冠儒, Huicheng Jiang and 8 more

Published Sep 30, 2026 · 0 citations · ▲ 146 on Hugging Face

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 1/5
80%Must read
?Must readVote to see the score

EVOKE: Eliciting World Knowledge in Agents for Transferable Decision-Making

EVOKE improves LLM agent transfer by ranking actions under diverse goals at fixed states to elicit pretrained world knowledge for robust decision-making.

Yuhan Guo, Jinming Liu, Liang Xu, Ziqiang Li and 7 more

Published Sep 29, 2026 · 0 citations · ▲ 78 on Hugging Face · Code ★ 7

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 0/5
76%Highly rated
?Highly ratedVote to see the score

Learning Meta-Skills for Agent Harness Design in Test-Time AI4AI

A Builder learns reusable meta-skills from target feedback to construct execution harnesses that boost target performance on unseen tasks. Meta-skills improve macro-average scores by 8.95 points over no-skill construction and 12.02 over direct delivery.

Cheng Qian, Kunlun Zhu, Beibin Li, Zhenhailong Wang and 1 more

Published Sep 29, 2026 · 0 citations · ▲ 83 on Hugging Face · Code ★ 9

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 0/5
Show 20 more papers