Good Papers

Showing LLM agents & planning Show all papers

80%Must read
?Must readVote to see the score

Beyond Semantic Similarity: Performance and Costs of Agentic Retrieval for Complex Tasks

Agentic retrieval combining LLM reasoning with dense retrieval improves nDCG@10 by 8.7 points over standard retrieval but requires 107 seconds and 764K input tokens per query.

Reza Esfandiarpoor, Radek Osmulski, Yauhen Babakhin, Gabriel de Souza P. Moreira and 7 more

Published Oct 5, 2026 · ▲ 7 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 1/5
78%Highly rated
?Highly ratedVote to see the score

EVISKILL: Grounding Skill Evolution in Replayable Evidence

EVISKILL grounds LLM skill evolution in replayable evidence cards linking edits to supporting contexts, using targeted replay for verification and global validation for incorporation.

Yan Zhou, Yili Wang, Yiwei Dai, Qinggang Zhang and 1 more

Published Oct 4, 2026 · ▲ 11 on Hugging Face · Code ★ 8

100% Readers1 of 1 upvoted
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 4/5
medium 4/10
strict 1/5
83%Must read
?Must readVote to see the score

ASCENT: Online Test-Time Training of Long-Horizon Agents via Self-Distillation of Verified Experience

ASCENT online test-time trains agents by self-distilling verified deployment trajectories into LoRA weights via a frozen hindsight model, improving long-horizon success and efficiency without external teachers or memory retrieval.

Haodong Lu, Dong Gong

Published Oct 4, 2026 · ▲ 17 on Hugging Face

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 0/5
71%Highly rated

Code2Games: Enabling Coding Agents for Gaming World Generation

Code2Games coordinates scene analysis and gameplay planning via shared representations to generate consistent gaming worlds and adapt them to Unreal Engine 5. The framework improves visual quality, interactive fidelity, and playable-game quality over direct coding-agent generation on the GameCode4D

Wei Wu, Ziyang Xu, Zeyu Zhang, Yang Zhao and 1 more

Published Oct 4, 2026 · ▲ 6 on Hugging Face · Code ★ 4

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 7 of 20 reviewers recommend it
lenient 4/5
medium 3/10
strict 0/5
80%Must read
?Must readVote to see the score

SearchJev: A Fast and Calibrated System-1 Model for Search Agents

SearchJev is a fast calibrated System-1 model that scores search decisions directly without autoregressive generation, improving decision quality, speed, and calibration over same-size language models.

Congfeng Cao, Lipeng Zuo, Konstantinos Papakostas, Qiwei Xu and 5 more

Published Oct 4, 2026 · ▲ 20 on Hugging Face · Code ★ 8

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 1/5
86%Must read
?Must readVote to see the score

Source Preference in the Wild: How LLM Agents Favor Items by Source, and How to Reduce It

LLM agents show strong source preferences across search domains that can override item quality, though supplying missing information or countering preconceptions reduces this bias.

Jonghyun Song, Haewon Park, Jeonghoon Shim, Woojung Song and 1 more

Published Oct 2, 2026 · 0 citations · ▲ 43 on Hugging Face

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read

Scaling Trajectories for Complex Tasks through Recursive Self-Rewrite

Recursive Self-Rewrite uses diverse harnesses and recursive revision to rewrite successful terminal trajectories for supervised fine-tuning, boosting pass@3 by up to 7.6x on hard benchmarks.

Zongxia Li, Yucheng Shi, Zhongzhi Li, Junyao Yang and 6 more

Published Oct 2, 2026 · 0 citations · ▲ 100 on Hugging Face

– ReadersNo votes yet. 1 from authors or colleagues not counted
13/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Harness-Aware Distillation for Small Language Model Agents

Harness-Aware Distillation focuses agent distillation on capabilities beyond the fixed harness via action preferences and validity checks, improving long-horizon agent performance without task rewards.

Moonseok Choi, Taehong Moon, Giung Nam, Juho Lee

Published Oct 2, 2026 · 0 citations · ▲ 4 on Hugging Face · Code

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

Training Numerical Intelligence via Auto-Diagnosis and Skill Discovery

ADSD links numerical diagnosis to reusable solver self-improvement, reducing mean solver error by nearly 71x across four challenging numerical domains.

Peter Chen, Wotao Yin

Published Oct 2, 2026 · ▲ 8 on Hugging Face

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 1/5
88%Must read
?Must readVote to see the score

VeriHarness: Scaling Agentic Verification for Long-Horizon Tasks

VeriHarness turns fixed base LLMs into agentic verifiers with workspaces and evidence tools, achieving top selection scores and 6.2, 6.4 point gains over single rollouts on long-horizon tasks.

Caiqi Zhang, Rujun Han, Zifeng Wang, Zoey CuiZhu and 3 more

Published Oct 1, 2026 · 0 citations · ▲ 55 on Hugging Face · Code ★ 54

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Selection-Based Structured Reasoning: Toward Efficient Multimodal Search Agents

Selection-based Structured Reasoning replaces open-ended reasoning with selection among reusable candidates, cutting per-turn latency over 90% while matching leading small-model search agents' success rates.

Feiyu Gavin Zhu, Xiaoyu Zhu, Jiqi Yang, Rui Yang and 8 more

Published Oct 1, 2026 · 0 citations · ▲ 3 on Hugging Face · Code

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

ActiveSaddler: Automated Curriculum Learning for Agent Harness Optimization

ActiveSaddler automates curriculum learning for agent harness optimization via non-stationary bandits that adapt training scenarios to evolving failure patterns, boosting Pass@1 by 4.4, 7.5 points.

Sungho Park, Wonjoong Kim, Jue Zhang, Wook-Shin Han and 7 more

Published Oct 1, 2026 · ▲ 82 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 0/5
86%Must read
?Must readVote to see the score

From Knowledge Access to Source Learning: Developing Source-Specific Competence

SourceLearn develops reusable source-specific competence via persistent source models and dual learning mechanisms, outperforming retrieval and memory baselines by up to 22.6 points.

Lucheng Fu, Kejing Xia, Yiyang Wang, Yiqiao Jin and 9 more

Published Oct 1, 2026 · 0 citations · ▲ 9 on Hugging Face · Code

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

Code Owns the Simulation, Jev Owns the Evaluation

Judgment models excel at evaluation but fail at simulation, yet pairing them with code simulation yields expert control.

Yaodong Yang, Hongyao Tang, Yi Ma, Xingyu Fan and 3 more

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Thinking Outside the Box: Can Language Models Rely on External Guidance Selectively?

Frontier models follow unreliable external guidance; training improves selective reliance, identifying it as a key agent reliability dimension.

Minghan Wang, Boyuan Wang, Jinhang Zuo, Yuxin Tao and 1 more

Published Sep 30, 2026 · 0 citations · ▲ 67 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents

Mid-Harness verifies candidate terminal actions at the model-harness boundary, raising TerminalBench-Lite Pass@1 from 50.00% to 68.03% and improving success at lower token cost than trajectory scaling alone.

Minki Kang, Ryo Hachiuma, Shaokun Zhang, Subhashree Radhakrishnan and 7 more

Published Sep 30, 2026 · 0 citations · ▲ 116 on Hugging Face

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

91%Must read

False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents

Self-evolving search agents suffer co-cheating where proposers and solvers mutually reinforce errors; CrossFit partitions sources to cross-fit agreement and cuts false agreement by over half, boosting downstream search by 8+ points.

Meijia Chen, Hao Li, Zheng Lu, Hongshan Lin and 11 more

Published Sep 30, 2026 · 0 citations · ▲ 672 on Hugging Face

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

GraphForge: Training Working Agents with Graph-Anchored Workspace Synthesis

GraphForge synthesizes workspace tasks and verifiers over real file evidence graphs to train working agents, and fine-tuning Qwen3.6-27B improves GDPVal, Workspace-Bench-Lite, and SpreadsheetBench II results.

Qisheng Su, Hanchen Wang, 朱冠儒, Huicheng Jiang and 8 more

Published Sep 30, 2026 · 0 citations · ▲ 146 on Hugging Face

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

EVOKE: Eliciting World Knowledge in Agents for Transferable Decision-Making

EVOKE improves LLM agent transfer by ranking actions under diverse goals at fixed states to elicit pretrained world knowledge for robust decision-making.

Yuhan Guo, Jinming Liu, Liang Xu, Ziqiang Li and 7 more

Published Sep 29, 2026 · 0 citations · ▲ 78 on Hugging Face · Code ★ 7

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score

Learning Meta-Skills for Agent Harness Design in Test-Time AI4AI

A Builder learns reusable meta-skills from target feedback to construct execution harnesses that boost target performance on unseen tasks. Meta-skills improve macro-average scores by 8.95 points over no-skill construction and 12.02 over direct delivery.

Cheng Qian, Kunlun Zhu, Beibin Li, Zhenhailong Wang and 1 more

Published Sep 29, 2026 · 0 citations · ▲ 83 on Hugging Face · Code ★ 9

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

72%Highly rated
?Highly ratedVote to see the score

AREX-2: Advancing Self-Improving Agents through Long-Horizon Reflective Tasks

AREX-2 synthesizes long-horizon reflective trajectories to train a Qwen3.8-27B agent that self-improves at test time, achieving strong results on MLE-bench, Frontier-CS, and deep research benchmarks while scaling with iteration budget.

Hongjin Qian, Chaofan Li, Kun Luo, Wenqing Wei and 10 more

Published Sep 29, 2026 · 0 citations · ▲ 139 on Hugging Face · Code ★ 31

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

Foundations of Proactive Agents: Principles, Technical Layers, and Proactivity-Gym

Proactive LLM agents need joint optimization of task capability, temporal compute allocation, and user trust, with Proactivity-Gym exposing evaluation gaps and human preference for unobtrusive assistance.

Jio Oh, Seunghyun Do, Youngjun Lee, Steven Euijong Whang and 1 more

Published Sep 29, 2026 · 0 citations · ▲ 28 on Hugging Face

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

72%Highly rated
?Highly ratedVote to see the score

Retrieval-Augmented Skill Optimization via Cross-Harness Adaptation

RASO retrieves and adapts external skills via cross-harness adaptation to initialize and iteratively update agent skills, outperforming non-retrieval baselines across benchmarks.

Jaewon Chu, Ji Soo Lee, Jihwan Park, Dohwan Ko and 7 more

Published Sep 29, 2026 · ▲ 63 on Hugging Face

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 8 of 20 reviewers recommend it
lenient 4/5
medium 4/10
strict 0/5
80%Must read
?Must readVote to see the score

AIM: Agentic Idea Management for Automated Research

AIM autonomously manages research ideas via Bayesian-inspired selection and auditing to outperform baselines by up to 4.9 points and speed up search up to 3.1x.

Hyeong Kyu Choi, Bhavana Dalvi Mishra, Jiefeng Chen, Mihir Parmar and 6 more

Published Sep 29, 2026 · 0 citations · ▲ 51 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

X-Tree: Tokenizing Reusable Experience for Efficient Agent Generalization

X-Tree learns reusable hierarchical skills from agent trajectories and improves success rates up to 5.8% across web and science benchmarks.

Sitao Cheng, Xunjian Yin, Zhiyuan Sun, Yuxuan Li and 3 more

Published Sep 26, 2026 · 0 citations · ▲ 76 on Hugging Face · Code ★ 2

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

PluginRSI: Recursive Improvement of Agent Harnesses with Reusable Plugins

PluginRSI evolves agent harnesses via reusable plugins, improving over existing methods and accelerating optimization on unseen tasks.

Yaorui Shi, Yuchun Miao, Yuxin Chen, Jiayuan Zhang and 4 more

Published Sep 26, 2026 · 0 citations · ▲ 7 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

From Anomalies to Failures: Constructing Causal Error Graphs for Agentic Trace Diagnosis

CEG-Agent introduces causal error graphs and a taxonomy separating anomalies, errors, and failures to diagnose agentic traces, achieving state-of-the-art results on the CEG-Bench benchmark.

Shu-Xun Yang, Yidong Wang, Zhuoer Feng, Bosi Wen and 6 more

Published Sep 26, 2026 · 0 citations

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

89%Must read
?Must readVote to see the score

RRSI: Regularized Recursive Self-Improvement of Agent Harnesses

RRSI regularizes recursive agent harness self-improvement via annealed edit budgets, trajectory exploration, and critical selection to boost out-of-distribution performance and reduce token use. It improves up to 14.1 points in-distribution and 4.7 points out-of-distribution while cutting policy tok

Peng Xia, Rujun Han, Zifeng Wang, Yanfei Chen and 10 more

Published Sep 21, 2026 · 0 citations · ▲ 222 on Hugging Face · Code ★ 1,293

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

72%Highly rated
?Highly ratedVote to see the score

NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness

NeoHorse-1 closes an evaluation-selection-update loop via harness-mediated routing, structured post-training, and capability-guided data allocation, raising 4B and 9B macro-averages by ~6 and ~3.4 points.

NeoHorse Team, Guoliang Cao, Guohao Dai, Tianyu Guo and 33 more

Published Sep 8, 2026 · 0 citations · ▲ 327 on Hugging Face · Code ★ 1,673

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

72%Highly rated
?Highly ratedVote to see the score

Apodex 1.1: Scaling Agentic Intelligence for Complex Work

Apodex 1.1 scales agentic intelligence via environment and coordination scaling to achieve leading complex-work performance with smaller models.

B. An, B. An, B. Wang, B. L. Wang and 36 more

Published Aug 24, 2026 · 0 citations · ▲ 212 on Hugging Face · Code ★ 5,146

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

91%Must read

StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling

StateM improves long-horizon agent accuracy via durable-state harness scaling without model changes, reaching 95.3% on Terminal-Bench 2.1 and cutting API costs to about $15 versus $574.68.

Ziheng Qin, Yaxin Lu, Zhangyang Atlas Wang, Kai Wang

Published Aug 15, 2026 · 0 citations · ▲ 452 on Hugging Face · Code ★ 1,312

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

SkillOpt: Executive Strategy for Self-Evolving Agent Skills

SkillOpt treats agent skills as external state optimized via bounded text edits validated on held-out scores, improving accuracy up to 24.8 points with stable transfer.

Yifan Yang, Ziyang Gong, Weiquan Huang, Qihao Yang and 11 more

Published May 22, 2026 · 2 citations · ▲ 267 on Hugging Face · Code ★ 18,090

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

RewardHarness: Self-Evolving Agentic Post-Training

RewardHarness evolves agentic evaluation tools from minimal preference data to judge image edits, surpassing GPT-5 accuracy with 0.05% training annotations.

Yuxuan Zhang, Penghui Du, Bo Li, Cong Wei and 10 more

Published May 9, 2026 · 0 citations · ▲ 244 on Hugging Face · Code ★ 69

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

72%Highly rated
?Highly ratedVote to see the score

SkillClaw: Let Skills Evolve Collectively with Agentic Evolver

SkillClaw enables collective skill evolution in multi-user LLM agent ecosystems by aggregating cross-user interaction trajectories and autonomously updating shared reusable skills, significantly improving real-world agent performance.

Ziyu Ma, Shidong Yang, Yuxiang Ji, Xucong Wang and 4 more

Published Apr 9, 2026 · 0 citations · ▲ 225 on Hugging Face · Code ★ 2,669

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

72%Highly rated
?Highly ratedVote to see the score

MiroThinker-1.7 & H1: Towards Heavy-Duty Research Agents via Verification

MiroThinker-H1 integrates local and global verification into reasoning for reliable multi-step problem solving and achieves state-of-the-art deep research performance.

MiroMind Team, S. Kamala Bai, L. Bing, L. Lei and 36 more

Published Mar 16, 2026 · 0 citations · ▲ 187 on Hugging Face

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

72%Highly rated
?Highly ratedVote to see the score

A Survey of Agentic Reasoning for Large Language Models: Towards Recursively Self-Improving and Collective Agents

This survey organizes LLM agentic reasoning into foundational, self-evolving, and collective layers, distinguishing in-context and post-training methods across applications while outlining open challenges.

Tianxin Wei, Ting-Wei Li, Zhining Liu, Xuying Ning and 25 more

Published Jan 18, 2026 · 1 citation · ▲ 208 on Hugging Face · Code ★ 1,398

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling

MiroThinker introduces interaction scaling to train open-source research agents for deeper tool use, achieving up to 81.9% on GAIA and rivaling commercial models.

MiroMind Team, Bai, Song, Lidong Bing, Chen, Carson and 36 more

Published Nov 14, 2025 · 0 citations · ▲ 197 on Hugging Face · Code ★ 8,419

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

70%Highly rated
?Highly ratedVote to see the score

ChatSOP: An SOP-Guided MCTS Planning Framework for Controllable LLM Dialogue Agents

ChatSOP introduces an SOP-guided MCTS framework that improves LLM dialogue agents' controllability and task success by following structured operational procedures.

Zhigen Li, Jianxiang Peng, Yanmeng Wang, Yong Cao and 12 more

Published 2025 · 3 citations

– ReadersNo votes yet
5/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

93%Must read
?Must readVote to see the score

Neglected Free Lunch from Post-training: Progress Advantage for LLM Agents

RL post-training yields progress advantage, a log-ratio that recovers optimal step-level advantage without dedicated reward models, outperforming trained alternatives across agent benchmarks.

Changdae Oh, Wendi Li, Seongheon Park, Samuel (Min-Hsuan) Yeh and 2 more

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026 · ▲ 13 on Hugging Face · Code ★ 12

100% Readers1 of 1 upvoted
18/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

StylePlan: Style-Conditioned Intent Planning for Zero-Shot Coordination

Xiao Su, Jian Zhou

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

The Scaling Laws of Skills in LLM Agent Systems

Qiguang Chen, Qiming Yu, Yuhang Gu, Zhuoye Huang and 11 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Refining Compositional Diffusion for Reliable Long-Horizon Planning

Kyowoon Lee, Yunhao Luo, Anh Tong, Jaesik Choi

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

PROBE: Learning to Audit Policy Compliance in Tool-Using LLM Agents

Kshitij Mishra, Abhijith Sharma, Nils Lukas, Salem Lahlou

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

67%Highly rated
?Highly ratedVote to see the score

Med-Agentic: Distilling Agentic Medical Reasoning with Internalized Meta-Capabilities

Yucheng Zhou, Junwei Sheng, Jianbing Shen

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Fast-WAM: Do World Action Models Need Test-time Future Imagination?

Tianyuan Yuan, Zibin Dong, Yicheng Liu, Hang Zhao

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Unifying Reasoning and Planning through Energy Minimization

Adrian Rodriguez, Angelica Kim, Yunhui Guo, Yilun Du

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Recursive Semantic Divergence for LLM Agent Consistency

Harshavardhan Abichandani, Penny Chong, Atin Ghosh, Daniel Dahlmeier

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

The Geometry of Agent Skills: Non-Commutative Composition in Representation Space

Junda Wu, Yifan Wang, Zihan Huang, Xunyi Jiang and 5 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

67%Highly rated
?Highly ratedVote to see the score

AgentAbstain: Do LLM Agents Know When Not to Act?

Xun Liu, Yi Evie Zhang, Vira Kasprova, Parisa Rabbani and 4 more

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 1/5
medium 1/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

AgentSSL: Can MLE Agents Leverage Unlabeled Data?

Akanksha Sarkar, Ethan Lin, Ziang Liu, Kristin Branson and 1 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

RADAR: Routing Agents via Difficulty-Aware Recovery

Haizhou Du, Jinze Zhao, Huaicheng Yan

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Lemon: Evidence-Risk-Aware Adaptive Organization for Long-Horizon LLM Agents

Zimo Yin, Haipeng Jiang, Kailong Ren, Zhetao Sun and 7 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

PathNavigate: A Training-Free Pathology Agent with Surprise-Guided Scan and Shared Slide Memory for Whole-Slide VQA

Chunze Yang, Qidong Liu, Wenjie Zhao, Yue Tang and 9 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

ExperiGen: Agentic Hypothesis Discovery from Observational Data

Jishu Sen Gupta, Harini S I, Somesh K Singh, Mohamad Tawseeq Syed and 5 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

HierFlow: Hierarchical Coupled Dual-Space Search for Automatic Agentic Workflow Generation

Dong Li, Yanchi Liu, Xujiang Zhao, Wei Cheng and 5 more

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

AgenTracer-v2: Agentic Failure Tracer for LLM Agentic Systems

Guibin Zhang, Haoyu Lu, Junhao Wang, He Zhu and 3 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Breadcrumbing Search Agents: Per-Turn Scheming Over Long-Horizon Trajectories

Xuebin Li, Hanqing Zhao, Siyuan Liang, Kejiang Chen and 3 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

VideoSailor: Navigating Video Deep Research via Trajectory-to-Policy Flywheel

Meng Cao, Pengfei Hu, Yingyao Wang, Chen Wang and 6 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

LLM-ACES: Closed-Loop Discovery of Dynamic Systems with LLM-Guided Adaptive Search

Nikhil Abhyankar, Sha Li, Sanchit Kabra, Naren Ramakrishnan and 2 more

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Shepherd: A Runtime Substrate Empowering Meta-Agents with a Formalized Execution Trace

Simon Yu, Derek Chong, Ananjan Nandi, Dilara Soylu and 3 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Selective Critique for Cost-Aware LLM Agents in Long-Horizon Decision Making

Heewon Park, Somin Im, Minhae Kwon

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Progress-Aware Distillation for Mitigating Stagnation in Small Language Model Agents

Yuanpu Cao, Saket Sathe, Hanyu Wang, Ziyi Yin and 2 more

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

67%Highly rated
?Highly ratedVote to see the score

Can AI Agents Synthesize Scientific Conclusions?

Hayoung Jung, Pedro V Diniz, José R Roveda, Abner F da Silva and 4 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 1/5
medium 1/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Agentic Video Editing from Underspecified Requests

Yongsheng Yu, Ziyun Zeng, Zhiyuan Xiao, Zhenghong Zhou and 3 more

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

67%Highly rated
?Highly ratedVote to see the score

MemeEconomy : Do LLM Agents Trade Ethics for Survival?

Syed Nazmus Sakib, Nafiul Haque, Ahnaf Manan, M. M MORSHED and 1 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

67%Highly rated
?Highly ratedVote to see the score

Calibration Is Not Control: Intervention Advantage for LLM-Agent Oversight

Chubin Zhang, Zhenglin Wan, Xingrui Yu, jingxuan wu and 4 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

67%Highly rated
?Highly ratedVote to see the score

Soteria: Formally Verified Planning with Runtime Enforcement for Safe LLM Agents

Deyuan (Mike) He, Ankush Desai, Sharad Malik, Aarti Gupta

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 1/5
medium 1/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

SeaPilot: Mobile Agent with Self-refining Environment Alignment

Zhigang Zuo, Senyao Li, Yufeng Jiang, Haozhao Wang and 4 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

67%Highly rated
?Highly ratedVote to see the score

MedZERO: Self-Evolving Agents for Open-Ended Medical Reasoning Through Controlled Knowledge Accumulation

Xilin Dang, Weilin Ruan, Xue Yang, Jinghao Wang and 3 more

Paris Poster Session 5, Fri, Dec 11, 11:30 AM–1:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 2/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Temperature Guidance For Robust Reward Conditioning In Diffusion Planning

Johannes V Busch, Roberto Calandra

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Automating ML for Science: Can Frontier Agents Climb Scientific Hills in the Wild?

Ming Zhong, Stacy Li, Nicholas Carlini, Matthew Jagielski

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

AIRA-Compose: Agentic Discovery of Neural Architectures

Alberto Pepe, Chien-Yu Lin, Despoina Magka, Bilge Acun and 4 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Plan2Sense: Open-World Task Planning in Epistemic States via Interleaved Ontic and Sensing Actions

Xiaotian Liu, Armin Toroghi, Jiazhou Liang, Ali Pesaranghader and 1 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Runtime Verification of Multiple Natural Language Criteria for Agent Governance

Silviu Pitis, Parand A. Alamdari, Jessica Tang, Toryn Klassen and 1 more

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

2-Step Agent: How a Bayesian Decision Maker Learns from AI-Decision Support

Otto Nyberg, Fausto Carcassi, Davide Tugnoli, Giovanni Cinà

Paris Poster Session 5, Fri, Dec 11, 11:30 AM–1:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

FEP-Agent: Grounding LLM Agent Self-Evolution in Active Inference with Semantic Memory

Minghao Chen, Xinyi Hu, Zhou Yu, Yufei Yin

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

RAOP: Step-Level Resource Orchestration for LLM Agents across Edge and Cloud

Jinze Li, Xin Yang, Shuo Yang, Jinfeng Xu and 1 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Sampling Is Not Curiosity: Why LLM Agents Should Investigate

Alfonso Amayuelas, Piotr Piękos, Xin Wang, William Yang Wang and 2 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

The Agentic Oversight Tax: Human Supervision of AI Agents Has a Cost that Must be Accounted For

Olivier Oullier

Paris Poster Session 6, Fri, Dec 11, 2:30 PM–4:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Treat Domain-Specific Languages as Design Variables in LLM Agents

Letian Li, Wenyuan Jiang, Xin Yang, Shuzhao Xie and 4 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Long-Horizon Agency Belongs in the Harness, Not the Context Window Only

Yi Han, YUANYUAN XU, jusheng zhang, Wenhao Wang

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Multimodal Foundation Agents Should Use Brain Data as Privileged Supervision

Richard Csaky

Paris Poster Session 1, Wed, Dec 9, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Agent-Native Research Artifacts

Jiachen Liu, Jiaxin Pei, Jintao Huang, Chenglei Si and 33 more

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

67%Highly rated
?Highly ratedVote to see the score

Agentic Abstention: Do Agents Know When to Stop Instead of Act?

Han Luo, Bingbing Wen, Lucy Lu Wang

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 1/5
medium 1/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Predicting What Changes: Causal Delta World Models for Risk-Aware LLM Agent Planning

Guo Yue, YANG LIU, Donghui Zhang, Tang Qingkang and 2 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

MatCurvs: Article Real-Coordinate Curve Extraction for Agent-Ready Materials Reasoning

Liang Yin, Zhan'ao yao, Jiahui Shi, Songlin Yu and 1 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Explanations over Graphs: An Agent Architecture for IT Enterprise Diagnostic Tasks

Saurabh Jha, Rohan R. Arora, Bhavya Bhavya, Noah Zheutlin and 5 more

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

PolyMind: Exploring Width Scaling for Reflective Reasoning in Language Agents

Heng Zhang, Chengyu Zhou, Jiajun Wu, Estella Liu and 7 more

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Who Wrote This Paper? Autonomous Scientific Discovery for 3DGS Research

Seemandhar Jain, Manmohan Chandraker

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

DoG: Sniffing Out Overconfidence in LLM Agents via Post-hoc Trajectory Restructuring

Hyunjun Jeon, Dongha Lim, Kunwoong Kim, Daewon Choi and 1 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Tokenizer Choice Shapes Generalization in State-Centric Learning for Planning

Vishal Pallagani, Nitin Gupta, John A Aydin, Biplav Srivastava

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

DrugSAGE: Self-evolving Agent Experience for Efficient State-of-the-Art Drug Discovery

Yikun Zhang, Xiwei Cheng, Tianyu Liu, Yuanqi Du and 1 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

AURA: An Autonomous Retouching Agent with Photographic Visual Thinking

Shuaizheng Liu, Fangzhou Han, Jiarong Liao, Yujing Sun and 7 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

AIRA 2: Overcoming Bottlenecks in AI Research Agents

Karen Hambardzumyan, Nicolas Baldwin, Edan Toledo, RISHI HAZRA and 21 more

Paris Poster Session 5, Fri, Dec 11, 11:30 AM–1:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

MultiSTEVE-1s: A Model Zoo and Interpretability Suite for Instruction-Following Vision Agents

Karolis Jucys, George Adamopoulos, Özgür Şimşek

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

When Further Realization Is Unnecessary: Amortized Reasoning for Long-Horizon LLM Agents

Rongzheng Wang, Jiakai Li, Renzhong Wang, Rongwei Wang and 5 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

MMGraph-Agent: Agentic Multimodal RAG via Cache-Inspired Multimodal Knowledge HyperGraphs

Haoran Luo, Ziyue Zhu, Shangyang Wu, LINHAO LUO and 7 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Who&When Pro: Can LLMs Really Attribute Failures in AI Agents?

Jiale Liu, Huajun Xi, Shaokun Zhang, Yifan Zeng and 5 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Agent Explorative Policy Optimization for Agentic Multimodal Reasoning

Minki Kang, Shizhe Diao, Ryo Hachiuma, Sung Ju Hwang and 3 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
86%Must read
?Must readVote to see the score

What Fits (Into Few Tokens) Doesn't Overfit: Compression and Generalization in ML Research Agents

LLM research agents find high-performance models via compressed prompts and feedback, supporting a description-length explanation for limited overfitting.

Martin Bertran, Aaron Roth, Steven Wu

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

74%Highly rated
?Highly ratedVote to see the score

How to Interpret Agent Behavior

ACT*ONOMY introduces a three-level taxonomy of 10 actions and 46 subactions for describing autonomous agent behavior at runtime, plus an open repository and automated analysis pipeline that compares behavioral profiles and surfaces failure patterns.

Sophia Gao, Kaiser Sun, Jen-Tse Huang, Katherine Van Koevering and 7 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 1 on Hugging Face

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

71%Highly rated
?Highly ratedVote to see the score

Agentic Multi-Turn Reasoning: A Fairness Approach

Fair-MPO improves multi-turn agentic reasoning via multi-level preference optimization and a fairness objective that fixes long-horizon credit assignment and data imbalance, achieving state-of-the-art benchmark results.

Thanh-Dat Truong, Sankalp Pandey, Hugh Churchill, Jackson Cothren and 2 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score

Reinforcement World Model Learning for LLM-based Agents

RWML learns action-conditioned world models for LLM agents via self-supervised sim-to-real alignment, outperforming direct task-success RL by up to 6.9 points without expert data.

Xiao Yu, Baolin Peng, Ruize Xu, yelong shen and 5 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 28 on Hugging Face

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

74%Highly rated
?Highly ratedVote to see the score

PhotoFlow: Agentic 3D Virtual Photography Missions

PhotoFlow uses a Director-Reviewer-Reflector agent for closed-loop camera search to generate language-conditioned virtual photographs in arbitrary 3D scenes, outperforming baselines on quality, alignment, and success rate.

Jiarui Guo, Haojia Wei, Yiming Zhang, Yifei Liu and 4 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 24 on Hugging Face · Code ★ 43

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

71%Highly rated
?Highly ratedVote to see the score

CTM-AI: A Blueprint for General AI Inspired by a Model of Consciousness

CTM-AI combines a consciousness model with foundation models to integrate diverse processors, achieving state-of-the-art results on multiple benchmarks.

Haofei Yu, Yining Zhao, Lenore Blum, Manuel Blum and 1 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Strategic Decision Support for AI Agents

Strategic decision support minimizes AI agent support usage via threshold policies controlling counterfactual missed-support error without distributional assumptions.

Shayan Kiyani, Sima Noorani, George J. Pappas, Hamed Hassani

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

OASES: Outcome-Aligned Search-Evaluation Co-Training for Agentic Search

OASES co-trains a search policy and adaptive evaluator to provide outcome-aligned process rewards, outperforming RL baselines on multi-hop QA benchmarks.

Erhan Zhang, Yiqun Chen, Zechun Niu, Wei Yang and 5 more

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

71%Highly rated
?Highly ratedVote to see the score

PhGPO: Pheromone-Guided Policy Optimization for Long-Horizon Tool Planning

PhGPO learns reusable tool-transition patterns from past trajectories via pheromone guidance to improve long-horizon tool planning.

Yu Li, Guangfeng Cai, shengtian yang, Han Luo and 4 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
6/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score

LensDesigner: A Self-Improving Agent for Optical Lens Design

LensDesigner is a self-improving autonomous agent that uses retrieval, simulation, and curriculum learning to design optical lenses, significantly outperforming baselines on 120 diverse tasks.

Lei Sun, Haoran Liang, Dannong Xu, Yao Gao and 5 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

91%Must read
?Must readVote to see the score

Beyond Suspicious Steps: Ontological Trust in Long-Horizon Agents

Ontological trust measures whether trajectory prefixes match authorized tasks; RGE detects long-horizon agent drift with over 93% F1 and above 95.8% benign coverage via deterministic Role, Goal, and Evidence checks.

一个 他, Yao Wang, Haibin Zhang

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
18/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score

Efficient Lookahead Encoding and Abstracted Width for Learning General Policies in Classical Planning

Holistic relational encoding and abstracted width enable GNN policies to learn general classical planning strategies efficiently, surpassing LAMA on IPC 2023 benchmarks.

Michael Aichmüller, Simon Ståhlberg, Martin Funkquist, Hector Geffner

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 2/5
medium 7/10
strict 1/5
91%Must read
?Must readVote to see the score

Towards On-Policy Data Evolution for Visual-Native Multimodal Deep Search Agents

A visual-native harness with an image bank and on-policy data evolution improves multimodal deep search agents, raising Qwen3-VL-8B to 39.0% average and surpassing Gemini-2.5 Pro.

Shijue Huang, Hangyu Guo, Guanting Dong, Chenxin Li and 7 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 21 on Hugging Face · Code ★ 30

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

HACO: Hedged Agent Computing for Reliable LLM Systems

HACO treats role requests as reliability-constrained selection over agent instances, adaptively hedging candidates via optimistic ranking plus conservative reliability accumulation to improve robustness at lower cost.

Enhan Li, Hongyang Du

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Sequential Behavioral Watermarking for LLM Agents

SeqWM embeds watermarks into history-conditioned agent transition patterns, enabling position-agnostic trajectory verification that resists corruption where step-indexed methods fail.

Hyeseon An, Shinwoo Park, Dongsu Kim, Yo-Sub Han

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 0/5
91%Must read
?Must readVote to see the score

Automata from Agent Traces: Failure and Next-Step Prediction

Trace corpora collapse into compact finite-state machines replaying held-out data at >=0.997 fitness, yielding state-context next-step prediction and 0.94 AUROC failure prediction for runtime monitoring.

Seonglae Cho, Franklin Cardenoso Fernandez, Umar Mohammed, Zekun Wu and 3 more

Paris Poster Session 2, Wed, Dec 9, 5:00 PM–7:00 PM, Paris Poster Hall · Published 2026 · ▲ 5 on Hugging Face

– ReadersNo votes yet
18/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

A Self-Evolving Framework for Efficient Terminal Agents via Observational Context Compression

TACO is a training-free framework that learns adaptive compression rules from terminal agent trajectories to filter noisy observations, improving accuracy by 1-4% and reducing token usage across benchmarks.

JinCheng Ren, Siwei Wu, Yizhi Li, Zhu and 9 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 24 on Hugging Face · Code ★ 47

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 1/5
91%Must read
?Must readVote to see the score

Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA

VTS frames grounded long-video QA as self-correcting search over an adaptive temporal tree with explicit backtracking, improving grounding and answer accuracy across benchmarks.

Ce Zhang, Ziyang Wang, Yulu Pan, Oluwatumininu Oguntola and 5 more

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score

Agentic AI Scientists Are Not Built For Autonomous Scientific Discovery

Agentic AI scientists serve as co-scientists but lack autonomous discovery due to flawed problem selection, missing tacit lab knowledge, compressed diversity, and inadequate benchmarks.

Harshit Bisht, Vinay Kumar, Kevin Maik Jablonka, Mausam and 1 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 5/5
medium 4/10
strict 1/5
88%Must read
?Must readVote to see the score

VisInteract: Towards Dynamic Interactive Text-to-Visualization under Imperfect Queries

VisInteract introduces interactive text-to-visualization with imperfect queries via VisInteract-Bench and Vis-MCTS, boosting success by over 13% versus interactive baselines.

Wenxin XU, Jinwei Lu, Hwanhee Kim, Chen J Zhang and 3 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

VeriGraph: Towards Verifiable Data-Analytic Agents

VeriGraph proposes a neuro-symbolic framework that builds explicit evidence DAGs to make LLM data-analysis agents verifiable, achieving 87.61% claim grounding.

Jiajie Jin, Zhao Yang, WenLe Liao, Yuyang Hu and 4 more

Paris Poster Session 4, Thu, Dec 10, 5:30 PM–7:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

GenEnv: Difficulty-Aligned Co-Evolution Between LLM Agents and Environment Simulators

GenEnv co-evolves LLM agents with generative simulators via difficulty-aligned curricula, improving 7B agents by up to 40.3% with 3.3x less data.

Jiacheng Guo, Ling Yang, Peter Chen, Qixin Xiao and 4 more

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026 · ▲ 19 on Hugging Face · Code ★ 67

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

Hyperagents

Hyperagents integrate editable task and meta agents to enable metacognitive self-modification, with DGM-H improving across domains and accumulating meta-level improvements.

Jenny Zhang, Bingchen Zhao, Wannan Yang, Jakob Foerster and 4 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 1/5
80%Must read
?Must readVote to see the score

Autonomous Scientific Discovery via Iterative Meta-Reflection

DiscoPER autonomously discovers scientific patterns via iterative meta-reflection and statistical testing, recovering 8 of 9 ecological patterns and outperforming baselines.

Bingchen Zhao, Sara Beery, Oisin Mac Aodha

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

91%Must read
?Must readVote to see the score

PACE: Two-Timescale Self-Evolution for Small Language Model Agents

PACE uses two-timescale self-evolution to let frozen small language models improve agents via validated prompt and control updates, outperforming baselines on 12 settings by up to 9.2%.

Chen Ling, Pei Chen, Xiangchen Guan, Jiaming Qu and 3 more

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
18/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Securing AI Agents with Information-Flow Control

Information-flow control secures AI agents via Fides, a planner with deterministic confidentiality and integrity tracking that completes diverse AgentDojo tasks with guarantees.

Manuel Costa, Boris Köpf, Aashish Kolluri, Andrew Paverd and 5 more

Paris Poster Session 3, Thu, Dec 10, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026 · ▲ 1 on Hugging Face · Code ★ 118

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

Lean4Agent: Formal Modeling and Verification for Agent Workflow and Trajectory

Lean4Agent uses Lean4 to formally model and verify agent workflows, with verified workflows outperforming failing ones by 11.94% and LeanEvolve improving SWE performance by 7.47%.

Ruida Wang, Jerry Huang, Pengcheng Wang, Xuanqing Liu and 2 more

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026 · ▲ 7 on Hugging Face · Code ★ 31

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 0/5
80%Must read
?Must readVote to see the score

Ares: Adaptive Reasoning Effort Selection for Efficient LLM Agents

Ares uses a lightweight router to select per-step reasoning effort for LLM agents, cutting reasoning tokens by up to 52.7% with minimal accuracy loss.

Jingbo Yang, Bairu Hou, Jiayun (Peter) Wang, Wei Wei and 2 more

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026 · ▲ 1 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Bian Que: An Agentic Framework with Flexible Skill Arrangement for Online System Operations

Bian Que is an agentic framework that arranges flexible skills for online system operations, reducing alerts by 75% and cutting resolution time by over 50%.

bochao liu, Zhipeng Qian, yang zhao, Xinyuan Jiang and 8 more

Paris Poster Session 1, Wed, Dec 9, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score

FlowSteer: Towards Agents Designing Agentic Workflows via Reinforced Progressive Canvas Editing

FlowSteer lets an agent design executable agentic workflows via a canvas environment that returns syntax-checked feedback for each edit, significantly outperforming baselines across twelve datasets.

Mingda Zhang, Wenjin Liu, Tiesunlong Shen, Qika Lin and 4 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

70%Highly rated
?Highly ratedVote to see the score

RelAgent: LLM Agents as Data Scientists for Relational Learning

RelAgent is an LLM agent that builds SQL feature queries and selects predictive models for relational learning, yielding fast, interpretable predictions deployable via standard databases.

Xingyue Huang, Louis Tichelman, Jinwoo Kim, Krzysztof Olejniczak and 1 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
5/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

69%Highly rated
?Highly ratedVote to see the score

Differentiable Learning of Lifted Action Schemas for Classical Planning

A neural architecture learns lifted STRIPS action schemas from fully observed state traces with hidden action arguments, recovering ground-truth domain structures robustly.

Jonas Reiter, Jakob Gebler, Hector Geffner

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
3/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Tracing Agentic Failure from the Flow of Success

OAT trains neural controlled differential equations on successful agent trajectories to detect failure steps without failure annotations, outperforming prompting baselines by up to 20% F1 with 200-5000x speedup.

Samuel (Min-Hsuan) Yeh, Yiwen Zhu, Shaleen Deep, Sharon Li

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026 · ▲ 16 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 1/5
86%Must read
?Must readVote to see the score

A Few Teacher Steps Go a Long Way: Cost-Efficient On-Policy Data Augmentation for Agent Post-Training

Few-step unfiltered teacher continuations at learner-induced contexts cost-efficiently outperform pure behavioral cloning and filtered long completions across agent benchmarks at matched budgets.

Junze Ye, Jiayi Cheng, Miao Lu, Michal Mankowski and 2 more

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 2/5
89%Must read
?Must readVote to see the score

Beyond Semantic Similarity: Rethinking Retrieval for Agentic Search via Direct Corpus Interaction

Direct corpus interaction uses terminal tools to search raw corpora directly, bypassing fixed retrieval interfaces and substantially outperforming sparse, dense, and reranking baselines on agentic search benchmarks.

Zhuofeng Li, Haoxiang Zhang, Cong Wei, Pan Lu and 14 more

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026 · ▲ 125 on Hugging Face · Code ★ 408

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 2/5
83%Must read
?Must readVote to see the score

Self-Improving World Modelling with Latent Actions

SWIRL learns world models from state-only sequences via latent actions and alternating forward/inverse dynamics, improving LLM/VLM reasoning benchmarks by up to 28%.

Yifu QIU, Zheng Zhao, Waylon Li, Yftah Ziser and 3 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 32 on Hugging Face · Code ★ 20

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 2/5
86%Must read
?Must readVote to see the score

Knowledge-Graph Paths as Intermediate Supervision for Self-Evolving Search Agents

Knowledge-graph paths provide intermediate supervision for self-evolving search agents, improving question validity via relational context and solver rewards via waypoint coverage, boosting multi-hop QA across benchmarks.

Huyu Wu, Jun Liu, Xiaochi Wei, Yan Gao and 2 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 1/5
83%Must read
?Must readVote to see the score

Harnessing Agentic Evolution

AEvo formulates agentic evolution as an interactive environment where a meta-agent edits the evolution procedure to steer long-horizon search, achieving up to 26% relative improvement over baselines.

Jiayi Zhang, Yongfeng Gu, Jianhao Ruan, Maojia Song and 8 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 2/5
88%Must read
?Must readVote to see the score

From History to State: Constant-Context Skill Learning for LLM Agents

Constant-context skill learning embeds recurring agent workflows into lightweight modules via step-level SFT and online RL, cutting prompt tokens 2-7x while matching state-of-the-art success on ALFWorld, WebShop, and SciWorld.

Haoyang Xie, Xinyuan Wang, Yancheng Wang, Puda Zhao and 1 more

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 2/5
74%Highly rated
?Highly ratedVote to see the score

Agent Security is a Systems Problem

Agent security requires systems-level invariants treating AI as untrusted, since model robustness alone cannot prevent real-world agent attacks.

Mihai Christodorescu, Earlence Fernandes, Ashish Hooda, Somesh Jha and 10 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 5/5
medium 3/10
strict 1/5
Show 20 more papers