Good Papers

Showing Tool use & function calling Show all papers

90%Must read
?Must readVote to see the score

From Evidence to Action: How Tool-Using Agents Fail

Tool-using agents often fail by acting before establishing required evidence or leaving multi-action workflow prerequisites unresolved, despite accurate static action assessment. SafeActBench reveals failures stem from how agents use established evidence during execution, not just missing informatio

Hongzhan Lin, Shidong Cao, Ziyang Luo, Wenhao Chai and 2 more

Published Oct 6, 2026 · ▲ 20 on Hugging Face · Code ★ 3

100% Readers1 of 1 upvoted
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 3/5
76%Highly rated
?Highly ratedVote to see the score

Hiding Tool Latency in On-Device Cascaded Voice Agent through Speculative Execution

Speculative tool execution predicts tool calls from partial ASR to run them during speech, cutting median voice-agent response latency from 5.79 s to 4.60 s.

Kyudan Jung, Hyunsin Park, Yoonhyung Lee, Jinhwan Park and 4 more

Published Oct 6, 2026 · ▲ 1 on Hugging Face

100% Readers1 of 1 upvoted
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 7 of 20 reviewers recommend it
lenient 4/5
medium 2/10
strict 1/5
86%Must read
?Must readVote to see the score

Judged Useless, Queried Anyway: Tool-Using Agents Rarely Turn Their Own Evidence Judgments into Stopping Decisions

Agents judge failing retrieval results useless but rarely stop; enforcing answers after five useless judgments improves success and fixes stopping.

Chubin Zhang, Zhenglin Wan, Xingrui Yu, Jingxuan Wu and 3 more

Published Oct 5, 2026 · ▲ 3 on Hugging Face · Code ★ 3

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 3/5
78%Highly rated
?Highly ratedVote to see the score

EvoDuet: Bilevel Co-Evolution of Web Searching and Task Solving for Scientific Discovery

EvoDuet co-evolves solutions and web queries via a retrieval gate to boost LLM discovery gains up to 82.3% across optimization tasks.

Young-Jun Lee, Jinheon Baek, Soyeong Jeong, Minki Kang and 4 more

Published Sep 30, 2026 · 0 citations · ▲ 109 on Hugging Face · Code ★ 4

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 5/5
medium 5/10
strict 1/5
86%Must read
?Must readVote to see the score

WEFT: Scaling Tool-Use Post-Training for General-Purpose Agents

WEFT evolves whole agentic interaction systems for tool-use post-training, outperforming environment-scaling baselines by up to 12.27 points across benchmarks.

Bo Mao, Hang He, Linting Wang, Lizhi Lin and 16 more

Published Sep 29, 2026 · 0 citations · ▲ 19 on Hugging Face

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 2/5
57%Worth a look
?Worth a lookVote to see the score

Optimizing Agent Tool-Use via Trajectory-based Insight Evolution

Hanchen Qiu, Haojia Zhu, Yifan Meng, Jiahui Jin

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Masked Diffusion Language Agents for Tool-Integrated Chemical Reasoning

Mengdi Liu, Chenghao Jia, Hong Chang, Shiguang Shan

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

I-PTC: Interactive Programmatic Tool Calling for Stateful Tool-Augmented Agents

Huanzhi Mao, Chengkun Cao, Shuo Yuan, Joseph Gonzalez

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

Thinking with Imitation: Adaptive Reinforcement-Imitation Learning for Tool-Augmented Scientific Reasoning

Fanrui Zhang, Hongmin Zhan, Qiang Zhang, Sizhuo Zhou and 4 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 2/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Choosing Before Acting: Comparative Value Estimation for Long-Horizon Tool-Use Agents

Yu Li, Zheng Zhang, Xin Liu, shengtian yang and 2 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Relevance Is Not Necessity: Selecting the Necessary API Set for Tool-Using LLMs

Chen Wang

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

The Scaling Paradox of Tool-Calling LLM Agents Under Realistic MCP Faults

Antonio Ken Iannillo, Joshua S Owotogbe, Roberto Natella, Francesco Avallone and 2 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 1/5
57%Worth a look
?Worth a lookVote to see the score

PORTool: Importance-Aware Policy Optimization with Rewarded Tree for Multi-Tool-Integrated Reasoning

Feijie Wu, Weiwu Zhu, Yuxiang Zhang, Soumya Chatterjee and 4 more

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Advancing Affordance-Grounded Creative Tool Use in Large Multimodal Models

Cheng Qian, Hyeonjeong Ha, Jiayu Liu, Jeonghwan Kim and 8 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
91%Must read
?Must readVote to see the score

LLM Agents Already Know When to Call Tools - Even Without Reasoning

When2Tool finds LLMs linearly encode tool necessity in hidden states, and Probe&Prefill uses this to cut unnecessary tool calls by 48% with minimal accuracy loss.

Chung-En Sun, Linbo Liu, Ge Yan, Zimo Wang and 1 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 2 on Hugging Face · Code ★ 16

– ReadersNo votes yet
18/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 18 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 4/5
88%Must read
?Must readVote to see the score

PEARL: Solver-in-the-Loop Interactive Optimization Modeling from Natural Language

PEARL integrates solvers into an interactive optimization modeling loop to iteratively revise formulations using execution feedback, substantially boosting verified solve rates and enabling a small model to outperform a much larger baseline.

Hongliang Lu, Zhong Li, Yuxuan Chen, Lan Yuan and 2 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 2/5
83%Must read
?Must readVote to see the score

SpecHop: Continuous Speculation for Accelerating Multi-Hop Retrieval Agents

SpecHop continuously speculates multi-hop retrieval trajectories via asynchronous verification and branch rollback, reducing latency by up to 40% losslessly.

Mehrdad Saberi, Keivan Rezaei, Soheil Feizi

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 1/5
83%Must read
?Must readVote to see the score

AgentBrew: Offline Tool-Use Agent Learning from Raw Real-World Trajectories

AgentBrew learns tool-use policies offline from raw real-world trajectories via retrospective task inference and PMI-based credit assignment, improving Qwen3-32B by +8.7 accuracy over larger baselines.

Zhiyi Lyu, Yewen Li, Longtao Zheng, shengtian yang and 6 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 1/5
78%Highly rated
?Highly ratedVote to see the score

OPERA: An Agent for Image Restoration with End-to-End Joint Planning–Execution Optimization

OPERA jointly optimizes restoration planning via reinforcement learning and tool execution via co-training to outperform existing methods on complex mixed degradations.

Feng Zhu, Shuyang Xie, Zeng Yihan, Ming Liu and 1 more

Paris Poster Session 1, Wed, Dec 9, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026 · Code ★ 25

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 0/5
76%Highly rated
?Highly ratedVote to see the score

Train the Agent, Not the Expert: Learning to Harness Heterogeneous Experts for Multi-Turn Visual Reasoning

VisHarness trains a visual agent to orchestrate heterogeneous experts for multi-turn reasoning, achieving strong results on segmentation, detection, and counting tasks.

Yaowu Fan, Tao Han, Dazhao Du, Jinhua Ma and 1 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 0/5
Show 20 more papers