Good Papers

AskToAct: Enhancing LLMs Tool Use via Self-Correcting Clarification

AskToAct improves LLM tool use via self-correcting clarification by automatically generating clarification training data from tool parameters and adding error-correction mechanisms, achieving over 57% intent recovery accuracy and 10.46% greater efficiency while generalizing to unseen APIs.

Xuan Zhang, Yongliang Shen, Zheng, Zhe, Linlin Wu, Wen Qi Zhang, Yuchen Yan, Qiuying Peng, Wang, Jun, Weiming Lü

Published Mar 3, 2025arXiv ↗

80%
OverallMust read
?
OverallMust readVote to see the score
Readers
–

Only vote on papers you've read. Sign in to vote.

AI panel12/20reviewers recommend it
lenient 5/5
medium 7/10
strict 0/5
AI panel?Vote to see what the 20 AI reviewers said

Abstract

Large language models (LLMs) have demonstrated remarkable capabilities in tool learning. In real-world scenarios, user queries are often ambiguous and incomplete, requiring effective clarification. However, existing interactive clarification approaches face two critical limitations: reliance on manually constructed datasets, which inherently constrains training data scale and diversity, and lack of error correction mechanisms during multi-turn clarification, leading to error accumulation that compromises both accuracy and efficiency. We present AskToAct, which addresses these challenges by exploiting the structural mapping between queries and their tool invocation solutions. Our key insight is that tool parameters naturally represent explicit user intents. By systematically removing key parameters from queries while retaining them as ground truth, we enable automated construction of high-quality training data. We further enhance model robustness through error-correction pairs and selective masking, enabling dynamic error detection during clarification interactions. Comprehensive experiments demonstrate that AskToAct significantly outperforms existing approaches, achieving above 57% accuracy in recovering critical unspecified intents and enhancing clarification efficiency by an average of 10.46% while maintaining high accuracy in tool invocation. Our framework exhibits robust performance across different model architectures and successfully generalizes to entirely unseen APIs without additional training, achieving performance comparable to GPT-4o with substantially fewer computational resources.