Hiding Tool Latency in On-Device Cascaded Voice Agent through Speculative Execution
Speculative tool execution predicts tool calls from partial ASR to run them during speech, cutting median voice-agent response latency from 5.79 s to 4.60 s.
Published Oct 6, 2026▲ 1 on Hugging FacearXiv ↗
Only vote on papers you've read. Sign in with GitHub to vote.
Speculative tool execution trims 1.2 seconds of median dead air on a live Android assistant by betting on partial ASR, but the rule-based validator and missing false-injection rates, energy costs, and dataset scope leave its worst-case…
Abstract
Tool-augmented speech assistants typically serialize automatic speech recognition, large language model inference, and external tool execution. As a result, tool latency is incurred only after the user has finished speaking and the LLM has identified the required tool calls. We present speculative tool execution for on-device cascaded voice agents, which predicts tool requests from partial ASR hypotheses and initiates tool execution while speech is still being received, thereby reducing end-to-end response latency. Our approach introduces a Predictor module that anticipates tool calls during speech recognition, executes them speculatively, and caches the results. The cached outputs are then injected into the LLM prompt, enabling faster responses. Additionally, to mitigate errors caused by user self-corrections during speech, we employ a rule-based validation mechanism that selectively injects only valid cached results. As a final safeguard, the LLM retains the ability to issue tool calls directly, ensuring that the latency of our framework is upper-bounded by the baseline serial execution pipeline in the worst case. We evaluate our method using live measurements from a fully implemented Android voice assistant. Our approach reduces the median time-to-first-audio from 5.79,s to 4.60,s and decreases the standard deviation from 3.49,s to 2.81,s, resulting in more predictable response latency.