Toward Elastic Speech Inference: Training-Free Wake-Word Detection from Pretrained ASR
Pretrained ASR backbones enable training-free wake-word detection, and PCA-based pruning retains performance at 50% encoder reduction.
Published Oct 1, 2026arXiv ↗

Only vote on papers you've read. Sign in with GitHub to vote.
An impressive training-free method achieves stable wake-word detection with 50% PCA pruning, though vague "relatively stable" metrics and missing acoustic-condition, clustering, and false-positive benchmarks leave the encoder's true phonetic resilience unproven.
Abstract
Recent ASR development has placed growing emphasis on generalization across diverse domains and acoustic conditions. Existing approaches typically adapt pretrained ASR models to front-end functions such as wake-up word (WuW) detection through additional training or task-specific modules. In this work, we explore the use of a shared pretrained ASR backbone for WuW detection without gradient-based fine-tuning and examine whether a compact encoder can be extracted using the PCA-based structured pruning approach of SliceGPT. Experiments with Parakeet-TDT-0.6B-v3 and Moonshine-base show that WuW detection performance remains relatively stable when the encoder channel dimension is reduced by 50%. These results suggest that task-relevant compact encoders can be derived from pretrained ASR models without fine-tuning.