Good Papers

Toward Elastic Speech Inference: Training-Free Wake-Word Detection from Pretrained ASR

Pretrained ASR backbones enable training-free wake-word detection, and PCA-based pruning retains performance at 50% encoder reduction.

Hwayeon Kim, Youngwon Choi, Hyeonyu Kim

Published Oct 1, 2026arXiv ↗

72%
OverallHighly rated
?
OverallHighly ratedVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel8/20reviewers recommend it
lenient 4/5
medium 2/10
strict 2/5
AI panel?Vote to see what the 20 AI reviewers said
Panel consensus
An impressive training-free method achieves stable wake-word detection with 50% PCA pruning, though vague "relatively stable" metrics and missing acoustic-condition, clustering, and false-positive benchmarks leave the encoder's true phonetic resilience unproven.

Abstract

Recent ASR development has placed growing emphasis on generalization across diverse domains and acoustic conditions. Existing approaches typically adapt pretrained ASR models to front-end functions such as wake-up word (WuW) detection through additional training or task-specific modules. In this work, we explore the use of a shared pretrained ASR backbone for WuW detection without gradient-based fine-tuning and examine whether a compact encoder can be extracted using the PCA-based structured pruning approach of SliceGPT. Experiments with Parakeet-TDT-0.6B-v3 and Moonshine-base show that WuW detection performance remains relatively stable when the encoder channel dimension is reduced by 50%. These results suggest that task-relevant compact encoders can be derived from pretrained ASR models without fine-tuning.