Good Papers

SITA: Learning Speaker-Invariant and Tone-Aware Speech Representations for Low-Resource Tonal Languages

SITA adapts wav2vec-style encoders via staged optimization to learn speaker-invariant, tone-aware representations, improving cross-gender lexical retrieval and tone separation on Hmong and Mandarin while preserving ASR accuracy.

Tianyi Xu, Xuan Ouyang, 姚斌伟, Shoua Xiong, Sara M. Misurelli, Maichou Lor, Junjie Hu

Published Jan 14, 2026arXiv ↗

75%
OverallHighly rated
?
OverallHighly ratedVote to see the score
Readers
–

Only vote on papers you've read. Sign in to vote.

AI panel10/20reviewers recommend it
lenient 5/5
medium 5/10
strict 0/5
AI panel?Vote to see what the 20 AI reviewers said

Abstract

Tonal low-resource languages are widely spoken but remain underserved by modern speech technologies. A central challenge is learning speech representations that are robust to nuisance variation, such as speaker gender, while preserving lexical tone, which carries word meaning. We propose SITA, a lightweight adaptation recipe for pretrained wav2vec-style self-supervised speech encoders. Rather than designing a new backbone or objective, SITA combines existing objectives in a staged optimization framework to reduce tone collapse while preserving ASR capability. Stage 1 improves speaker invariance without erasing tonal contrasts by combining a cross-gender contrastive loss with a tone-repulsive loss that separates same-word, different-tone realizations. Stage 2 restores recognition-oriented linguistic information through CTC fine-tuning and knowledge distillation on upper encoder layers. We evaluate SITA primarily on Hmong, a tonal language with limited digital resources and a small speaker pool. Against multilingual, speaker-adversarial, label-aware, and semi-supervised baselines, SITA achieves the best trade-off between cross-gender lexical retrieval and tone separation, while maintaining ASR accuracy close to an ASR-adapted XLS-R teacher. Results on Mandarin show consistent gains, suggesting that SITA is a general plug-in recipe for tonal speech representation learning.