Good Papers

WAXAL: A Large-Scale Multilingual African Language Speech Corpus

WAXAL introduces an open 1,250-hour ASR and 235-hour TTS speech corpus for 24 African languages to advance inclusive speech technology.

MohamedElfatih MohamedKhair, Emmanuel Asiedu Brempong, Subhashini Venugopalan, Abdoulaye Diack, Perry Nelson, Mireku, Tavonga Siyavora, Pooja Rao, Uche Okonkwo, Angela Nakalembe, Abhishek Bapna, Aisha Walcott-Bryant

Published 2026Sydney Poster Session 4 · Wed, Dec 9, 5:00 PM–8:00 PM local time · Hall 1-4▲ 4 on Hugging FacearXiv ↗OpenReview ↗

74%
OverallHighly rated
?
OverallHighly ratedVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel9/20reviewers recommend it
lenient 5/5
medium 3/10
strict 1/5
AI panel?Vote to see what the 20 AI reviewers said

Abstract

The advancement of speech technology has predominantly favored high-resource languages, creating a significant digital divide for speakers of most Sub-Saharan African languages. To address this gap, we introduce WAXAL, a large-scale, openly accessible speech dataset for 24 languages representing over 100 million speakers. The collection consists of two main components: an Automated Speech Recognition (ASR) dataset containing approximately 1,250 hours of transcribed, natural speech from a diverse range of speakers, and a Text-to-Speech (TTS) dataset with around 235 hours of high-quality, single-speaker recordings reading phonetically balanced scripts. This paper details our methodology for data collection, annotation, and quality control, which involved partnerships with four African academic and community organizations. We provide a detailed statistical overview of the dataset and discuss its potential limitations and ethical considerations. The WAXAL datasets are released at https://huggingface.co/datasets/google/WaxalNLP under the permissive CC-BY-4.0 license to catalyze research, enable the development of inclusive technologies, and serve as a vital resource for the digital preservation of these languages.