Yo-ByT5: Efficient and High-Fidelity Diacritic Restoration for Yorùbá
Yo-ByT5 is a byte-level Yorùbá diacritic restoration model matching mT5-base accuracy with half the parameters and superior text fidelity.
Published Oct 1, 2026arXiv ↗

Only vote on papers you've read. Sign in with GitHub to vote.
Yo-ByT5 delivers efficient, byte-level Yorùbá diacritic restoration with half the parameters of mT5-base and genuinely deployable code, though its 10% DER and unverified fidelity gains depend on a benchmark the authors admit is too small to…
Abstract
Yorùbá is a widely spoken tonal language that depends on diacritics to avoid lexical ambiguity. However, it is often written without these diacritics, thereby hindering downstream Natural Language Processing (NLP) tasks. In this paper, we introduce Yo-ByT5, a byte-level Automatic Diacritic Restoration (ADR) model fine-tuned from ByT5-small. We evaluate Yo-ByT5 alongside five publicly released Yorùbá ADR models and one open-weight large language model (LLM) on the YAD benchmark under a consistent protocol. Our results demonstrate that Yo-ByT5 matches the performance of the strongest existing model, mT5-base, with a DER of 10.14% and a CER of 3.48%. Furthermore, it exhibits superior text fidelity despite using approximately half the parameter count of mT5-base. We also release our training code and model outputs, as well as call for the development of a larger, purpose-built benchmark for Yorùbá diacritic restoration.