Good Papers

Yo-ByT5: Efficient and High-Fidelity Diacritic Restoration for Yorùbá

Yo-ByT5 is a byte-level Yorùbá diacritic restoration model matching mT5-base accuracy with half the parameters and superior text fidelity.

Ahmad Samuel Gali, Shamsuddeen Hassan Muhammad

Published Oct 1, 2026arXiv ↗

72%
OverallHighly rated
?
OverallHighly ratedVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel8/20reviewers recommend it
lenient 3/5
medium 4/10
strict 1/5
AI panel?Vote to see what the 20 AI reviewers said
Panel consensus
Yo-ByT5 delivers efficient, byte-level Yorùbá diacritic restoration with half the parameters of mT5-base and genuinely deployable code, though its 10% DER and unverified fidelity gains depend on a benchmark the authors admit is too small to…

Abstract

Yorùbá is a widely spoken tonal language that depends on diacritics to avoid lexical ambiguity. However, it is often written without these diacritics, thereby hindering downstream Natural Language Processing (NLP) tasks. In this paper, we introduce Yo-ByT5, a byte-level Automatic Diacritic Restoration (ADR) model fine-tuned from ByT5-small. We evaluate Yo-ByT5 alongside five publicly released Yorùbá ADR models and one open-weight large language model (LLM) on the YAD benchmark under a consistent protocol. Our results demonstrate that Yo-ByT5 matches the performance of the strongest existing model, mT5-base, with a DER of 10.14% and a CER of 3.48%. Furthermore, it exhibits superior text fidelity despite using approximately half the parameter count of mT5-base. We also release our training code and model outputs, as well as call for the development of a larger, purpose-built benchmark for Yorùbá diacritic restoration.