NaijaVoices Dataset
AI Resources
Speech Dataset
Docs live
Approval required
NaijaVoices is a large-scale speech dataset of about 1,800 hours from over 5,000 speakers with expert-curated transcripts in Igbo, Hausa and Yoruba, roughly 600 hours per language. It is designed for building ASR and speech AI for Nigerian languages and improves Whisper and MMS fine-tuning performance. It is available on HuggingFace behind a free registration and powers models like AfriHuBERT and SBPN.
- Category
- AI Resources
- Pricing
- Free for non-commercial research (registration required)
- Country
- 🇳🇬 Nigeria
- Last verified
- 25 Aug 2026
{ "name": "NaijaVoices Dataset", "slug": "naijavoices-dataset", "category": "AI", "country": "NG", "docs_status": "LIVE", "licensing_required": "APPROVAL", "verified": false, "last_verified": "2026-08-25", "website": "https://huggingface.co/datasets/naijavoices/naijavoices-dataset", "documentation_url": "https://arxiv.org/abs/2505.20564"}Verification history
- 25 Aug 2026 · live
- 22 Aug 2026 · live
- 19 Aug 2026 · live
- 16 Aug 2026 · live
- 13 Aug 2026 · live
- 10 Aug 2026 · live
- 7 Aug 2026 · live
- 4 Aug 2026 · live
Automated checks run every few days. See all recent status changes
Tags
yoruba
hausa
igbo
nigerian-languages
speech-dataset