Verified AI Resources in Nigeria
17 verified ai resources for building in Nigeria. Every one has had its documentation, licensing and live status checked, not just listed. Last verified 25 Aug 2026.
NaijaML
Production-ready NLP library (v0.2.1, Mar 2026) covering all 4 Nigerian languages, running in 4GB RAM with no GPU, with PII masking for NIN/phone. The most practical Nigerian AI resource for developers.
YarnGPT2b
Nigerian-accented English TTS and ASR model (Jan 2025) with 11 voices, trained on Nollywood and podcast audio.
Africa GPU Hub
First Nigerian GPU rental marketplace (Udutech, Lagos), offering GPU cloud compute at under $1/hr.
Awarri
Awarri is a Lagos-based AI and robotics company that built N-ATLaS, Nigeria's first government-backed open-source multilingual LLM (Llama-3 8B fine-tuned on ~392M tokens of English, Hausa, Igbo and Yoruba), in partnership with NCAIR and the Federal Ministry of Communications, Innovation & Digital Economy. It also operates the LangEasy data-collection platform.
EqualyzAI
EqualyzAI is a voice-first agentic AI company building speech recognition, text-to-speech and voice agents for African languages and dialects (Yoruba, Igbo, Hausa, Pidgin with code-switching), with products including VoiceMaker, VoiceAgent and VoiceBridge plus datasets/APIs. Operates from Lagos, Nigeria and Washington DC.
Intron Health
Intron Health is a Nigerian voice-AI company whose Sahara-v2 models deliver clinical/medical speech recognition (and TTS) optimized natively for African accents and dialects, trained on Africa's largest clinical speech database (millions of clips across 200+ accents). Serves healthcare, call-centre, legal and biometrics use cases via STT/TTS/voice-bot APIs.
Lanfrica
Lanfrica is a catalog/registry mapping African language resources (datasets, models, papers and policies) via its African AI Atlas, positioning itself as 'the evidence layer for African AI.' Built by Lanfrica Labs with partners including Meta, Mozilla, Masakhane and Lacuna Fund.
LangEasy
LangEasy is Awarri's crowdsourced data-collection platform (smartphone app) that lets anyone contribute voice and text in Nigerian languages (Yoruba, Hausa, Igbo, Ibibio, Pidgin and accented English) to build the training data behind Nigeria's national LLM, N-ATLaS.
N-ATLaS
Nigeria's first government-backed multilingual LLM (Sep 2025): a Llama-3 8B fine-tuned on 400M+ tokens across 4 Nigerian languages. Produced by NCAIR/NITDA and Awarri.
Prembly FraudLens
Open-source fraud-intelligence dataset for Nigerian fintechs (Mar 2026), the first of its kind. Relevant to payment fraud, KYC fraud and account-takeover detection.
Spitch
Spitch is a Lagos-based voice-AI company (founded Oct 2024) providing ASR (speech-to-text), TTS (text-to-speech) and translation APIs/SDKs for Nigerian languages (Yoruba, Igbo, Hausa, Nigerian-accented English, plus Amharic) so teams can add local-language voice to call centres, media and learning tools.
AfriqueLLM Collection
[ingest-scout] Suite of 10 open large language models from McGill-NLP adapted for 20–50 African languages via continued pre-training on 26–35B tokens; includes AfriqueLlama-8B, AfriqueQwen variants (4B–14B), and AfriqueGemma variants (4B–12B). Backed by ACL 2026 research paper (arXiv:2601.06395); models updated June 2026 with active downloads. Covers languages not served by registry entries (Amharic, Tigrinya, Oromo, Wolof, Bambara, Moroccan/Egyptian/Tunisian Arabic, Malagasy, Afrikaans, etc.).
Imported from a community submission, needs review.
DigitalUmuganda Mbaza-ASR-Afrivoice-660h
[ingest-scout] Rwanda-based NLP organization Digital Umuganda's Conformer-CTC ASR model trained on 660 hours of Kinyarwanda speech (Afrivoice dataset); last updated December 2 2025, 52 monthly downloads, CC-BY-4.0 license, built on NVIDIA NeMo. Distinct from the already-inventoried mbaza-whisper-small-kinyarwanda — this is a larger NeMo-based model; fills Rwanda AI infrastructure representation.
Imported from a community submission, needs review.
Ethio-ASR Amharic
[ingest-scout] Ethiopian language ASR model fine-tuned from Facebook w2v-bert-2.0 (600M params) on Amharic speech, achieving 22.37 WER on the WAXAL Amharic test set; paper published March 2026 (arXiv:2603.23654), dataset updated June 2026, 1.03k monthly downloads, CC-BY-4.0. Developed by the Ethio-ASR research team — a verified live model filling the Ethiopia AI gap in the registry.
Imported from a community submission, needs review.
NaijaVoices Dataset
NaijaVoices is a large-scale speech dataset of about 1,800 hours from over 5,000 speakers with expert-curated transcripts in Igbo, Hausa and Yoruba, roughly 600 hours per language. It is designed for building ASR and speech AI for Nigerian languages and improves Whisper and MMS fine-tuning performance. It is available on HuggingFace behind a free registration and powers models like AfriHuBERT and SBPN.
SabiYarn-125M
SabiYarn-125M is a 125M-parameter decoder-only foundation model pretrained on Nigerian-language text, the first in the SabiYarn series. It supports English, Yoruba, Hausa, Igbo and Nigerian Pidgin plus Fulfulde, Efik and Urhobo, with fine-tuned variants for translation, NER, sentiment and diacritization. It was built by Aletheia.ai Research Lab and presented at the AfricaNLP 2025 workshop.
Simba African Speech AI Suite
[ingest-scout] Open-source African speech AI ecosystem from UBC's Deep Learning & NLP Lab: 5 ASR models covering 43 languages, 7 TTS models covering 7 languages, spoken language identification for 49 languages, and SimbaBench evaluation dataset. Introduced at EMNLP 2025 (paper: 'Voice of a Continent'); models hosted on HuggingFace under UBC-NLP. Confirmed live at github.com/UBC-NLP/simba with CC-BY-4.0 license. Distinct from african-whisper (ASR only), bibletts (TTS only), and spitch (commercial) already in registry.
Imported from a community submission, needs review.