Skip to content
All Nigeria resources

Verified Datasets in Nigeria

32 verified datasets for building in Nigeria. Every one has had its documentation, licensing and live status checked, not just listed. Last verified 28 Aug 2026.

Population estimates (GRID3 v3.0)

Datasets

Modelled gridded population raster (100m, GeoTIFF, Aug 2025) covering all of Nigeria, aggregatable to LGA/ward level.

An estimate, not a census, the last census was 2006 and the 2023 census did not happen.

Docs live
Population
Verified Aug 2026Free

States/LGAs JSON

Datasets

Zero-dependency JavaScript package (naija-state-local-government) providing Nigerian states, LGAs and senatorial districts; some versions add wards and polling units.

Docs live
Geographic
Verified Jun 2026Free (npm)

Ward Boundaries (GRID3)

Datasets

Ward boundary polygons in GeoJSON/SHP, no auth.

v1.0 covers all states (approximate); high-accuracy v2.0 covers only 15 of 36 states.

Docs live
Geographic
Verified Aug 2026Free

WorldPop Nigeria Population Counts 2020

Datasets

Gridded population count raster for Nigeria at ~100m (3 arc-second) resolution, 2020, in GeoTIFF, UN-adjusted and constrained to building footprints. Single-file download ~58.67 MB from WorldPop (University of Southampton), CC BY 4.0.

Docs live
Geographic
Verified Aug 2026Free / CC-BY 4.0

Yoruba Speech-Text Parallel Corpus

Datasets

Large Yoruba parallel speech-text corpus of 1,647,022 audio-text pairs (~21.5 GB, WAV) aligned with the MMS-300M Forced Aligner for ASR and TTS, with clips of 0.04-12 seconds.

Docs live
Speech
Verified Aug 2026Free / CC-BY 4.0

AfroBench

Datasets

[ingest-scout] Large-scale benchmark from McGill-NLP for evaluating large language models across African languages; covers diverse NLP tasks and language families. Active GitHub repository under the same McGill-NLP lab that produced AfriqueLLM. Complements IrokoBench (already in registry) but targets broader language coverage and LLM evaluation. Confirmed live on GitHub under McGill-NLP organisation.

Imported from a community submission, needs review.

Docs live
Verified Aug 2026

HarvestStat Africa

Datasets

[ingest-scout] Open-access harmonized subnational crop statistics for 33 Sub-Saharan African countries from 1980–2022, containing 574,204 records across 94 crop types at Admin-1 and Admin-2 levels (area, production, yield). Published in Scientific Data 2025 (doi:10.1038/s41597-025-05001-z), GitHub repo confirmed live at v1.2. Critical food-security and agriculture dataset not yet in the registry.

Imported from a community submission, needs review.

Docs live
Verified Aug 2026

WURA

Datasets

[ingest-scout] 49 GB document-level pretraining corpus spanning 16 African languages (Yoruba, Amharic, Egyptian Arabic, Hausa, and others) plus English/French/Arabic/Portuguese, built by auditing mC4 and crawling verified news sources; fetched HuggingFace page and confirmed 239 monthly downloads, Apache 2.0 license, collection updated June 2025, and use as the training corpus for AfriTeVa V2. High-quality complement to Wycord's existing NLP datasets.

Imported from a community submission, needs review.

Docs live
Verified Aug 2026