Verified Datasets in Nigeria
32 verified datasets for building in Nigeria. Every one has had its documentation, licensing and live status checked, not just listed. Last verified 28 Aug 2026.
Population estimates (GRID3 v3.0)
Modelled gridded population raster (100m, GeoTIFF, Aug 2025) covering all of Nigeria, aggregatable to LGA/ward level.
An estimate, not a census, the last census was 2006 and the 2023 census did not happen.
States/LGAs JSON
Zero-dependency JavaScript package (naija-state-local-government) providing Nigerian states, LGAs and senatorial districts; some versions add wards and polling units.
Ward Boundaries (GRID3)
Ward boundary polygons in GeoJSON/SHP, no auth.
v1.0 covers all states (approximate); high-accuracy v2.0 covers only 15 of 36 states.
WorldPop Nigeria Population Counts 2020
Gridded population count raster for Nigeria at ~100m (3 arc-second) resolution, 2020, in GeoTIFF, UN-adjusted and constrained to building footprints. Single-file download ~58.67 MB from WorldPop (University of Southampton), CC BY 4.0.
Yoruba Speech-Text Parallel Corpus
Large Yoruba parallel speech-text corpus of 1,647,022 audio-text pairs (~21.5 GB, WAV) aligned with the MMS-300M Forced Aligner for ASR and TTS, with clips of 0.04-12 seconds.
AfroBench
[ingest-scout] Large-scale benchmark from McGill-NLP for evaluating large language models across African languages; covers diverse NLP tasks and language families. Active GitHub repository under the same McGill-NLP lab that produced AfriqueLLM. Complements IrokoBench (already in registry) but targets broader language coverage and LLM evaluation. Confirmed live on GitHub under McGill-NLP organisation.
Imported from a community submission, needs review.
HarvestStat Africa
[ingest-scout] Open-access harmonized subnational crop statistics for 33 Sub-Saharan African countries from 1980–2022, containing 574,204 records across 94 crop types at Admin-1 and Admin-2 levels (area, production, yield). Published in Scientific Data 2025 (doi:10.1038/s41597-025-05001-z), GitHub repo confirmed live at v1.2. Critical food-security and agriculture dataset not yet in the registry.
Imported from a community submission, needs review.
WURA
[ingest-scout] 49 GB document-level pretraining corpus spanning 16 African languages (Yoruba, Amharic, Egyptian Arabic, Hausa, and others) plus English/French/Arabic/Portuguese, built by auditing mC4 and crawling verified news sources; fetched HuggingFace page and confirmed 239 monthly downloads, Apache 2.0 license, collection updated June 2025, and use as the training corpus for AfriTeVa V2. High-quality complement to Wycord's existing NLP datasets.
Imported from a community submission, needs review.