Nigeria
32 verified resources in Datasets for building in Nigeria.
- Resources
- 208
- Docs live
- 95%
- Companies (HQ)
- 75
- Last verified
- 25 Aug 2026
Population estimates (GRID3 v3.0)
Modelled gridded population raster (100m, GeoTIFF, Aug 2025) covering all of Nigeria, aggregatable to LGA/ward level.
An estimate, not a census, the last census was 2006 and the 2023 census did not happen.
States/LGAs JSON
Zero-dependency JavaScript package (naija-state-local-government) providing Nigerian states, LGAs and senatorial districts; some versions add wards and polling units.
Ward Boundaries (GRID3)
Ward boundary polygons in GeoJSON/SHP, no auth.
v1.0 covers all states (approximate); high-accuracy v2.0 covers only 15 of 36 states.
WorldPop Nigeria Population Counts 2020
Gridded population count raster for Nigeria at ~100m (3 arc-second) resolution, 2020, in GeoTIFF, UN-adjusted and constrained to building footprints. Single-file download ~58.67 MB from WorldPop (University of Southampton), CC BY 4.0.
Yoruba Speech-Text Parallel Corpus
Large Yoruba parallel speech-text corpus of 1,647,022 audio-text pairs (~21.5 GB, WAV) aligned with the MMS-300M Forced Aligner for ASR and TTS, with clips of 0.04-12 seconds.
AfroBench
[ingest-scout] Large-scale benchmark from McGill-NLP for evaluating large language models across African languages; covers diverse NLP tasks and language families. Active GitHub repository under the same McGill-NLP lab that produced AfriqueLLM. Complements IrokoBench (already in registry) but targets broader language coverage and LLM evaluation. Confirmed live on GitHub under McGill-NLP organisation.
Imported from a community submission, needs review.
HarvestStat Africa
[ingest-scout] Open-access harmonized subnational crop statistics for 33 Sub-Saharan African countries from 1980–2022, containing 574,204 records across 94 crop types at Admin-1 and Admin-2 levels (area, production, yield). Published in Scientific Data 2025 (doi:10.1038/s41597-025-05001-z), GitHub repo confirmed live at v1.2. Critical food-security and agriculture dataset not yet in the registry.
Imported from a community submission, needs review.
WURA
[ingest-scout] 49 GB document-level pretraining corpus spanning 16 African languages (Yoruba, Amharic, Egyptian Arabic, Hausa, and others) plus English/French/Arabic/Portuguese, built by auditing mC4 and crawling verified news sources; fetched HuggingFace page and confirmed 239 monthly downloads, Apache 2.0 license, collection updated June 2025, and use as the training corpus for AfriTeVa V2. High-quality complement to Wycord's existing NLP datasets.
Imported from a community submission, needs review.