Skip to content
Datasets

MasakhaNER 2.0

Verified
Datasets
Language / NLP
Docs live

Largest high-quality named-entity-recognition corpus for 20 African languages (incl. Nigerian Pidgin, Hausa, Igbo, Yoruba) with PER/ORG/LOC/DATE tags over news-domain text, totaling ~152,786 rows. Built by the Masakhane community.

Category
Datasets
Pricing
Free / CC-BY-NC 4.0
Country
🌍 Pan-African
Last verified
25 Aug 2026
{
"name": "MasakhaNER 2.0",
"slug": "masakhaner-20",
"category": "DATASET",
"country": "Pan-African",
"docs_status": "LIVE",
"licensing_required": "NONE",
"verified": true,
"last_verified": "2026-08-25",
"website": "https://huggingface.co/datasets/masakhane/masakhaner2",
"documentation_url": "https://huggingface.co/datasets/masakhane/masakhaner2"
}
get_resource("masakhaner-20") — via the Wycord MCP server

Verification history

  • 25 Aug 2026 · live
  • 22 Aug 2026 · live
  • 19 Aug 2026 · live
  • 16 Aug 2026 · live
  • 13 Aug 2026 · live
  • 10 Aug 2026 · live
  • 7 Aug 2026 · live
  • 4 Aug 2026 · live

Automated checks run every few days. See all recent status changes

Tags

nlp
ner
named-entity-recognition
african-languages
token-classification

Compare MasakhaNER 2.0

Side-by-side, verified specs against its closest language / nlp alternatives.

Related in Datasets