AfriHate Hate Speech Datasets
Multilingual collection of hate speech and abusive language datasets covering 15 African languages, built from tweets annotated by native speakers. Each instance carries labels from 3 to 4 annotators with anonymous annotator IDs, downloadable on HuggingFace. Published at NAACL 2025.
- Category
- Datasets
- Pricing
- Free / open
- Country
- 🌍 Pan-African
- Last verified
- 25 Aug 2026
{ "name": "AfriHate Hate Speech Datasets", "slug": "afrihate-hate-speech-datasets", "category": "DATASET", "country": "Pan-African", "docs_status": "LIVE", "licensing_required": "NONE", "verified": false, "last_verified": "2026-08-25", "website": "https://huggingface.co/datasets/afrihate/afrihate", "documentation_url": "https://aclanthology.org/2025.naacl-long.92/"}Verification history
- 25 Aug 2026 · live
- 22 Aug 2026 · live
- 19 Aug 2026 · live
- 16 Aug 2026 · live
- 13 Aug 2026 · live
- 10 Aug 2026 · live
- 7 Aug 2026 · live
- 4 Aug 2026 · live
Automated checks run every few days. See all recent status changes
Tags
Compare AfriHate Hate Speech Datasets
Side-by-side, verified specs against its closest language / nlp alternatives.
Related in Datasets
MasakhaNER 2.0
Largest high-quality named-entity-recognition corpus for 20 African languages (incl. Nigerian Pidgin, Hausa, Igbo, Yoruba) with PER/ORG/LOC/DATE tags over news-domain text, totaling ~152,786 rows. Built by the Masakhane community.
MasakhaNEWS
News-topic-classification dataset for 16 widely spoken African languages (incl. Hausa, Igbo, Yoruba, Nigerian Pidgin), ~31,088 rows in CSV/Parquet with train/val/test splits across seven topic categories. Built by the Masakhane community.
Hausa Visual Genome (HausaVG)
Multimodal Hausa-English dataset of 32,923 images with paired English/Hausa region descriptions (train/dev/test/challenge splits), post-edited by HausaNLP and Bayero University Kano translators for English-to-Hausa machine translation and image description.