AfriQA: Cross-lingual Open-Retrieval Question Answering for African Languages
VerifiedAfriQA is the first cross-lingual open-retrieval question answering benchmark for African languages, with more than 12,000 XOR-QA examples across 10 African languages. The paper shows that current automatic translation and multilingual retrieval methods perform poorly for these languages, where in-language digital content is scarce.
- Category
- Research
- Pricing
- Free / open
- Country
- 🌍 Pan-African
- Last verified
- 25 Aug 2026
{ "name": "AfriQA: Cross-lingual Open-Retrieval Question Answering for African Languages", "slug": "afriqa-cross-lingual-open-retrieval-question-answering-for-african-languages", "category": "RESEARCH", "country": "Pan-African", "docs_status": "LIVE", "licensing_required": "NONE", "verified": true, "last_verified": "2026-08-25", "website": "https://arxiv.org/abs/2305.06897", "documentation_url": null}Verification history
- 25 Aug 2026 · live
- 22 Aug 2026 · live
- 19 Aug 2026 · live
- 16 Aug 2026 · live
- 13 Aug 2026 · live
- 10 Aug 2026 · live
- 7 Aug 2026 · live
- 4 Aug 2026 · live
Automated checks run every few days. See all recent status changes
Tags
Compare AfriQA: Cross-lingual Open-Retrieval Question Answering for African Languages
Side-by-side, verified specs against its closest nlp benchmark alternatives.
Related in Research
AfriSenti: A Twitter Sentiment Analysis Benchmark for African Languages
AfriSenti is a sentiment analysis benchmark of more than 110,000 tweets in 14 African languages spanning four language families, annotated by native speakers. It underpinned SemEval-2023 Task 12, a shared task that attracted more than 200 participants, and documents data collection, annotation and baseline methods for low-resource languages.
AfriHate: A Multilingual Collection of Hate Speech and Abusive Language Datasets for African Languages
AfriHate is a multilingual benchmark of hate speech and abusive language datasets covering 15 African languages, annotated by native speakers. The paper contributes classification baselines and hate speech and offensive language lexicons, and analyses why keyword-based moderation fails for low-resource African languages. It was released on arXiv in January 2025.
MasakhaNER 2.0: Africa-centric Transfer Learning for Named Entity Recognition
MasakhaNER 2.0 introduces the largest human-annotated named entity recognition dataset for 20 African languages and studies Africa-centric cross-lingual transfer learning. The paper reports that choosing the best transfer language improves zero-shot F1 by an average of 14 points across the 20 languages compared with transferring from English.