Learn Before
HotpotQA Multi-Hop QA Benchmark
HotpotQA is a large-scale multi-hop question answering dataset introduced by Yang et al. at EMNLP 2018. It contains approximately 113{,}000 English question-answer pairs built over Wikipedia, where each question is constructed to require reasoning over two or more supporting paragraphs rather than a single passage. Every question is annotated at the sentence level with supporting facts that justify the answer, enabling joint evaluation of answer correctness and evidence selection. The benchmark defines two standard settings: a distractor setting, in which a system must answer each question given paragraphs ( gold plus lexically related distractors), and a fullwiki open-domain setting, in which the supporting paragraphs must be retrieved from the entire Wikipedia corpus. The release also includes comparison questions that contrast attributes of two entities. Standard metrics are Exact Match and F1 for both the predicted answer and the predicted supporting-fact set, with retrieval components typically scored by Recall@ over the fullwiki candidate pool.
0
1
Tags
Science
Related
Disciplinary Research
Legal Research
Historical Research
Scientific Research
Research References
Research Methods
Research Philosophy
Research Center
Empirical Research Report
Reference: What Should I Learn First: Introducing LectureBank for NLP Education and Prerequisite Chain Learning
Reference: R-VGAE: Relational-variational Graph Autoencoder for Unsupervised Prerequisite Chain Learning
Reference: ojs.aaai.org
Reference: Prerequisite Relation Learning for Concepts in MOOCs
Reference: Course Prerequisite Relation (MOOC prerequisite dataset release page)
Reference: MOOCCube: A Large-scale Data Repository for NLP Applications in MOOCs
Reference: QASC: A Dataset for Question Answering via Sentence Composition
Reference: QASC: A Dataset for Question Answering via Sentence Composition (arXiv preprint)
Reference: ColBERTv2: Effective and Efficient Retrieval via Lightweight Late Interaction
Reference: ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT
Reference: arxiv.org
Reference: arxiv.org
Reference: Introduction to Information Retrieval
Reference: Evaluation measures (information retrieval)
Reference: REPLUG: Retrieval-Augmented Black-Box Language Models
Reference: arxiv.org
Reference: HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering
Reference: HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering (arXiv preprint)
Reference: HotpotQA Official Dataset and Leaderboard
Reference: Dense Passage Retrieval for Open-Domain Question Answering
Reference: Dense Passage Retrieval for Open-Domain Question Answering (arXiv preprint)
LectureBank Dataset
MOOC-CS Prerequisite Benchmark
QASC Question Answering Benchmark
Late-Interaction Neural Retrieval
Recall@k Retrieval Metric
RePlug Retrieval-Augmented Black-Box Language Model
HotpotQA Multi-Hop QA Benchmark
Single-Vector Dense Passage Retrieval
Reference: Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
Reference: How Significant Are the Real Performance Gains? An Unbiased Evaluation Framework for GraphRAG
Reference: arxiv.org
Unbiased GraphRAG Evaluation Framework (Zeng et al., 2025)
Reference: RAG vs. GraphRAG: A Systematic Evaluation and Key Insights
Reference: arxiv.org
RAG vs Graph-RAG Controlled Comparison (Han et al., 2025)
Reference: Controlled Retrieval-augmented Context Evaluation for Long-form RAG
Reference: Controlled Retrieval-augmented Context Evaluation for Long-form RAG (ACL Anthology)
CRUX Controlled RAG Context Evaluation (Ju et al., 2025)
Reference: Anytime Heuristic Search
Reference: The Anatomy of a Large-Scale Hypertextual Web Search Engine
Learn After
Induced HotpotQA Slice as Auxiliary Stress Test in Supplement
HotpotQA FullWiki-1k Boundary Probe: Flat 93.4, Hierarchical 92.9, Adaptive 94.0 R@10
HotpotQA External-Validity Probe: Adaptive Depth Does Not Transfer to a Denser Non-Prerequisite Graph (FullWiki-1k: Flat 93.4 / Hier 92.9 / Adaptive 94.0 R@10)