Learn Before
Reproducible IR Benchmarking and Evaluation-Variance Control
Reproducible IR benchmarking and evaluation-variance control is a methodological tradition in information retrieval and adjacent ML evaluation that treats how a result is produced and reported as part of the result itself. Four influential strands define this tradition. (i) Reproducible reference implementations and statistical analyses on shared benchmarks: Kamalloo et al. (SIGIR 2024) release reference dense and sparse retrievers on BEIR, together with effect-size meta-analyses and an official leaderboard, so that compared systems share the same artifacts and statistical conventions. (ii) Improved reporting of experimental results: Dodge et al. (EMNLP 2019) show that test-set numbers alone are insufficient and prescribe reporting validation performance as a function of compute budget, since conclusions about which model is best can flip with hyperparameter search effort. (iii) Split design as a controlled variable: Gorman and Bedrick (ACL 2019) show that single standard train/test splits can produce unstable system rankings, and argue for randomized splits to test whether reported gains are reproducible across resamplings. (iv) Hidden benchmark assumptions: Dehghani et al. (2021) formalize the benchmark lottery, demonstrating that the choice of benchmark tasks (and other unstated reporting choices) can flip the apparent ordering of algorithms. Together these works establish that retrieval and ML claims must be reported with enough artifact, split, and configuration detail for variance and benchmark choices to be audited — the reporting line that downstream claim-level traceability discipline extends.
0
1
Tags
Science
Auditable Strict-Parity Evaluation of Prerequisite-Graph Retrieval for RAG under Leakage Controls
Related
Disciplinary Research
Legal Research
Historical Research
Scientific Research
Research References
Research Methods
Research Philosophy
Research Center
Empirical Research Report
Reference: What Should I Learn First: Introducing LectureBank for NLP Education and Prerequisite Chain Learning
Reference: R-VGAE: Relational-variational Graph Autoencoder for Unsupervised Prerequisite Chain Learning
Reference: ojs.aaai.org
Reference: Prerequisite Relation Learning for Concepts in MOOCs
Reference: Course Prerequisite Relation (MOOC prerequisite dataset release page)
Reference: MOOCCube: A Large-scale Data Repository for NLP Applications in MOOCs
Reference: QASC: A Dataset for Question Answering via Sentence Composition
Reference: QASC: A Dataset for Question Answering via Sentence Composition (arXiv preprint)
Reference: ColBERTv2: Effective and Efficient Retrieval via Lightweight Late Interaction
Reference: ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT
Reference: arxiv.org
Reference: arxiv.org
Reference: Introduction to Information Retrieval
Reference: Evaluation measures (information retrieval)
Reference: REPLUG: Retrieval-Augmented Black-Box Language Models
Reference: arxiv.org
Reference: HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering
Reference: HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering (arXiv preprint)
Reference: HotpotQA Official Dataset and Leaderboard
Reference: Dense Passage Retrieval for Open-Domain Question Answering
Reference: Dense Passage Retrieval for Open-Domain Question Answering (arXiv preprint)
LectureBank Dataset
MOOC-CS Prerequisite Benchmark
QASC Question Answering Benchmark
Late-Interaction Neural Retrieval
Recall@k Retrieval Metric
RePlug Retrieval-Augmented Black-Box Language Model
HotpotQA Multi-Hop QA Benchmark
Single-Vector Dense Passage Retrieval
Reference: Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
Reference: How Significant Are the Real Performance Gains? An Unbiased Evaluation Framework for GraphRAG
Reference: arxiv.org
Unbiased GraphRAG Evaluation Framework (Zeng et al., 2025)
Reference: RAG vs. GraphRAG: A Systematic Evaluation and Key Insights
Reference: arxiv.org
RAG vs Graph-RAG Controlled Comparison (Han et al., 2025)
Reference: Controlled Retrieval-augmented Context Evaluation for Long-form RAG
Reference: Controlled Retrieval-augmented Context Evaluation for Long-form RAG (ACL Anthology)
CRUX Controlled RAG Context Evaluation (Ju et al., 2025)
Reference: Anytime Heuristic Search
Reference: The Anatomy of a Large-Scale Hypertextual Web Search Engine