Learn Before
Auditable-Artifact Reporting Standards: Data Statements, Datasheets, Model Cards, Reproducibility Checklists
Auditable-artifact reporting standards are a tradition in machine learning, NLP, and adjacent fields that attaches standardized disclosure documents to the artifacts a research community produces — datasets, trained models, and paper submissions — so that downstream users can independently audit how each artifact was made, what it contains, and what claims it can support. Four founding works define the tradition: (i) data statements (Bender and Friedman, TACL 2018), structured documents attached to NLP datasets that disclose curation rationale, language variety, and speaker and annotator demographics; (ii) datasheets for datasets (Gebru et al., CACM 2021), a 57-question form covering motivation, composition, collection, preprocessing, uses, distribution, and maintenance of any dataset; (iii) model cards for model reporting (Mitchell et al., FAT* 2019), short documents accompanying trained models that report intended use, factors, evaluation and training data, subgroup-disaggregated metrics, and ethical considerations; and (iv) the NeurIPS 2019 reproducibility program and ML Reproducibility Checklist (Pineau et al., JMLR 2021), which attaches an explicit reproducibility checklist to conference submissions and pairs it with a code-submission policy. The shared discipline across all four is that each artifact ships with the documentation a third party needs to audit it, rather than relying on prose buried inside the paper. This tradition is methodologically distinct from the IR-evaluation reproducibility line (reference implementations, statistical controls, split-design hygiene): it is about per-artifact disclosure rather than benchmark-level evaluation variance.
0
1
Tags
Science
Auditable Strict-Parity Evaluation of Prerequisite-Graph Retrieval for RAG under Leakage Controls
Related
Disciplinary Research
Legal Research
Historical Research
Scientific Research
Research References
Research Methods
Research Philosophy
Research Center
Empirical Research Report
Reference: What Should I Learn First: Introducing LectureBank for NLP Education and Prerequisite Chain Learning
Reference: R-VGAE: Relational-variational Graph Autoencoder for Unsupervised Prerequisite Chain Learning
Reference: ojs.aaai.org
Reference: Prerequisite Relation Learning for Concepts in MOOCs
Reference: Course Prerequisite Relation (MOOC prerequisite dataset release page)
Reference: MOOCCube: A Large-scale Data Repository for NLP Applications in MOOCs
Reference: QASC: A Dataset for Question Answering via Sentence Composition
Reference: QASC: A Dataset for Question Answering via Sentence Composition (arXiv preprint)
Reference: ColBERTv2: Effective and Efficient Retrieval via Lightweight Late Interaction
Reference: ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT
Reference: arxiv.org
Reference: arxiv.org
Reference: Introduction to Information Retrieval
Reference: Evaluation measures (information retrieval)
Reference: REPLUG: Retrieval-Augmented Black-Box Language Models
Reference: arxiv.org
Reference: HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering
Reference: HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering (arXiv preprint)
Reference: HotpotQA Official Dataset and Leaderboard
Reference: Dense Passage Retrieval for Open-Domain Question Answering
Reference: Dense Passage Retrieval for Open-Domain Question Answering (arXiv preprint)
LectureBank Dataset
MOOC-CS Prerequisite Benchmark
QASC Question Answering Benchmark
Late-Interaction Neural Retrieval
Recall@k Retrieval Metric
RePlug Retrieval-Augmented Black-Box Language Model
HotpotQA Multi-Hop QA Benchmark
Single-Vector Dense Passage Retrieval
Reference: Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
Reference: How Significant Are the Real Performance Gains? An Unbiased Evaluation Framework for GraphRAG
Reference: arxiv.org
Unbiased GraphRAG Evaluation Framework (Zeng et al., 2025)
Reference: RAG vs. GraphRAG: A Systematic Evaluation and Key Insights
Reference: arxiv.org
RAG vs Graph-RAG Controlled Comparison (Han et al., 2025)
Reference: Controlled Retrieval-augmented Context Evaluation for Long-form RAG
Reference: Controlled Retrieval-augmented Context Evaluation for Long-form RAG (ACL Anthology)
CRUX Controlled RAG Context Evaluation (Ju et al., 2025)
Reference: Anytime Heuristic Search
Reference: The Anatomy of a Large-Scale Hypertextual Web Search Engine
Learn After
Claim-Level Traceability in Graph-RAG Evaluation
Assistive-Tool-Only Deployment Recommendation for Hierarchical Prerequisite Graph RAG
Dataset License and Attribution Disclosure via Released Data Cards
Non-Privacy Ethical Dimensions Deferred to Released Data Cards (Licenses, Redistribution, Instructor Attribution)