Dataset catalog
Filter by type, access, and pricing. Specs show before you open the product page.
MMLU-Pro
A harder, reasoning-focused successor to MMLU: 12,032 multiple-choice questions across 14 subjects with up to 10 options each, scored by exact match on the gold answer letter.
GPQA
GPQA is a gated benchmark of 448 expert-written, 'Google-proof' graduate-level multiple-choice questions in biology, physics, and chemistry, built for reasoning evaluation and scalable-oversight research.
BIG-Bench Hard (BBH)
A 6,511-example reasoning benchmark of 23 hard BIG-Bench tasks (input/target pairs) used to test chain-of-thought prompting, graded by exact match on the gold answer.
MATH (Hendrycks)
12,500 competition mathematics problems (7,500 train / 5,000 test) with full step-by-step LaTeX solutions whose final answer is wrapped in \boxed{}.
AI2 ARC
A dataset of 7,787 genuine grade-school-level multiple-choice science questions, split into a harder Challenge Set and an Easy Set, for evaluating advanced question answering and reasoning.
GSM8K
GSM8K is a dataset of 8.5K human-written grade-school math word problems, each paired with a multi-step natural-language solution that ends in a single final numeric answer marked by '####'.
StrategyQA
A yes/no question-answering benchmark whose questions require implicit multi-step (strategy) reasoning, graded by boolean exact-match against a gold true/false answer.
DROP
A crowdsourced reading-comprehension benchmark whose questions require discrete reasoning (addition, counting, sorting, comparison) over Wikipedia-derived paragraphs.
HotpotQA
A Wikipedia-based multi-hop question-answering dataset of 113,000+ QA pairs that require reasoning across multiple supporting documents and provide sentence-level supporting facts.