Convex Markets / Datasets / Reasoning QA

Dataset catalog

Filter by type, access, and pricing. Specs show before you open the product page.

9 results
M
External · RQA-5009

MMLU-Pro

A harder, reasoning-focused successor to MMLU: 12,032 multiple-choice questions across 14 subjects with up to 10 options each, scored by exact match on the gold answer letter.

Type Reasoning QAVolume 12032 multiple-choice questions (test split; plus 70 validation CoT few-shot examples)Format ParquetAccess PUBLIC LICENSE
Free
open dataset (HF, MIT)
View
G
External · RQA-5004

GPQA

GPQA is a gated benchmark of 448 expert-written, 'Google-proof' graduate-level multiple-choice questions in biology, physics, and chemistry, built for reasoning evaluation and scalable-oversight research.

Type Reasoning QAVolume 448 multiple-choice questions (GPQA Main set; also GPQA Diamond=198, GPQA Extended=546)Format CSVAccess GATED
Free
open dataset (HF, gated)
View
BH
External · RQA-5008

BIG-Bench Hard (BBH)

A 6,511-example reasoning benchmark of 23 hard BIG-Bench tasks (input/target pairs) used to test chain-of-thought prompting, graded by exact match on the gold answer.

Type Reasoning QAVolume 6,511 examples (test split; across 27 task configs / 23 BBH tasks)Format parquet (Hugging Face); per-task JSON in the source GitHub repoAccess PUBLIC LICENSE
Free
open dataset (HF)
View
M(
External · RQA-5002

MATH (Hendrycks)

12,500 competition mathematics problems (7,500 train / 5,000 test) with full step-by-step LaTeX solutions whose final answer is wrapped in \boxed{}.

Type Reasoning QAVolume 12,500 problems (7,500 train / 5,000 test)Format Parquet (Hugging Face mirror, 7 subject configs); original repository ships per-problem JSON files with fields problem/level/type/solutionAccess PUBLIC LICENSE
Free
open dataset (HF)
View
AA
External · RQA-5003

AI2 ARC

A dataset of 7,787 genuine grade-school-level multiple-choice science questions, split into a harder Challenge Set and an Easy Set, for evaluating advanced question answering and reasoning.

Type Reasoning QAVolume 7787 questionsFormat parquetAccess PUBLIC LICENSE
Free
open dataset (HF)
View
G
External · RQA-5001

GSM8K

GSM8K is a dataset of 8.5K human-written grade-school math word problems, each paired with a multi-step natural-language solution that ends in a single final numeric answer marked by '####'.

Type Reasoning QAVolume 8792 math word problems (7473 train / 1319 test, main config)Format parquet (Hugging Face); JSONL in the source GitHub repoAccess PUBLIC LICENSE
Free
open dataset (HF)
View
S
External · RQA-5005

StrategyQA

A yes/no question-answering benchmark whose questions require implicit multi-step (strategy) reasoning, graded by boolean exact-match against a gold true/false answer.

Type Reasoning QAVolume 2,290 questions (1,603 train + 687 test)Format parquetAccess PUBLIC LICENSE
Free
open-source license (MIT)
View
D
External · RQA-5007

DROP

A crowdsourced reading-comprehension benchmark whose questions require discrete reasoning (addition, counting, sorting, comparison) over Wikipedia-derived paragraphs.

Type Reasoning QAVolume 86,935 questions (77,400 train + 9,535 validation)Format Parquet (Hugging Face); originally distributed as JSONAccess PUBLIC LICENSE
Free
open dataset (HF)
View
H
External · RQA-5006

HotpotQA

A Wikipedia-based multi-hop question-answering dataset of 113,000+ QA pairs that require reasoning across multiple supporting documents and provide sentence-level supporting facts.

Type Reasoning QAVolume 113,000+ multi-hop question-answer pairsFormat Parquet (Hugging Face); original distribution JSONAccess PUBLIC LICENSE
Free
open dataset (HF)
View