Dataset catalog
Filter by type, access, and pricing. Specs show before you open the product page.
Nemotron-Post-Training-Dataset-v1
NVIDIA 25.6M-row post-training corpus (chat/code/math/stem/tool_calling) — includes a rare 310K genuine tool-calling split with real tool_calls schemas; CC-BY-4.0.
OpenThoughts-114k
114K verified DeepSeek-R1 reasoning traces over math, science, code and puzzles; Open Thoughts, Apache-2.0.
OpenMathReasoning
~5.68M math solutions over 306k unique AoPS problems, split into CoT, tool-integrated reasoning (Python code) and GenSelect.
Llama-Nemotron-Post-Training-Dataset (v1.1)
NVIDIA post-training corpus for Llama-Nemotron models spanning math, code, science, instruction-following, chat and safety; CC-BY-4.0.
OpenCodeReasoning (OCR-1)
735K competitive-programming reasoning samples (Python) with R1-generated chain-of-thought over 28,319 unique questions; NVIDIA, CC-BY-4.0.
SmolTalk
Hugging Face TB's ~1.04M-row synthetic SFT mixture (the 'all' config) used to train the SmolLM2 instruct models.
Bespoke-Stratos-17k
16.7K reasoning traces (~10K math, ~5K code, ~1K science/puzzle) distilled from DeepSeek-R1 via the Sky-T1 pipeline; Bespoke Labs, Apache-2.0.
Tulu 3 SFT Mixture
Allen AI's 939k-row supervised fine-tuning mixture combining public and synthetic instruction data used to train the Tulu 3 models.
Orca AgentInstruct 1M v1
1,046,410 synthetic instruction/response conversations generated by Microsoft's AgentInstruct agentic pipeline across 15 task-type splits.
OpenMathInstruct-2
14M math problem-solution pairs (GSM8K/MATH augmentation) generated by Llama-3.1-405B-Instruct; text chain-of-thought solutions.
Infinity-Instruct
BAAI's large-scale open SFT collection; multi-config (7M chat, 3M foundational, plus dated Gen sets). Gated on Hugging Face.
WildChat-1M
Real user–ChatGPT (GPT-3.5/GPT-4) conversations collected by Ai2 with metadata and moderation labels; current train split ~838K conversations.
No Robots
10,000 human-written instruction-and-demonstration pairs across 10 task categories, created by skilled human annotators (no model-generated data) for supervised fine-tuning.
CodeFeedback-Filtered-Instruction
156.5K high-quality single-turn code instructions filtered (complexity 4-5 via Qwen-72B-Chat) from four open code-instruction sources; M-A-P.
orca-math-word-problems-200k
200k grade-school math word problems with GPT-4-Turbo-generated worked solutions; English, text explanations only.
Aya Dataset
204,112 human-authored multilingual instruction prompt-completion pairs across 65 languages, curated by native/fluent speakers via Cohere Labs' Aya Annotation Platform for multilingual instruction tuning.
OpenAssistant Conversations v2 (OASST2)
A human-generated, human-annotated corpus of multilingual assistant-style conversation trees released by the OpenAssistant community for supervised fine-tuning and alignment research.
Magicoder OSS-Instruct 75K & Evol-Instruct 110K
ISE-UIUC's paired Magicoder code-instruction datasets: OSS-Instruct (75K, seeded from open-source snippets) and Evol-Instruct (110K, decontaminated evol-codealpaca).
SlimOrca
Open-Orca's ~518k-row curated subset of OpenOrca — GPT-4 FLAN reasoning traces with GPT-4 verification against human annotations.
OpenHermes 2.5
Teknium's ~1M-row compilation of primarily GPT-4-generated instruction, chat, coding and reasoning data in ShareGPT (from/value) format.