Dataset catalog
Filter by type, access, and pricing. Specs show before you open the product page.
VisualWebArena
Multimodal, visually grounded web-agent benchmark: 910 tasks over self-hosted Classifieds, Shopping and Reddit sites.
WorkArena / WorkArena++ (ServiceNow)
Enterprise knowledge-work benchmark on live ServiceNow instances: 33 L1 atomic tasks (19,912 instances) plus 682 L2/L3 compositional tasks.
MineRL Human Demonstrations
A large-scale dataset of over 60 million automatically-annotated human state-action pairs recorded while people play Minecraft across a set of related item-acquisition and navigation tasks.
PKU-SafeRLHF
A large human-annotated safety-preference dataset where each question has two model responses ranked separately for helpfulness and harmlessness, plus per-response safety meta-labels across 19 harm categories.
Infinity-Instruct
BAAI's large-scale open SFT collection; multi-config (7M chat, 3M foundational, plus dated Gen sets). Gated on Hugging Face.
tau-bench (τ-bench)
Tool-agent-user benchmark of 165 customer-service tasks (retail 115, airline 50) with policy-following and pass@k evaluation.
WildChat-1M
Real user–ChatGPT (GPT-3.5/GPT-4) conversations collected by Ai2 with metadata and moderation labels; current train split ~838K conversations.
WebArena
Realistic self-hosted web environment: 812 long-horizon tasks over shopping, forum, GitLab, CMS and maps with execution-based evaluation.
No Robots
10,000 human-written instruction-and-demonstration pairs across 10 task categories, created by skilled human annotators (no model-generated data) for supervised fine-tuning.
CodeFeedback-Filtered-Instruction
156.5K high-quality single-turn code instructions filtered (complexity 4-5 via Qwen-72B-Chat) from four open code-instruction sources; M-A-P.
orca-math-word-problems-200k
200k grade-school math word problems with GPT-4-Turbo-generated worked solutions; English, text explanations only.
Aya Dataset
204,112 human-authored multilingual instruction prompt-completion pairs across 65 languages, curated by native/fluent speakers via Cohere Labs' Aya Annotation Platform for multilingual instruction tuning.
WebLINX
Expert demonstrations of conversational, multi-turn website navigation, where a navigator agent must predict the next web action (click/say/load/submit/change) from dialogue history and the page DOM.
UltraFeedback
A large-scale, fine-grained preference dataset of ~64k prompts, each with 4 model completions rated by GPT-4 across four aspects, for training reward and critique models.
OpenAssistant Conversations v2 (OASST2)
A human-generated, human-annotated corpus of multilingual assistant-style conversation trees released by the OpenAssistant community for supervised fine-tuning and alignment research.
AI2 ARC
A dataset of 7,787 genuine grade-school-level multiple-choice science questions, split into a harder Challenge Set and an Easy Set, for evaluating advanced question answering and reasoning.
Mind2Web
2,350 human-demonstrated web-navigation task trajectories across 137 real websites, each pairing a natural-language instruction with a step-by-step sequence of DOM-grounded CLICK/TYPE/SELECT actions and full HTML snapshots.
Magicoder OSS-Instruct 75K & Evol-Instruct 110K
ISE-UIUC's paired Magicoder code-instruction datasets: OSS-Instruct (75K, seeded from open-source snippets) and Evol-Instruct (110K, decontaminated evol-codealpaca).
Intel Orca DPO Pairs
A ~12.9K-example preference dataset in Direct Preference Optimization (DPO) format, derived from Open-Orca/OpenOrca, pairing a 'chosen' and 'rejected' response for each instruction prompt.
OpenHermes 2.5
Teknium's ~1M-row compilation of primarily GPT-4-generated instruction, chat, coding and reasoning data in ShareGPT (from/value) format.