Convex Markets / Datasets / All departments

Dataset catalog

Filter by type, access, and pricing. Specs show before you open the product page.

110 results
V
External · RLE-1102

VisualWebArena

Multimodal, visually grounded web-agent benchmark: 910 tasks over self-hosted Classifieds, Shopping and Reddit sites.

Type RL environmentsVolume 910 tasks (Classifieds 234, Shopping 466, Reddit 210)Format JSON task configs + Python (gym) + self-hosted Docker sitesAccess PUBLIC LICENSE
Free
open-source license
View
W/
External · RLE-1104

WorkArena / WorkArena++ (ServiceNow)

Enterprise knowledge-work benchmark on live ServiceNow instances: 33 L1 atomic tasks (19,912 instances) plus 682 L2/L3 compositional tasks.

Type RL environmentsVolume 33 L1 atomic tasks (19,912 instances) + 682 WorkArena++ L2/L3 tasksFormat Python package (pip install browsergym-workarena) + live ServiceNow instance + PlaywrightAccess GATED
Free
open-source license
View
MH
External · DEM-6008

MineRL Human Demonstrations

A large-scale dataset of over 60 million automatically-annotated human state-action pairs recorded while people play Minecraft across a set of related item-acquisition and navigation tasks.

Type DemonstrationsVolume 60,000,000+ state-action pairsFormat Per-task compressed archives distributed for download; loaded via the minerl.data API as OpenAI Gym Dict observation/action tuples yielded as (current_state, action, reward, next_state, done) by BufferedBatchIterAccess PUBLIC LICENSE
Free
open dataset (non-commercial license)
View
P
External · PRF-3004

PKU-SafeRLHF

A large human-annotated safety-preference dataset where each question has two model responses ranked separately for helpfulness and harmlessness, plus per-response safety meta-labels across 19 harm categories.

Type Preference dataVolume 83.4K preference entries (Q-A pairs, each with two ranked responses)Format JSONL (per-model train.jsonl / test.jsonl; auto-converted to Hugging Face Parquet)Access PUBLIC LICENSE
Free
open dataset (HF)
View
I
External · SFT-2104

Infinity-Instruct

BAAI's large-scale open SFT collection; multi-config (7M chat, 3M foundational, plus dated Gen sets). Gated on Hugging Face.

Type SFT datasetVolume 7449106 examplesFormat Parquet (Hugging Face)Access GATED
Free
open-source license
View
t(
External · RLE-1206

tau-bench (τ-bench)

Tool-agent-user benchmark of 165 customer-service tasks (retail 115, airline 50) with policy-following and pass@k evaluation.

Type RL environmentsVolume 165 tasks (retail 115 + airline 50)Format Python package + JSON domain dataAccess PUBLIC LICENSE
Free
open-source license
View
W
External · SFT-2107

WildChat-1M

Real user–ChatGPT (GPT-3.5/GPT-4) conversations collected by Ai2 with metadata and moderation labels; current train split ~838K conversations.

Type SFT datasetVolume 837989 conversationsFormat Parquet (Hugging Face)Access PUBLIC LICENSE
Free
open-source license
View
W
External · RLE-1101

WebArena

Realistic self-hosted web environment: 812 long-horizon tasks over shopping, forum, GitLab, CMS and maps with execution-based evaluation.

Type RL environmentsVolume 812 tasks (from 241 intent templates)Format JSON task configs + Python (gym) + self-hosted Docker sitesAccess PUBLIC LICENSE
Free
open-source license
View
NR
External · SFT-2404

No Robots

10,000 human-written instruction-and-demonstration pairs across 10 task categories, created by skilled human annotators (no model-generated data) for supervised fine-tuning.

Type SFT datasetVolume 10,000 instruction-demonstration pairsFormat parquetAccess PUBLIC LICENSE
Free
open dataset (HF)
View
C
External · SFT-2306

CodeFeedback-Filtered-Instruction

156.5K high-quality single-turn code instructions filtered (complexity 4-5 via Qwen-72B-Chat) from four open code-instruction sources; M-A-P.

Type SFT datasetVolume 156526 rowsFormat Parquet (Hugging Face)Access PUBLIC LICENSE
Free
open-source license
View
o
External · SFT-2206

orca-math-word-problems-200k

200k grade-school math word problems with GPT-4-Turbo-generated worked solutions; English, text explanations only.

Type SFT datasetVolume 200k word-problem QA pairsFormat Parquet (Hugging Face)Access PUBLIC LICENSE
Free
open-source license
View
AD
External · SFT-2407

Aya Dataset

204,112 human-authored multilingual instruction prompt-completion pairs across 65 languages, curated by native/fluent speakers via Cohere Labs' Aya Annotation Platform for multilingual instruction tuning.

Type SFT datasetVolume 204,112 prompt-completion pairs (202,362 train + 1,750 test)Format ParquetAccess PUBLIC LICENSE
Free
open dataset (HF)
View
W
External · AGT-4008

WebLINX

Expert demonstrations of conversational, multi-turn website navigation, where a navigator agent must predict the next web action (click/say/load/submit/change) from dialogue history and the page DOM.

Type Agent tracesVolume 2,300 expert demonstrations (multi-turn web-navigation episodes; ~100K interaction steps; HF 'chat' config exposes 24,418 train turn-level rows / 58.7k total rows)Format JSON (gzip-compressed .json.gz; auto-converted to Parquet on Hugging Face)Access PUBLIC LICENSE
Free
open dataset (Hugging Face, CC BY-NC-SA 4.0, non-commercial)
View
U
External · PRF-3001

UltraFeedback

A large-scale, fine-grained preference dataset of ~64k prompts, each with 4 model completions rated by GPT-4 across four aspects, for training reward and critique models.

Type Preference dataVolume 63967 prompts (train split rows; 256k completions total)Format Parquet (HF datasets); JSON in source repoAccess PUBLIC LICENSE
Free
open dataset (HF)
View
OC
External · SFT-2406

OpenAssistant Conversations v2 (OASST2)

A human-generated, human-annotated corpus of multilingual assistant-style conversation trees released by the OpenAssistant community for supervised fine-tuning and alignment research.

Type SFT datasetVolume 135,174 messages (128,575 train + 6,599 validation)Format ParquetAccess PUBLIC LICENSE
Free
open dataset (HF)
View
AA
External · RQA-5003

AI2 ARC

A dataset of 7,787 genuine grade-school-level multiple-choice science questions, split into a harder Challenge Set and an Easy Set, for evaluating advanced question answering and reasoning.

Type Reasoning QAVolume 7787 questionsFormat parquetAccess PUBLIC LICENSE
Free
open dataset (HF)
View
M
External · AGT-4007

Mind2Web

2,350 human-demonstrated web-navigation task trajectories across 137 real websites, each pairing a natural-language instruction with a step-by-step sequence of DOM-grounded CLICK/TYPE/SELECT actions and full HTML snapshots.

Type Agent tracesVolume 2,350 human-demonstrated web task trajectories (1,009 train + 1,341 test: 252 Cross-Task + 177 Cross-Website + 912 Cross-Domain)Format JSON (task records; each action embeds raw/cleaned HTML snapshots and DOM element candidates)Access PUBLIC LICENSE
Free
open dataset (Hugging Face, CC-BY-4.0)
View
MO
External · SFT-2301

Magicoder OSS-Instruct 75K & Evol-Instruct 110K

ISE-UIUC's paired Magicoder code-instruction datasets: OSS-Instruct (75K, seeded from open-source snippets) and Evol-Instruct (110K, decontaminated evol-codealpaca).

Type SFT datasetVolume 186380 rowsFormat Parquet (Hugging Face)Access PUBLIC LICENSE
Free
open-source license
View
IO
External · PRF-3008

Intel Orca DPO Pairs

A ~12.9K-example preference dataset in Direct Preference Optimization (DPO) format, derived from Open-Orca/OpenOrca, pairing a 'chosen' and 'rejected' response for each instruction prompt.

Type Preference dataVolume 12,859 preference pairs (DPO triples: prompt + chosen + rejected)Format JSONL (single file orca_rlhf.jsonl; also auto-converted to Parquet on Hugging Face)Access PUBLIC LICENSE
Free
open-source license (Apache-2.0)
View
O2
External · SFT-2002

OpenHermes 2.5

Teknium's ~1M-row compilation of primarily GPT-4-generated instruction, chat, coding and reasoning data in ShareGPT (from/value) format.

Type SFT datasetVolume 1001551 rowsFormat Parquet (Hugging Face)Access PUBLIC LICENSE
Free
open-source license
View