Convex Markets / Datasets / All departments

Dataset catalog

Filter by type, access, and pricing. Specs show before you open the product page.

37 results for “Agent”
T
External · RLE-1204

Terminal-Bench

Benchmark of 89 hard, realistic command-line tasks run in Docker; agents issue tmux/bash keystrokes verified by outcome tests.

Type RL environmentsVolume 89 tasksFormat Docker + YAML/dir task specs + Python harnessAccess PUBLIC LICENSE
Free
open-source license
View
A(
External · RLE-1305

AgentGym (14 environments)

Unified framework of 14 interactive environments across 7 scenario types with a common HTTP/ReAct interface, plus trajectory datasets and the AgentEval benchmark.

Type RL environmentsVolume 14 environmentsFormat Python + HTTP env servers (agentenv)Access PUBLIC LICENSE
Free
open-source license
View
t(
External · RLE-1207

tau2-bench (τ²-bench)

Dual-control tool-agent benchmark (278 tasks: retail 114, telecom 114, airline 50) where both agent and user can call tools.

Type RL environmentsVolume 278 tasks (retail 114 + telecom 114 + airline 50)Format Python package + JSON domain dataAccess PUBLIC LICENSE
Free
open-source license
View
L
External · DEM-6005

LIBERO

Human-teleoperated demonstration data for the LIBERO lifelong robot-manipulation benchmark: 6,500 successful trajectories (50 per task) across 130 language-conditioned tasks in four suites, stored as robomimic-format HDF5.

Type DemonstrationsVolume 6500 human-teleoperated demonstrations (130 tasks x 50 demos)Format HDF5 (robomimic-format demonstrations) with BDDL task-definition filesAccess PUBLIC LICENSE
Free
open dataset (CC BY 4.0)
View
RE
External · RLE-1401

RAGEN Environments

Reinforcement-learning framework with 10 stylized interactive environments (Sokoban, FrozenLake, Bandit, Countdown, Sudoku, WebShop, etc.) for training multi-turn reasoning agents.

Type RL environmentsVolume 10 environmentsFormat Python (Gym-compatible interface; installed from source via setup script)Access PUBLIC LICENSE
Free
open-source license
View
T
External · RLE-1306

TextArena

Open collection of 100+ competitive/cooperative text games with an OpenAI-Gym-style interface, online play, and a TrueSkill leaderboard for LLM agents.

Type RL environmentsVolume 100+ gamesFormat Python (pip: textarena) + Gym-style APIAccess PUBLIC LICENSE
Free
open-source license
View
R
External · RLE-1202

R2E-Gym

Procedurally-curated executable SWE gym of 8,135+ Dockerized Python bug-fix environments with unit tests for training agents.

Type RL environmentsVolume 8135 executable environmentsFormat Hugging Face dataset + Docker images (Python/Gym)Access PUBLIC LICENSE
Free
open-source license
View
xF
External · AGT-4005

xLAM Function-Calling 60k

60,000 verifiable single/multi/parallel function-calling instances (query + available tools + reference tool-call answers) generated and triple-verified by Salesforce's APIGen pipeline.

Type Agent tracesVolume 60000 function-calling instancesFormat JSON (single file xlam_function_calling_60k.json; config 'dataset', split 'train')Access GATED
Free
gated open dataset (HF, auto-approval)
View
B
External · RLE-1107

BrowserGym

Unified Gymnasium harness for web agents aggregating MiniWoB, WebArena, VisualWebArena, WorkArena, AssistantBench, WebLINX and more.

Type RL environmentsVolume aggregate harness (no single fixed count; sums member benchmarks e.g. WebArena 812, VisualWebArena 910, WorkArena 33/682, MiniWoB 128)Format Python packages (pip install browsergym / browsergym-core + per-benchmark extras) + PlaywrightAccess PUBLIC LICENSE
Free
open-source license
View
S
External · RLE-1201

SWE-Gym

Executable RL environment of 2,438 real Python GitHub-issue tasks with runtimes and unit tests for training SWE agents.

Type RL environmentsVolume 2438 task instancesFormat Hugging Face dataset + Docker imagesAccess PUBLIC LICENSE
Free
open-source license
View
GA
External · AGT-4003

Gorilla APIBench

Instruction-to-API-call dataset over HuggingFace, TorchHub, and TensorHub APIs, used to train and evaluate LLMs that write correct API calls (the Gorilla / APIBench benchmark).

Type Agent tracesVolume 17003 instruction-API-call pairs (15,218 train + 1,785 eval across HuggingFace/TorchHub/TensorHub), over 1,726 reference APIsFormat JSON Lines (one JSON object per line) - {hub}_train.json / {hub}_eval.json for instruction pairs; {hub}_api.jsonl for the reference API databaseAccess PUBLIC LICENSE
Free
open-source dataset (Apache-2.0)
View
O
External · RLE-1106

OSWorld

Real-computer benchmark: 369 open-ended desktop/web tasks on Ubuntu (also Windows/macOS) with execution-based, script evaluation.

Type RL environmentsVolume 369 tasks (execution-based; 134 unique evaluators)Format Python + downloadable VM images (Docker / VMware / VirtualBox); JSON task configsAccess PUBLIC LICENSE
Free
open-source license
View
OA
External · SFT-2105

Orca AgentInstruct 1M v1

1,046,410 synthetic instruction/response conversations generated by Microsoft's AgentInstruct agentic pipeline across 15 task-type splits.

Type SFT datasetVolume 1046410 examplesFormat Parquet (Hugging Face)Access PUBLIC LICENSE
Free
open-source license
View
T
External · AGT-4009

ToolACE

11,300 synthetic multi-turn function-calling dialogues generated over a self-evolved pool of 26,507 APIs and filtered by a dual-layer (rule-based + model-based) verification pipeline.

Type Agent tracesVolume 11,300 function-calling dialogues (multi-turn conversations)Format JSON (single data.json; each record = {system: string, conversations: list of {from, value}})Access PUBLIC LICENSE
Free
open dataset (HF)
View
S/
External · RLE-1203

SWE-bench / SWE-bench Verified

Benchmark of real GitHub-issue tasks (2,294 full; 500 human-verified) evaluated by FAIL_TO_PASS/PASS_TO_PASS unit tests in Docker.

Type RL environmentsVolume 2294 task instancesFormat Hugging Face dataset (Parquet/JSON) + DockerAccess PUBLIC LICENSE
Free
open-source license
View
A
External · RLE-1205

AppWorld

Simulated world of 9 apps and 457 APIs with 750 interactive coding tasks evaluated by state-based unit tests.

Type RL environmentsVolume 750 tasksFormat Python package + SQLite DBs + JSON task specsAccess PUBLIC LICENSE
Free
open-source license
View
V
External · RLE-1102

VisualWebArena

Multimodal, visually grounded web-agent benchmark: 910 tasks over self-hosted Classifieds, Shopping and Reddit sites.

Type RL environmentsVolume 910 tasks (Classifieds 234, Shopping 466, Reddit 210)Format JSON task configs + Python (gym) + self-hosted Docker sitesAccess PUBLIC LICENSE
Free
open-source license
View
W/
External · RLE-1104

WorkArena / WorkArena++ (ServiceNow)

Enterprise knowledge-work benchmark on live ServiceNow instances: 33 L1 atomic tasks (19,912 instances) plus 682 L2/L3 compositional tasks.

Type RL environmentsVolume 33 L1 atomic tasks (19,912 instances) + 682 WorkArena++ L2/L3 tasksFormat Python package (pip install browsergym-workarena) + live ServiceNow instance + PlaywrightAccess GATED
Free
open-source license
View
t(
External · RLE-1206

tau-bench (τ-bench)

Tool-agent-user benchmark of 165 customer-service tasks (retail 115, airline 50) with policy-following and pass@k evaluation.

Type RL environmentsVolume 165 tasks (retail 115 + airline 50)Format Python package + JSON domain dataAccess PUBLIC LICENSE
Free
open-source license
View
W
External · RLE-1101

WebArena

Realistic self-hosted web environment: 812 long-horizon tasks over shopping, forum, GitLab, CMS and maps with execution-based evaluation.

Type RL environmentsVolume 812 tasks (from 241 intent templates)Format JSON task configs + Python (gym) + self-hosted Docker sitesAccess PUBLIC LICENSE
Free
open-source license
View