Convex Markets / Datasets / All departments

Dataset catalog

Filter by type, access, and pricing. Specs show before you open the product page.

27 results for “RL environments”
PI
External · RLE-1405

Prime Intellect Environments Hub / verifiers

Community hub and Python registry of open-source RL environments, built on the MIT-licensed verifiers library; site headlines 2,500+ environments as of Jul 2026.

Type RL environmentsVolume 2500+ environmentsFormat Python wheels via Prime CLI (uv / pyproject); verifiers library (MIT)Access PUBLIC LICENSE
Free
open-source license
View
T
External · RLE-1204

Terminal-Bench

Benchmark of 89 hard, realistic command-line tasks run in Docker; agents issue tmux/bash keystrokes verified by outcome tests.

Type RL environmentsVolume 89 tasksFormat Docker + YAML/dir task specs + Python harnessAccess PUBLIC LICENSE
Free
open-source license
View
RG
External · RLE-1403

Reasoning Gym

Python library of 100+ procedural dataset generators with algorithmic verifiers for RL with verifiable rewards; generates virtually unlimited reasoning problems with adjustable difficulty.

Type RL environmentsVolume 100+ generatorsFormat Python (pip: reasoning-gym)Access PUBLIC LICENSE
Free
open-source license
View
S
External · RLE-1404

SkyRL-Gym

Gymnasium-API library of tool-use environments (math, code, search, text-to-SQL) for LLM post-training, part of the SkyRL RL stack from NovaSky.

Type RL environmentsVolume tool-use environment library (math/code/search/SQL)Format Python (Gymnasium API; pip)Access PUBLIC LICENSE
Free
open-source license
View
A(
External · RLE-1305

AgentGym (14 environments)

Unified framework of 14 interactive environments across 7 scenario types with a common HTTP/ReAct interface, plus trajectory datasets and the AgentEval benchmark.

Type RL environmentsVolume 14 environmentsFormat Python + HTTP env servers (agentenv)Access PUBLIC LICENSE
Free
open-source license
View
t(
External · RLE-1207

tau2-bench (τ²-bench)

Dual-control tool-agent benchmark (278 tasks: retail 114, telecom 114, airline 50) where both agent and user can call tools.

Type RL environmentsVolume 278 tasks (retail 114 + telecom 114 + airline 50)Format Python package + JSON domain dataAccess PUBLIC LICENSE
Free
open-source license
View
S
External · RLE-1402

Search-R1

Open-source RL framework that trains LLMs to interleave reasoning with live search-engine calls; the retriever is treated as part of the RL environment.

Type RL environmentsVolume retrieval-augmented QA training frameworkFormat Python (built on veRL; installed from source via conda/pip)Access PUBLIC LICENSE
Free
open-source license
View
RE
External · RLE-1401

RAGEN Environments

Reinforcement-learning framework with 10 stylized interactive environments (Sokoban, FrozenLake, Bandit, Countdown, Sudoku, WebShop, etc.) for training multi-turn reasoning agents.

Type RL environmentsVolume 10 environmentsFormat Python (Gym-compatible interface; installed from source via setup script)Access PUBLIC LICENSE
Free
open-source license
View
T
External · RLE-1306

TextArena

Open collection of 100+ competitive/cooperative text games with an OpenAI-Gym-style interface, online play, and a TrueSkill leaderboard for LLM agents.

Type RL environmentsVolume 100+ gamesFormat Python (pip: textarena) + Gym-style APIAccess PUBLIC LICENSE
Free
open-source license
View
R
External · RLE-1202

R2E-Gym

Procedurally-curated executable SWE gym of 8,135+ Dockerized Python bug-fix environments with unit tests for training agents.

Type RL environmentsVolume 8135 executable environmentsFormat Hugging Face dataset + Docker images (Python/Gym)Access PUBLIC LICENSE
Free
open-source license
View
B
External · RLE-1107

BrowserGym

Unified Gymnasium harness for web agents aggregating MiniWoB, WebArena, VisualWebArena, WorkArena, AssistantBench, WebLINX and more.

Type RL environmentsVolume aggregate harness (no single fixed count; sums member benchmarks e.g. WebArena 812, VisualWebArena 910, WorkArena 33/682, MiniWoB 128)Format Python packages (pip install browsergym / browsergym-core + per-benchmark extras) + PlaywrightAccess PUBLIC LICENSE
Free
open-source license
View
S
External · RLE-1201

SWE-Gym

Executable RL environment of 2,438 real Python GitHub-issue tasks with runtimes and unit tests for training SWE agents.

Type RL environmentsVolume 2438 task instancesFormat Hugging Face dataset + Docker imagesAccess PUBLIC LICENSE
Free
open-source license
View
O
External · RLE-1106

OSWorld

Real-computer benchmark: 369 open-ended desktop/web tasks on Ubuntu (also Windows/macOS) with execution-based, script evaluation.

Type RL environmentsVolume 369 tasks (execution-based; 134 unique evaluators)Format Python + downloadable VM images (Docker / VMware / VirtualBox); JSON task configsAccess PUBLIC LICENSE
Free
open-source license
View
G(
External · RLE-1406

Gymnasium (Farama)

The standard Python RL API and reference environment suite (successor to OpenAI Gym) covering Classic Control, Box2D, Toy Text, MuJoCo, and Atari.

Type RL environmentsVolume dozens reference environmentsFormat Python (pip: gymnasium; Gym reset/step API)Access PUBLIC LICENSE
Free
open-source license
View
S/
External · RLE-1203

SWE-bench / SWE-bench Verified

Benchmark of real GitHub-issue tasks (2,294 full; 500 human-verified) evaluated by FAIL_TO_PASS/PASS_TO_PASS unit tests in Docker.

Type RL environmentsVolume 2294 task instancesFormat Hugging Face dataset (Parquet/JSON) + DockerAccess PUBLIC LICENSE
Free
open-source license
View
V
External · RLE-1102

VisualWebArena

Multimodal, visually grounded web-agent benchmark: 910 tasks over self-hosted Classifieds, Shopping and Reddit sites.

Type RL environmentsVolume 910 tasks (Classifieds 234, Shopping 466, Reddit 210)Format JSON task configs + Python (gym) + self-hosted Docker sitesAccess PUBLIC LICENSE
Free
open-source license
View
A
External · RLE-1205

AppWorld

Simulated world of 9 apps and 457 APIs with 750 interactive coding tasks evaluated by state-based unit tests.

Type RL environmentsVolume 750 tasksFormat Python package + SQLite DBs + JSON task specsAccess PUBLIC LICENSE
Free
open-source license
View
W/
External · RLE-1104

WorkArena / WorkArena++ (ServiceNow)

Enterprise knowledge-work benchmark on live ServiceNow instances: 33 L1 atomic tasks (19,912 instances) plus 682 L2/L3 compositional tasks.

Type RL environmentsVolume 33 L1 atomic tasks (19,912 instances) + 682 WorkArena++ L2/L3 tasksFormat Python package (pip install browsergym-workarena) + live ServiceNow instance + PlaywrightAccess GATED
Free
open-source license
View
t(
External · RLE-1206

tau-bench (τ-bench)

Tool-agent-user benchmark of 165 customer-service tasks (retail 115, airline 50) with policy-following and pass@k evaluation.

Type RL environmentsVolume 165 tasks (retail 115 + airline 50)Format Python package + JSON domain dataAccess PUBLIC LICENSE
Free
open-source license
View
W
External · RLE-1101

WebArena

Realistic self-hosted web environment: 812 long-horizon tasks over shopping, forum, GitLab, CMS and maps with execution-based evaluation.

Type RL environmentsVolume 812 tasks (from 241 intent templates)Format JSON task configs + Python (gym) + self-hosted Docker sitesAccess PUBLIC LICENSE
Free
open-source license
View