Dataset catalog
Filter by type, access, and pricing. Specs show before you open the product page.
Prime Intellect Environments Hub / verifiers
Community hub and Python registry of open-source RL environments, built on the MIT-licensed verifiers library; site headlines 2,500+ environments as of Jul 2026.
Craftax
Craftax is a lightning-fast, JAX-native open-ended RL benchmark that reimplements and extends Crafter with NetHack-inspired roguelike mechanics, using ~67 unlockable achievements as sparse-reward tasks.
Terminal-Bench
Benchmark of 89 hard, realistic command-line tasks run in Docker; agents issue tmux/bash keystrokes verified by outcome tests.
Reasoning Gym
Python library of 100+ procedural dataset generators with algorithmic verifiers for RL with verifiable rewards; generates virtually unlimited reasoning problems with adjustable difficulty.
SkyRL-Gym
Gymnasium-API library of tool-use environments (math, code, search, text-to-SQL) for LLM post-training, part of the SkyRL RL stack from NovaSky.
AgentGym (14 environments)
Unified framework of 14 interactive environments across 7 scenario types with a common HTTP/ReAct interface, plus trajectory datasets and the AgentEval benchmark.
tau2-bench (τ²-bench)
Dual-control tool-agent benchmark (278 tasks: retail 114, telecom 114, airline 50) where both agent and user can call tools.
Search-R1
Open-source RL framework that trains LLMs to interleave reasoning with live search-engine calls; the retriever is treated as part of the RL environment.
Meta-World
Open-source benchmark of 50 simulated Sawyer-arm robotic manipulation tasks (Gymnasium API) with a shared 4-D continuous action space, dense shaped rewards, and a per-task binary success metric for multi-task and meta-RL.
RAGEN Environments
Reinforcement-learning framework with 10 stylized interactive environments (Sokoban, FrozenLake, Bandit, Countdown, Sudoku, WebShop, etc.) for training multi-turn reasoning agents.
TextArena
Open collection of 100+ competitive/cooperative text games with an OpenAI-Gym-style interface, online play, and a TrueSkill leaderboard for LLM agents.
R2E-Gym
Procedurally-curated executable SWE gym of 8,135+ Dockerized Python bug-fix environments with unit tests for training agents.
BrowserGym
Unified Gymnasium harness for web agents aggregating MiniWoB, WebArena, VisualWebArena, WorkArena, AssistantBench, WebLINX and more.
SWE-Gym
Executable RL environment of 2,438 real Python GitHub-issue tasks with runtimes and unit tests for training SWE agents.
OSWorld
Real-computer benchmark: 369 open-ended desktop/web tasks on Ubuntu (also Windows/macOS) with execution-based, script evaluation.
Gymnasium (Farama)
The standard Python RL API and reference environment suite (successor to OpenAI Gym) covering Classic Control, Box2D, Toy Text, MuJoCo, and Atari.
SWE-bench / SWE-bench Verified
Benchmark of real GitHub-issue tasks (2,294 full; 500 human-verified) evaluated by FAIL_TO_PASS/PASS_TO_PASS unit tests in Docker.
VisualWebArena
Multimodal, visually grounded web-agent benchmark: 910 tasks over self-hosted Classifieds, Shopping and Reddit sites.
AppWorld
Simulated world of 9 apps and 457 APIs with 750 interactive coding tasks evaluated by state-based unit tests.
WorkArena / WorkArena++ (ServiceNow)
Enterprise knowledge-work benchmark on live ServiceNow instances: 33 L1 atomic tasks (19,912 instances) plus 682 L2/L3 compositional tasks.