SKU RLE-1204 · Sold by External

Terminal-Bench

Product specifications

SKURLE-1204
Data typeRL environments
Volume89 tasks
Size on diskNot published by source
FormatDocker + YAML/dir task specs + Python harness
Access modelPUBLIC LICENSE
PricingFree · open-source license
Quality score
LicenseApache-2.0
Terminal-Bench evaluates AI agents on hard, realistic command-line tasks (software engineering, security, scientific computing, data science, ML, sysadmin, debugging) executed inside Docker containers. Agents issue tmux keystrokes / bash commands via a ReAct-style loop (Terminus reference agent) and are scored by outcome-based test scripts checking the final container state. Terminal-Bench 2.0 comprises 89 curated tasks (from 229 created by 93 contributors); 2.1 is the current release. Each task is a directory with an instruction, a Dockerized environment, a test script, and a reference solution.