SKU RLE-1203 · Sold by External
SWE-bench / SWE-bench Verified
Product specifications
| SKU | RLE-1203 |
|---|---|
| Data type | RL environments |
| Volume | 2294 task instances |
| Size on disk | Not published by source |
| Format | Hugging Face dataset (Parquet/JSON) + Docker |
| Access model | PUBLIC LICENSE |
| Pricing | Free · open-source license |
| Quality score | — |
| License | MIT |
SWE-bench evaluates LLMs/agents on real-world software issues collected from GitHub: given a repository and an issue, the agent must produce a patch that passes hidden regression tests. The full test set has 2,294 instances; key subsets include SWE-bench Verified (500 human-validated), Lite (300), Multimodal, and Multilingual, plus a large train split without executable environments. Evaluation runs in sandboxed Docker with FAIL_TO_PASS/PASS_TO_PASS test checks. Data is distributed as public Hugging Face datasets (SWE-bench, _Verified, _Lite, _Multimodal).