SKU RLE-1102 · Sold by External
VisualWebArena
Product specifications
| SKU | RLE-1102 |
|---|---|
| Data type | RL environments |
| Volume | 910 tasks (Classifieds 234, Shopping 466, Reddit 210) |
| Size on disk | Not published by source |
| Format | JSON task configs + Python (gym) + self-hosted Docker sites |
| Access model | PUBLIC LICENSE |
| Pricing | Free · open-source license |
| Quality score | — |
| License | MIT |
VisualWebArena extends the WebArena framework to visually grounded tasks that require understanding image content and image-based goals across three self-hosted sites: Classifieds, Shopping, and Reddit. Agents use the same gym-style browser action space (click, type, scroll) augmented with set-of-marks visual grounding, and receive multimodal observations (screenshot plus accessibility tree/DOM). It provides 910 human-authored tasks (234 Classifieds, 466 Shopping, 210 Reddit) with execution-based evaluation (string_match, url_match, program_html) and per-task difficulty labels.