Convex Markets / Datasets / Demonstrations

Dataset catalog

Filter by type, access, and pricing. Specs show before you open the product page.

8 results
L
External · DEM-6005

LIBERO

Human-teleoperated demonstration data for the LIBERO lifelong robot-manipulation benchmark: 6,500 successful trajectories (50 per task) across 130 language-conditioned tasks in four suites, stored as robomimic-format HDF5.

Type DemonstrationsVolume 6500 human-teleoperated demonstrations (130 tasks x 50 demos)Format HDF5 (robomimic-format demonstrations) with BDDL task-definition filesAccess PUBLIC LICENSE
Free
open dataset (CC BY 4.0)
View
OX
External · DEM-6001

Open X-Embodiment

Open X-Embodiment is an aggregated robot-learning dataset that pools over 1 million real robot demonstration trajectories from 22 embodiments across 60 datasets into a single standardized RLDS/TFDS format.

Type DemonstrationsVolume 1,000,000+ real robot trajectories (RLDS episodes)Format RLDS episode format stored as TFDS / TFRecord (loadable via tensorflow_datasets)Access PUBLIC LICENSE
Free
open dataset (RLDS/TFDS, public GCS)
View
D
External · DEM-6003

DROID

DROID is a large-scale in-the-wild robot manipulation dataset of 76,000 human-teleoperated Franka Panda demonstration trajectories (350 hours) with multi-view stereo RGB, robot state/action, and natural-language task instructions.

Type DemonstrationsVolume 76,000 teleoperated demonstration trajectories (episodes)Format RLDS / TensorFlow Datasets (tfds.load("droid")); also raw HDF5 and LeRobot parquet+mp4 conversionsAccess PUBLIC LICENSE
Free
open dataset (CC-BY 4.0)
View
MH
External · DEM-6008

MineRL Human Demonstrations

A large-scale dataset of over 60 million automatically-annotated human state-action pairs recorded while people play Minecraft across a set of related item-acquisition and navigation tasks.

Type DemonstrationsVolume 60,000,000+ state-action pairsFormat Per-task compressed archives distributed for download; loaded via the minerl.data API as OpenAI Gym Dict observation/action tuples yielded as (current_state, action, reward, next_state, done) by BufferedBatchIterAccess PUBLIC LICENSE
Free
open dataset (non-commercial license)
View
L
External · DEM-6006

Language-Table

A large suite of human-collected, language-conditioned tabletop block-manipulation demonstrations from Robotics at Google, where a robot pushes colored blocks in response to natural-language instructions, released in RLDS/TFDS format.

Type DemonstrationsVolume 442,226 real-robot demonstration episodes (language_table split; 1,639,544 episodes total across all 9 released variants)Format RLDS / TFDS (TFRecord); episodes as sequences of stepsAccess PUBLIC LICENSE
Free
open-source license (Apache-2.0)
View
BV
External · DEM-6002

BridgeData V2

60,096 real WidowX 250 robot manipulation trajectories (teleoperated + scripted) across 24 environments and 13 skills, each labeled with natural-language task instructions.

Type DemonstrationsVolume 60096 trajectoriesFormat Per-trajectory raw files (obs_dict.pkl, policy_out.pkl, agent_data.pkl, lang.txt, images0/im_*.jpg 640x480 JPEG); also NumPy and TFRecord/RLDS (TFDS, 256x256) conversionsAccess PUBLIC LICENSE
Free
open dataset (CC-BY-4.0)
View
A/
External · DEM-6007

ALOHA / ACT Demonstrations

Human-teleoperated bimanual fine-manipulation demonstrations collected with the low-cost open-source ALOHA hardware for the ACT paper, recorded as multi-camera RGB video plus 14-DoF joint states/actions in per-episode HDF5 files.

Type DemonstrationsVolume 50 human-teleoperated demonstrations per task (Thread Velcro: 100; 4 simulated task configs of 50 episodes each also released)Format HDF5 (one .hdf5 file per episode)Access PUBLIC LICENSE
Free
open-source license
View
C
External · DEM-6004

CALVIN

CALVIN is an open-source simulated benchmark of ~24 hours of human teleoperated play data (6 h in each of 4 environments A/B/C/D) with crowd-sourced natural-language annotations for learning long-horizon, language-conditioned robot manipulation over 34 tasks.

Type DemonstrationsVolume 24 hours of teleoperated play data (6 h in each of 4 environments A/B/C/D); 34 annotated tasksFormat Per-timestep NumPy .npz frames + .npy language-annotation files (episode_*.npz, lang_annotations/auto_lang_ann.npy); PyBullet simulationAccess PUBLIC LICENSE
Free
open-source license (MIT)
View