LeRobot Datasets for VLA Training: 5 Real and Synthetic Sets
Blog post from Voxel51
A Voxel51 field guide compares five LeRobot v3.0 manipulation datasets for vision-language-action training: two real teleoperated datasets, fmb_multi on a Franka Panda and rh20t_cfg1 on a Flexiv arm, and three synthetic datasets, robocasa-MG_100, robotwin_unified, and InternData-A1. It emphasizes that real robot data is costly and slow to collect because it requires operators, hardware, prepared scenes, and safety measures, leaving the real sources with only 1,804 and 4,258 episodes, while InternData-A1 contains 637,498 simulated trajectories. The synthetic datasets represent distinct strategies: MimicGen recombines 1,250 human demonstrations into varied simulated scenes for RoboCasa, RoboTwin uses an LLM to generate task code and a VLM to detect and repair failures without seed demonstrations, and InternData-A1 programmatically composes skills, assets, and scenes without demonstrations. Each dataset provides different robot embodiments, camera streams, state representations, action spaces, licenses, and export sizes, with several FiftyOne releases being subsets of larger upstream collections. InternData-A1 is the only dataset discussed whose paper reports zero-shot sim-to-real transfer after synthetic-only pretraining, although the guide notes that this evidence is limited and not independently reproduced, while it also cautions that nominal frame rates in the real datasets are inferred rather than derived from per-frame timestamps.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| LLM | 5 | 747 | 162 | 79 | -85% |
| AI Guardrails | 1 | 35 | 22 | 12 | -94% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.