VLANeXt: A Simple and Research-Oriented Codebase for Robotics Research
Blog post from Hugging Face
VLANeXt is an open, research-oriented codebase and set of design recipes for vision-language-action robotics models, developed through more than 500 experiments examining model architecture, perception, action representation, and training choices. Starting from an RT-2-style baseline, the authors found that a dedicated policy module, action chunking, continuous flow-matching action generation, stronger vision-language backbones, soft VLM-policy connections, multi-view camera input, VLM-side proprioception, and frequency-domain action regularization substantially improved performance, while visual history offered limited benefits and world modeling increased training cost. The resulting 2.5B-parameter baseline reportedly achieved strong results on LIBERO and improved robustness on LIBERO-plus, with real-robot demonstrations on cleaning, drawer-opening, basket-lifting, and bimanual tasks. The expanded codebase also provides baselines for latent action pretraining, smaller and larger backbones, latent-space future prediction, and world-action modeling, enabling comparisons across model scales and objectives; reported experiments suggest that VQ-VAE latent actions, DINO-based predictive learning, and video-generation-based world-action modeling can further improve results, although larger models do not always perform better when fine-tuning data is limited.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| AI Model Fine-tuning | 6 | 516 | 143 | 56 | -47% |
| LLM | 1 | 4,718 | 960 | 222 | -38% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.