FLUX 3 Action: a world action model you can fine-tune
Blog post from Hugging Face
FLUX 3 Action is an open-weights, 7-billion-parameter world action model from Black Forest Labs that jointly predicts future video frames and action sequences from camera input, state information, and text instructions. Fine-tuned on DROID, it achieved a 42.92% success rate on the RoboLab benchmark, exceeding several larger or proprietary policies, and is released under the FLUX Kommunity License with code, weights, and LeRobot integration for real-arm deployment. The diffusion-transformer model supports multiple robot embodiments through shared visual and language encoders alongside embodiment-specific action projections, producing 32 planned actions per call and enabling iterative observation and replanning. The authors demonstrate parameter-efficient adaptation on an SO-101 robotic arm using roughly 200 teleoperated pick-and-place episodes, reporting generalization to unfamiliar objects, containers, partial occlusions, camera positions, and recovery from errors. Beyond robotics, the same model was fine-tuned using 800 scripted episodes per task for a shooter game, a racing game, and an indoor drone simulator, where it showed competitive gameplay, road navigation, and language-guided movement in unseen room layouts.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| AI Model Fine-tuning | 6 | 139 | 28 | 14 | -75% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.