hn.today

Flux 3 Action: A 7B open-weight world action model for robots

huggingface.co12 points0 comments
Screenshot of Flux 3 Action: A 7B open-weight world action model for robots

FLUX 3 Action is an open-weights 7B "world action model" that jointly predicts future video frames and multi-step actions using a diffusion transformer. Visual input is encoded by a frozen video VAE and instructions by a frozen Qwen3-VL-4B text encoder; actions are represented as a second token stream interleaved with frame tokens and denoised together. Each call receives composited camera frames, a state vector (e.g., joint positions), and a caption, and returns 32 future actions and optionally 32 decoded frames. At control time practitioners skip frame decoding, execute the first few predicted actions, observe, and replan. Training used NVIDIA GB200 hardware with custom kernels and PEFT fine-tuning recipes; weights are released under the FLUX Kommunity License v1.0 and code is on GitHub.

Fine-tuned on the DROID dataset, FLUX 3 Action tops the RoboLab leaderboard with 42.92% success versus 36.8% for a 16B Cosmos3 policy, and a DROID and SO-101 arm checkpoint are integrated into LeRobot. Practical evaluations include adapting a policy to an SO-101 arm from ~200 teleoperated episodes (recovering from mistakes, handling unseen containers and camera shifts), training a single joint model on two games (GRUNT shooter and VECTOR racer) from 800 scripted episodes each with near-bot performance, and controlling an indoor drone from 800 Isaac Sim flights while generalizing to novel rooms and phrasing. Documentation, fine-tuning guides, and examples for robotics and games are provided.

Read on huggingface.co0 comments on Hacker News

Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.

More in AI

The daily digest

Today's best Hacker News stories, summarized and screenshotted, one email a day.