OpenWAM is an open, composable system that unifies video prediction and robot control on a shared causal video backbone, enabling flexible ordering and interaction between predicted videos and actions. It adapts a Wan2.2-5B visual expert via causal robot-video pretraining on roughly 3.34 million trajectories (≈14.6K hours) and pairs it with a 2B action expert in a Mixture-of-Transformers architecture that preserves separate normalizations and feed-forward layers while sharing attention. The architecture supports four interaction programs - video-then-action (VTA), action-then-video (ATV), joint, and decoupled - uses latent flow matching (four latent frames per 16-step action chunk) with proprioception per chunk, and exposes independently trainable local-context inverse and forward dynamics models (IDM/FDM) that take only current observation, proprioception, and a proposed future trajectory.
Empirical results show dramatic gains from causal robot-video training and counterfactual supervision. VTA success on LIBERO-Long rises from 68.4% to 97.8% after adaptation. A local-context IDM trained on demonstrations plus 32K counterfactual segments (LIBERO-LONG-CF) achieves 84.0% mean success on four held-out LIBERO-90 tasks versus 47.0% for a full-context IDM and 21.5% for a local-context model trained only on demonstrations. For FDMs, counterfactual data cuts RGB prediction error by 34.5% and boosts outcome identification from 21.1% to 71.3% among 16 same-state outcomes. OpenWAM serves as a standardized testbed to compare world-action interaction designs and to study dynamics learning from video beyond successful demonstrations.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.