Helix 2.5 is a neural-policy humanoid system pretrained on Index, a massive dataset of human behavior, and evaluated for zero-shot whole-body autonomy across 30 unseen Bay Area homes. A single foundation model was fine-tuned into three long-horizon behaviors - tidying living rooms (picking 13-15 scattered toys into a basket), folding towels, and making beds (placing pillows and smoothing the comforter) - that require integrated perception, locomotion, bimanual manipulation, and active perception. Evaluation used no data from the homes or objects, a single fixed checkpoint per task, and strict pass/fail criteria that granted success only for fully completed tasks; no environment-specific adaptation or further tuning was allowed.
Index pretraining was the dominant factor driving transfer: otherwise-identical policies trained from scratch achieved 9% zero-shot success, while Index-initialized policies reached 56%. Helix 2.5 matched prior Helix 02 success using half the task-specific data and displayed robust whole-body self-correction (repositioning, stance changes, moving around obstacles). Pretraining scale followed a precise human-to-humanoid transfer law: doubling Index data predictably reduced downstream action-prediction loss and allowed forecasting the largest run’s loss to four decimal places (forecast error 0.54% of observed variation). Figure positions this as first evidence that broad human pretraining can enable generalizable humanoid behavior and is committing major data and compute to scale further.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.