hn.today

HomeBody: A humanoid that explores, remembers, and acts on its own

tml.stanford.edu23 points2 comments
Screenshot of HomeBody: A humanoid that explores, remembers, and acts on its own

HomeBody equips a frontier vision-language model (Astra) to directly command a humanoid by combining persistent spatial memory with a library of composable loco-manipulation skills. Rather than inserting a learned mid-level policy between high-level reasoning and motion control, HomeBody lets a replaceable VLM select skills like navigate, pick, place, open-drawer and pick-from-drawer, then receives execution outcomes and visual feedback to revise plans. In demonstrations on a Unitree G1 in previously unseen kitchens, HomeBody performs long-horizon tasks - multi-object cleanup, occluded-object retrieval, and sequencing bimanual actions - without environment-specific policy training. Memory keyframes let the system find out-of-view items and coordinate repeated trips, grasps and placements across the room.

Deployment uses a three-step pipeline: explore (collect iPhone video, D435i stereo, LiDAR, joint poses and waypoints), Real2Sim (Astra builds a digital twin in Isaac Sim grounded in SLAM geometry), and tasking (send everyday instructions). Technical components align SLAM and simulation via Super Odometry and ICP, convert VLM-selected image points to 3D with segmentation and Fast-FoundationStereo, predict analytic grasps, plan spline reference trajectories with minimum-jerk timing, solve inverse kinematics and check collisions. Skills expose a common interface so new learned or classical controllers plug in; local retries and visual checks correct failures, enabling robust, long-range humanoid autonomy driven by a high-level VLM.

Read on tml.stanford.edu2 comments on Hacker News

Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.

More in AI

The daily digest

Today's best Hacker News stories, summarized and screenshotted, one email a day.