AI / Embodied AI

HomeBody gives a humanoid spatial memory for long kitchen tasks

Stanford and Caltech researchers connected GPT-6 Astra to a reusable skill library and persistent spatial map, allowing a Unitree G1 to tidy an unfamiliar kitchen and retrieve an out-of-view object. The demonstration is promising, but it remains a research prototype with substantial setup, latency and hardware constraints.

INNOVOX News DeskSep 27, 2026 · 7 min read
A white Unitree G1 humanoid robot standing on display at UAV Expo 2024
Sayanesy / Wikimedia Commons · CC0 1.0 public-domain dedication

The story

A humanoid robot has cleaned an unfamiliar kitchen and recovered medicine from a drawer using a system that gives a frontier AI model access to persistent spatial memory and a library of physical skills. The project, called HomeBody, was developed by researchers affiliated with Stanford University and Caltech and uses GPT-6 Astra to plan work for a Unitree G1. Independent coverage of the demonstration appeared on September 27, bringing new attention to an approach that sits between fully scripted robots and end-to-end learned control.

HomeBody does not send language-model output straight to the robot's motors. Instead, Astra receives visual observations, remembered locations, gripper state and the result of the previous action, then selects a skill and a target through a structured tool call. The available skills include navigation, picking, placing, opening a drawer and retrieving an item from an open drawer. Lower-level software handles motion planning, inverse kinematics, collision checking, balance and high-frequency arm and hand control.

That distinction matters because many embodied-AI systems use three layers: a vision-language model for reasoning, a learned vision-language-action policy for translating intent into commands, and a controller for coordinated motion. HomeBody removes the middle learned policy from this particular stack. Its researchers ask whether a more capable general model can instead orchestrate reliable, reusable skills, receiving feedback when execution fails and revising the plan without retraining a policy for each environment.

Before performing the kitchen tasks, the robot explores the space. The system records camera observations, LiDAR scans, SLAM geometry, joint poses and selected waypoints. Astra then helps build a digital twin in Nvidia Isaac Sim, aligning that reconstruction with measured geometry. Camera frames and descriptions are stored in the same coordinate system, allowing the planner to reason about objects that are no longer visible and to return to locations observed earlier.

The team shows two long-horizon demonstrations. In one, the G1 gathers coffee bags on an island and discards spoiled milk and orange-juice cartons, coordinating repeated navigation, grasping and placement steps. In the other, it interprets an underspecified request for forgotten medicine, recalls that the item is in a drawer, opens the drawer, retrieves the medicine and throws away a bad carton. The system can switch hands when objects sit on different sides of the body and can return local failure information to the model for replanning.

The implementation combines learned and conventional robotics rather than replacing either. Object selection starts from an image point chosen by the model; segmentation and stereo depth estimate the object's three-dimensional position. The arm planner creates a timed path, solves inverse kinematics and checks clearance. Visual tracking corrects alignment as the robot approaches, while bounded local retries handle failed grasps. A pretrained whole-body controller coordinates the legs with upper-body motion at faster control rates than a remote language model could sustain.

The evidence is still limited. The project page presents compelling task videos and a detailed technical description, but it does not report a broad benchmark, a comparison across several homes or a statistically meaningful task-success rate. Setup requires a Real2Sim reconstruction, which adds time and API cost. Astra runs remotely and introduces pauses between skills, while local perception and planning require a laptop with an RTX 4090 GPU. The researchers also report finger-servo overheating during extended operation and physical limits in reach, manipulation and endurance.

There is also a transparency gap to close. The project's public GitHub repository identifies the authors and currently says that code is coming soon. The Decoder reported the demonstration and described its latency, heat and compute constraints, but its statement that the code is available should be read cautiously: as of publication, the repository is visible, while the implementation itself is not yet released. Reproducibility will depend on code, configuration details, task logs and evaluation protocols becoming available.

INNOVOX analysis: HomeBody points toward a practical modular architecture for general-purpose robots. A high-level model can focus on interpreting goals, remembering context and sequencing actions, while specialized components remain responsible for safety-critical geometry and control. Modular skills may be easier to test and replace than a single opaque policy. The tradeoff is that capability depends heavily on the quality and coverage of the skill library, the map and the feedback loop; apparent generality can shrink quickly when a task demands an action the library does not contain.

The next step is not a more polished kitchen video but harder measurement. Researchers need to publish repeated trials, failure categories, task duration, human interventions and performance in changed layouts with unfamiliar objects. They should test whether the system refuses unsafe or ambiguous commands and whether spatial memory stays accurate after people move objects. If HomeBody can preserve its planning advantages while reducing setup, latency and hardware demands, it could help shift humanoids from isolated demonstrations toward adaptable machines whose decisions and physical competencies can be evaluated separately.

INNOVOX analysis

HomeBody's most important contribution is architectural rather than theatrical. It treats a frontier model as a high-level coordinator that chooses grounded tools while classical and learned controllers handle precise motion. That division could make robot behavior easier to extend and inspect, but it also means the demonstration is not evidence that a general model can directly control a humanoid end to end.

What to watch

Watch for release of the promised code, repeated trials across multiple homes and layouts, quantitative success and recovery rates, lower-latency local models, safety constraints for ambiguous instructions and evidence that new skills can be added without destabilizing existing behavior.