Stanford Lets GPT-6 Astra Direct a Humanoid Robot to Tidy an Unfamiliar Kitchen With No Specially Trained Action Model in Between, but Reports No Success Rates

Researchers at Stanford's Movement Lab and Caltech released the HomeBody project in September. The authors are Gio Huh (Caltech), Cayden Gu, Takara E. Truong, C. Karen Liu and Guy Tevet (Stanford, with the last two as equal advisers). The common humanoid architecture has three layers: a vision-language model for reasoning, a vision-language-action (VLA) model for actions, and a low-level controller. HomeBody removes the specially trained VLA in the middle and has the frontier vision-language model GPT Astra compose five skills directly through structured tool calls: navigation, picking, placing, opening a drawer and picking from a drawer. The robot first explores the environment, collecting video, depth camera, LiDAR and SLAM data; Astra helps build a digital twin of the kitchen in Nvidia Isaac Sim, and the objects and locations observed are written into a persistent spatial memory, so the robot can find items that have left its field of view. The project page says that in a previously unseen kitchen, a Unitree G1 guided by Astra can carry out room-wide cleanup such as gathering coffee bags in the middle and throwing away spoiled milk and orange juice cartons, and retrieve a remembered medicine from an underspecified request, without environment-specific training data or additional policy learning. The page reports no success rates or systematic evaluation, only demonstrations. The limitations the authors list include setup time and API costs from the reconstruction, pauses between skills caused by Astra's reasoning latency, finger servos overheating during extended operation, and a local stack that requires an RTX 4090 laptop GPU. The GitHub repository currently holds only a README and figures, says "Code coming soon," and has no license.

The question: is the middle layer still needed?

For the past two years, the mainstream humanoid approach has used three layers: a large model on top that understands instructions and breaks down tasks; a specially trained VLA model in the middle that turns "pick up that cup" into concrete motions; and a controller at the bottom handling joints and balance. The middle layer is the expensive one, because it needs large amounts of collected robot demonstration data. HomeBody asks the question directly: **if the model on top is smart enough, can it skip the middle layer and call a set of prebuilt skills?** The skills number just five: navigate, pick, place, open a drawer, and take something out of a drawer, running on an off-the-shelf whole-body controller and classical motion planning. GPT Astra is called remotely and picks skills step by step as if calling tools; execution results come back to it, and it replans when something fails.

What may really do the work is memory

A smart model alone isn't enough. Once the robot turns around, what it just saw is out of view; in a larger room, it loses track of where it is. HomeBody first has the robot walk the kitchen, collecting video, depth and LiDAR data, reconstructs a digital twin in Isaac Sim, and writes "what is where" into a persistent spatial memory. Later, given an instruction like "get my medicine" that doesn't say where it is, the robot can check its memory, know the medicine is in a particular drawer, and go open it. **That step turns "can understand what it sees" into "can remember what it saw," and for room-wide chores like housework the latter is often the real bottleneck.**

Why the limitations are worth reading in full

This is a demonstration, not an evaluation result. The project page reports no success rates and no systematic comparison against VLA approaches, so "the middle layer isn't needed" can for now only be read as "it wasn't needed in these demos." The authors' own list of limitations is specific: every new environment needs a reconstruction first, with setup time and API costs; Astra has to think at each step, leaving pauses between skills; the finger servos overheat over long runs, limiting task length; and the local stack needs an RTX 4090 laptop GPU. **Put together, it is a long way from something you could bring home and use; it reads more like a feasibility test of the "frontier model + a few skills + spatial memory" route.** One more clarification: some coverage says the code has been open-sourced, but the GitHub repository currently holds only a README and figures, says "Code coming soon," and has no license file.

What this means for readers

For robotics and embodied-AI teams, the interesting part is the architectural trade-off: replacing a VLA that needs large amounts of demonstration data with a frontier model's reasoning plus a small set of reliable skills. If that route works, the data barrier drops sharply; the cost is that every step depends on the latency and price of a remote large model. Whether it can be reproduced is best judged once the code and success rates are public.

via: HomeBody project page, HomeBody GitHub repository, The Decoder