The question: is the middle layer still needed?
For the past two years, the mainstream humanoid approach has used three layers: a large model on top that understands instructions and breaks down tasks; a specially trained VLA model in the middle that turns "pick up that cup" into concrete motions; and a controller at the bottom handling joints and balance. The middle layer is the expensive one, because it needs large amounts of collected robot demonstration data. HomeBody asks the question directly: **if the model on top is smart enough, can it skip the middle layer and call a set of prebuilt skills?** The skills number just five: navigate, pick, place, open a drawer, and take something out of a drawer, running on an off-the-shelf whole-body controller and classical motion planning. GPT Astra is called remotely and picks skills step by step as if calling tools; execution results come back to it, and it replans when something fails.
What may really do the work is memory
A smart model alone isn't enough. Once the robot turns around, what it just saw is out of view; in a larger room, it loses track of where it is. HomeBody first has the robot walk the kitchen, collecting video, depth and LiDAR data, reconstructs a digital twin in Isaac Sim, and writes "what is where" into a persistent spatial memory. Later, given an instruction like "get my medicine" that doesn't say where it is, the robot can check its memory, know the medicine is in a particular drawer, and go open it. **That step turns "can understand what it sees" into "can remember what it saw," and for room-wide chores like housework the latter is often the real bottleneck.**
Why the limitations are worth reading in full
This is a demonstration, not an evaluation result. The project page reports no success rates and no systematic comparison against VLA approaches, so "the middle layer isn't needed" can for now only be read as "it wasn't needed in these demos." The authors' own list of limitations is specific: every new environment needs a reconstruction first, with setup time and API costs; Astra has to think at each step, leaving pauses between skills; the finger servos overheat over long runs, limiting task length; and the local stack needs an RTX 4090 laptop GPU. **Put together, it is a long way from something you could bring home and use; it reads more like a feasibility test of the "frontier model + a few skills + spatial memory" route.** One more clarification: some coverage says the code has been open-sourced, but the GitHub repository currently holds only a README and figures, says "Code coming soon," and has no license file.
What this means for readers
For robotics and embodied-AI teams, the interesting part is the architectural trade-off: replacing a VLA that needs large amounts of demonstration data with a frontier model's reasoning plus a small set of reliable skills. If that route works, the data barrier drops sharply; the cost is that every step depends on the latency and price of a remote large model. Whether it can be reproduced is best judged once the code and success rates are public.
via: HomeBody project page, HomeBody GitHub repository, The Decoder