Figure Releases Helix 2.5, Robots Zero-Shot Tidy 30 Unseen Homes
Figure has released Helix 2.5, a new model that allows robots to perform household chores autonomously. For the test, the company rented 30 houses in the San Francisco Bay Area. The robots did not undergo additional training; once they arrived at each location, they began tidying the home on their own.
Helix 2.5 was built to answer a single question: can a humanoid enter a home it has never seen before and immediately work autonomously? According to Figure, the way this question is framed reveals more about what the company wants to validate than any technical benchmark does.
Previously, robots operating in new environments usually needed pre-built maps, collected data, or scene-specific fine-tuning before they could work reliably. Helix 2.5 is intended to test zero-shot generalization: the robot enters an unfamiliar household, has no data from that environment, and starts working directly.
The 30-home test is notable because it does not rely on carefully staged demonstrations in a single lab. Each home presents different furniture, layouts, objects, and lighting. Zero-shot operation means the robot cannot rely on memorized maps or pre-collected trajectories, so success depends on transferring general skills to new physical settings.
A single foundation model produces three behaviors. Figure pretrained Helix 2.5 on Index, a global dataset of human behavior. From this one foundation model, the robot generated three distinct behaviors: tidying a living room, folding towels, and making a bed. Together, these tasks cover locomotion, rigid and deformable object manipulation, two-handed coordination, and active perception.
The three behaviors are not trivial. Tidying a living room involves navigation and object rearrangement; folding towels involves deformable objects; making a bed requires bimanual coordination and active perception. By adapting one foundation model to all three, Figure aims to show that a general policy can serve multiple household tasks rather than requiring separate systems for each one.
Figure then brought the robot into 30 households across the Bay Area without collecting any prior data from those homes. During testing, one Index-pretrained foundation model was adapted to the three behaviors. The key design choice is that the foundation model stays general, while only small task-specific adjustments are made. Instead of training a separate model from scratch for every chore, Figure reuses the same pretrained base.
The measured success rate rose from 8% to 56%. According to the reported results, Index pretraining significantly improved the humanoid's zero-shot whole-body generalization. With task-specific data, architecture, training, and evaluation all held constant, the pretraining variable lifted task success from 8% to 56%. The 48-percentage-point gain represents the generalization benefit contributed by pretraining data.
Compared with the previous neural network, Helix 02, Helix 2.5 needed only half the task-specific data to generalize the behavior to 30 previously unseen environments.
Figure says Helix 2.5 proves for the first time that whole-body intelligence can learn from human experience and transfer to new scenes without being rebuilt each time. A 56% success rate also means 44% of tasks still fail. Under zero-shot conditions, however, that figure already indicates a substantial improvement in the foundation model's ability to generalize to new environments.
The next questions are whether the success rate remains stable after longer stays in real homes and whether household tasks beyond the three tested behaviors can generalize in the same way.