Figure AI Launches Helix 2.5: First Large-Scale Robot Deployment in Real Homes

- Release Date: September 2026 (date not specified, inferred from news publication)
- New Version: Helix 2.5
- Test Environment: 30 real residential homes in Bay Area (no prior data collection)
- Test Tasks: Living room cleanup (13-15 toys into basket), towel folding (all 4 folded), bed making (position, direction, coverage requirements)
- Test Count: 420 total, 140 per task
- Core Metric: 56% full task success rate (237/420); bed making 67%, towel folding 62%, toy cleanup 40%
- Zero-Shot Capability: Model not fine-tuned for specific homes or items; fixed version used throughout
Zero-Shot Generalization: Contrasting Claims and Critiques

Helix 2.5 represents a shift in learning approach. Previous Helix 02 trained on environment-specific robot data, limiting transferability. Helix 2.5 first pretrains on Figure AI’s human behavior dataset Index, then fine-tunes for three household tasks. The breakthrough allows deployment without any adaptation to住宅 layout or specific items (e.g., towel material, bed height, toy placement)—true zero-shot generalization to unseen homes.
Figure AI claims this demonstrates the first measured “human-to-robot transfer scaling law.” When pretraining data increases from 1x to 8x, action prediction error declines steadily. Control models without Index pretraining achieved only 9% success versus 56% for Helix 2.5.
The key tension lies in reliability. Sunday Robotics co-founder Tony Zhao (ACT/ALOHA system author) argues: “Useful work = generalization + reliability.” Nearly half the failures (183/420) make Helix 2.5 unsuitable for unsupervised deployment. Sunday’s ACT-2 achieved 99.1% success (778/785) on single-item folding, though Figure’s towel test requires concurrent processing of four items.
Stealth founder Nikolai Ensslen challenges the fundamental approach: “Generalization means working every time—Helix 2.5 has not achieved this.” He contends the success rate proves the imitation learning system cannot truly generalize.
Scaling Law: Figure AI’s Bet and Industry Divide

Figure AI now adopts the scaling law narrative—previously the domain of large language models. Robotics has lacked predictable performance trajectories: data is costly, environments vary, and hardware heterogeneity means most capabilities must be rebuilt per scenario.
Helix 2.5’s technical blog asserts human-to-robot transfer scaling, suggesting data, model size, and compute can linearly improve performance—just as with LLMs. If valid, current 56% success merely reflects insufficient scale.
To validate this, Figure AI commitment includes $3.5 billion in initial compute with Nscale, targeting up to 100,000 NVIDIA Vera Rubin GPUs. Valued at $3.9 billion (world’s most valuable unnamed humanoid robotics company), Figure has produced over 350 Figure 03 units with hourly production capacity.
| Metric | Figure AI Helix 2.5 | Sunday Robotics ACT-2 (Reference Comparison) |
|---|---|---|
| Test Environment | 30 real homes (zero-shot) | Lab-controlled environment |
| Success Rate | 56% (237/420) | 99.1% (778/785) |
| Task Unit | Entire home task (multi-item continuous) | Single-item folding |
| Generalization Method | Zero-shot household transfer | Task-specific fine-tuning |
| Data Scale | Multi-version Index dataset | Not disclosed |
Who Should Care?

- Enterprise buyers (e.g., senior care facilities, hotels) may track long-term progress but should wait for 70%+ success before piloting; individual consumers should delay until continuous-task failure rates drop below 30%.
- Researchers and developers: Helix 2.5’s zero-shot test framework and published metrics offer valuable research benchmarks; monitor potential Index dataset openness.
Final Thoughts
Helix 2.5 matters less for its 56% figure than for advancing the industry’s debate from “can it demo” to “can it be trusted.” If scaling holds, 56% becomes the springboard to 90%; if success plateaus, the 44% failures represent not scale issues but fundamental methodological limits.
