Featured image of post Figure AI Unveils Helix 2.5 with 56% Success Rate: Can Scaling Laws Deliver Real Robot Utility?

Figure AI Unveils Helix 2.5 with 56% Success Rate: Can Scaling Laws Deliver Real Robot Utility?

Figure AI launches Helix 2.5 achieving 56% success rate in 30 unseen homes, sparking debate over robot generalization and scaling laws.

Figure AI Launches Helix 2.5: First Large-Scale Robot Deployment in Real Homes

Figure AI Launches Helix 2.5: First Large-Scale Robot Deployment in Real Homes
Figure AI Launches Helix 2.5: First Large-Scale Robot Deployment in Real Homes|News screenshot

  • Release Date: September 2026 (date not specified, inferred from news publication)
  • New Version: Helix 2.5
  • Test Environment: 30 real residential homes in Bay Area (no prior data collection)
  • Test Tasks: Living room cleanup (13-15 toys into basket), towel folding (all 4 folded), bed making (position, direction, coverage requirements)
  • Test Count: 420 total, 140 per task
  • Core Metric: 56% full task success rate (237/420); bed making 67%, towel folding 62%, toy cleanup 40%
  • Zero-Shot Capability: Model not fine-tuned for specific homes or items; fixed version used throughout

Zero-Shot Generalization: Contrasting Claims and Critiques

Zero-Shot Generalization: Contrasting Claims and Critiques
Zero-Shot Generalization: Contrasting Claims and Critiques|News screenshot

Helix 2.5 represents a shift in learning approach. Previous Helix 02 trained on environment-specific robot data, limiting transferability. Helix 2.5 first pretrains on Figure AI’s human behavior dataset Index, then fine-tunes for three household tasks. The breakthrough allows deployment without any adaptation to住宅 layout or specific items (e.g., towel material, bed height, toy placement)—true zero-shot generalization to unseen homes.

Figure AI claims this demonstrates the first measured “human-to-robot transfer scaling law.” When pretraining data increases from 1x to 8x, action prediction error declines steadily. Control models without Index pretraining achieved only 9% success versus 56% for Helix 2.5.

The key tension lies in reliability. Sunday Robotics co-founder Tony Zhao (ACT/ALOHA system author) argues: “Useful work = generalization + reliability.” Nearly half the failures (183/420) make Helix 2.5 unsuitable for unsupervised deployment. Sunday’s ACT-2 achieved 99.1% success (778/785) on single-item folding, though Figure’s towel test requires concurrent processing of four items.

Stealth founder Nikolai Ensslen challenges the fundamental approach: “Generalization means working every time—Helix 2.5 has not achieved this.” He contends the success rate proves the imitation learning system cannot truly generalize.

Scaling Law: Figure AI’s Bet and Industry Divide

Scaling Law: Figure AI’s Bet and Industry Divide
Scaling Law: Figure AI’s Bet and Industry Divide|News screenshot

Figure AI now adopts the scaling law narrative—previously the domain of large language models. Robotics has lacked predictable performance trajectories: data is costly, environments vary, and hardware heterogeneity means most capabilities must be rebuilt per scenario.

Helix 2.5’s technical blog asserts human-to-robot transfer scaling, suggesting data, model size, and compute can linearly improve performance—just as with LLMs. If valid, current 56% success merely reflects insufficient scale.

To validate this, Figure AI commitment includes $3.5 billion in initial compute with Nscale, targeting up to 100,000 NVIDIA Vera Rubin GPUs. Valued at $3.9 billion (world’s most valuable unnamed humanoid robotics company), Figure has produced over 350 Figure 03 units with hourly production capacity.

MetricFigure AI Helix 2.5Sunday Robotics ACT-2 (Reference Comparison)
Test Environment30 real homes (zero-shot)Lab-controlled environment
Success Rate56% (237/420)99.1% (778/785)
Task UnitEntire home task (multi-item continuous)Single-item folding
Generalization MethodZero-shot household transferTask-specific fine-tuning
Data ScaleMulti-version Index datasetNot disclosed

Who Should Care?

Who Should Care?
Who Should Care?|News screenshot

  • Enterprise buyers (e.g., senior care facilities, hotels) may track long-term progress but should wait for 70%+ success before piloting; individual consumers should delay until continuous-task failure rates drop below 30%.
  • Researchers and developers: Helix 2.5’s zero-shot test framework and published metrics offer valuable research benchmarks; monitor potential Index dataset openness.

Final Thoughts

Helix 2.5 matters less for its 56% figure than for advancing the industry’s debate from “can it demo” to “can it be trusted.” If scaling holds, 56% becomes the springboard to 90%; if success plateaus, the 44% failures represent not scale issues but fundamental methodological limits.