Humanoid robotics has spent a decade producing results that are individually astonishing and collectively useless for planning. A video of a robot folding laundry tells you nothing about what happens in the next house, or the one after that. On September 17, Figure published Helix 2.5, and the announcement deserves attention not primarily for the demo, a humanoid tidying living rooms, folding towels and making beds in 30 Bay Area homes it had never seen, but for two numbers buried under it: a controlled ablation showing that pretraining alone took zero-shot success from 9% to 56%, and a scaling experiment clean enough that Figure could forecast its largest training run’s loss to four decimal places before running it.

One of those numbers is a robotics result. The other is an industry-changing result, and it is worth being precise about the difference.

What Was Actually Demonstrated

The setup: Figure pretrained Helix 2.5 on Index, its proprietary dataset of recorded human behavior, then adapted that single foundation model to three whole-body behaviors. The behaviors were chosen to span the discipline’s hard problems at once: tidying a living room (locomotion plus rigid manipulation), folding towels (deformable manipulation), and making a bed (bimanual coordination plus whole-body reaching). The robot then walked into 30 real homes it had never encountered, with objects it had never touched, and worked.

The zero-shot claim is more disciplined than most. No data was collected in any evaluation home. No evaluation toy, towel or bedding appeared in the training data, a constraint verified by an automated check followed by human review. Every behavior ran from a single fixed checkpoint across all 30 homes: no per-site fine-tuning, no environment-specific adaptation, and no cherry-picking of checkpoints based on evaluation performance. Trials were graded blind against rubrics fixed in advance, success required completing the entire task with no partial credit, and any human safety intervention aborted the rollout and counted as a failure. The rubrics themselves are strict enough to be worth reading: every one of 13 to 15 scattered toys in the basket, towels folded with all four corners within an inch of touching, both pillows and both comforter corners placed and smoothed, with per-object timeouts that abort slow trials.

Helix 02, Figure’s previous system, already demonstrated whole-body coordination over long horizons, from unloading dishwashers to running a logistics task autonomously for 200 hours. But those policies learned from data collected in the environments where the robots would operate. That is the deployment model every serious humanoid program has used until now: instrument the site, collect site data, tune, deploy. Helix 2.5’s wager is that the model arrives knowing enough about physical common sense that the site itself contributes nothing.

The Ablation That Matters

The most important experiment in the release is also the least cinematic. Figure trained two policies on identical task-specification data. One started from random initialization; the other started from the Index-pretrained Helix 2.5 weights. Architecture, optimization, hyperparameters, downstream data and evaluation were all held fixed. The only variable was pretraining.

The from-scratch policy succeeded on 9% of zero-shot trials. The Index-pretrained policy succeeded on 56%. Success, again, meant the whole task, every toy, every towel, the entire bed. That 47-point gap is a direct measurement of what broad human-experience pretraining contributes to a humanoid, isolated from every confound Figure could control. The company also notes that no single evaluation task makes up more than 1.90% of the Index pretraining data, which matters because it rules out the boring explanation: the model did not simply memorize these specific chores at scale.

For a field that has spent three years asserting that foundation-model pretraining would transfer to manipulation, this is the cleanest evidence yet published by a frontier lab. It converts a thesis into a measurement.

Fifty-six percent full-task success is simultaneously impressive and nowhere near a product. Inverted, it means the robot fails to complete nearly half of the tasks it attempts in houses full of its own company’s sensors and graders. The honest reading: pretraining solved the generalization gap in these settings, and what remains is the gap between 56% and whatever number a paying household tolerates, which for a machine folding your towels is likely north of 95% with graceful failure modes below it.

The Scaling Law Is the Real Story

Tucked beneath the headline results is the finding with the longest half-life. Figure trained four models on nested subsets of Index spanning an eightfold increase in pretraining data, holding model size and downstream training fixed. Held-out robot-action-prediction loss fell predictably with each doubling. The relationship was smooth enough that, using only the smaller runs, Figure forecast its largest run’s test loss to four decimal places before training began, with forecasting error of just 0.54% of the variation across the full data range. Figure describes it as the first human-to-robot transfer scaling law measured on a humanoid.

To understand why this matters more than any single demo, look at what scaling laws did for language models. Before Kaplan and Chinchilla, large model training was gambling: you spent millions, crossed your fingers, and evaluated after the fact. Once loss became forecastable from small runs, training runs became capital projects with predictable returns. That predictability is arguably the single enabling fact behind the tens of billions of dollars subsequently invested in frontier AI. Investors do not fund intelligence; they fund predictable improvement.

Humanoid robotics has never had that. Every capability jump has been a surprise, announced after the fact in a video. If Figure’s result holds, the next doubling of Index data has a price tag attached to an expected capability delta, which means the 100,000-GPU Nscale compute commitment starting in late 2027 stops being a moonshot and starts looking like a priced instrument. That is what “time to scale up” at the end of Figure’s post actually means: the curve is smooth, and they can afford the next several doublings.

Two caveats keep this from being a solved problem. First, the scaling law is measured on action-prediction loss, not task success. Loss curves that forecast cleanly do not automatically tell you when the 56% becomes 90%; the loss-to-success mapping is the unknown that decides whether this is Chinchilla or a mirage. Second, smooth scaling in data does not preclude capability plateaus that only appear at larger scale. The LLM precedent cuts both ways.

Data Economics: The Flywheel and Its Marginal Cost

The scaling law also sharpens the strategic value of Index itself. Unlike LLM labs, which pretrained on the internet for free, Figure manufactures its data: a network of paid contributors recording structured human activity, now generating roughly 35 minutes of new human experience per second, with more than 69,000 weekly active contributors reported as of early September. At that rate Index ingests on the order of 840 hours of experience per day, every day, against the $3.5 billion of committed compute Figure is pointing at Helix training.

This is a moat with a cost structure, and the scaling law is what tells you the spend converts. Each doubling of Index is purchasable: more contributors, more recording hours, more categories. If downstream capability keeps tracking the curve, Figure has effectively built an assembly line for capability, in an industry where rivals are still hand-building it. The competitive question for everyone else, from Tesla to Physical Intelligence to 1X, is whether they can assemble a comparable data engine or whether they will be buying capability at retail while Figure produces it wholesale.

The Home Race: Two Opposite Bets on the Same Market

Helix 2.5 lands the same season as the first consumer humanoid deliveries. 1X is shipping Neo to its first US customers at $20,000 or $499 per month, and the machine’s open secret is that its early autonomy is partly human-assisted: 1X operators can watch and intervene through the robot’s own cameras. Whatever the privacy discourse around that arrangement, it is a coherent strategy: ship now, let teleoperation carry the failure cases, and let in-home data accumulate.

Figure’s Helix 2.5 is the opposite bet. No teleoperator in the loop, no data collected in the deployment home, no adaptation after arrival. It is slower to market and scientifically riskier, but it attacks the exact cost that breaks the 1X model at scale: per-site human labor. A home humanoid that requires unseen-household competence to be manufactured on site, by teleop or by data collection, has a marginal cost per customer that never reaches zero. A model that generalizes from pooled human experience has a marginal cost per customer that approaches the hardware.

The likely equilibrium is sequential. Teleop-assisted robots ship first and generate revenue and data; pretrained generalization quietly lowers the teleop fraction until the training wheels come off. The metric that will decide the home market is not any single demo but the slope of that teleop-fraction curve, and Helix 2.5 is the first published evidence that the slope can be engineered deliberately.

The Skeptic’s Ledger

An analyst’s confidence should be rationed precisely. The evaluation was designed, run and graded by Figure itself, with no third-party replication yet. The 30 homes are Bay Area homes: modern construction, generous layouts by global standards, and a population unusually tolerant of robots in their living rooms. The scenes were staged to spec, 13 to 15 toys scattered by a procedure, not the chaos of a household with toddlers and pets. Timing data, cycle times and intervention rates beyond the pass/fail rubric were not published. The self-correction behaviors Figure highlights, stepping back, repositioning, walking around a bed to fix a fold, are qualitative, and nobody outside the company has stress-tested them at scale.

And 56% remains 56%. As Humanoids Daily notes, the next commercial question is whether learned behaviors survive repeated work over days, interrupted chores and occupied rooms, conditions no rubric fully anticipates.

None of this diminishes the result. It defines what kind of result it is: a research milestone with industrial implications, not a shipping product.

What to Watch

First, replication. Zero-shot whole-body generalization at this scope will not stay unclaimed; expect competing evaluations from 1X, Tesla and the VLA laboratories within quarters, and expect the definitional haggling over “zero-shot” to become the field’s new sideshow. Second, the loss-to-success curve. The first lab to publish task success as a function of pretraining scale, not just loss, converts its scaling law into a roadmap. Third, the compute supercycle. Figure’s Nscale deployment begins in late 2027; if Helix’s improvement curve holds through 2026, that build-out arrives pre-justified. Fourth, the teleop fraction on shipped home robots, the single number that quietly tracks how much of this is real.

The robotics industry has had miracles before. It has never had a forecast. Figure just published the first one, and the companies that dismiss it as a laundry demo are making the same mistake the ones who dismissed GPT-2 made: watching the task, and missing the curve.