Every field has its eras, and you can usually date them by how new capabilities are installed. Language modeling had its BERT era: capable models that still needed task-specific data collection and fine-tuning for every new job. Then GPT-3 arrived and the paradigm flipped. You did not retrain anything. You described what you wanted in a prompt, and the model, having internalized how to learn from examples during pretraining, simply complied. That single transition, from fine-tuning to in-context learning, is what turned language models from research infrastructure into a consumer industry.
Robotics, by the explicit argument of the people building its frontier models, has been stuck in the BERT era. A humanoid can learn a complex behavior, but deploying it somewhere new has meant collecting tens to hundreds of hours of data in deployment conditions and fine-tuning a specialist policy for each task, in each place. September 2026 delivered three independent results, from Figure, Skild AI and Generalist, that attack this bottleneck from different directions and converge on the same conclusion: the cost of teaching a robot a new task is collapsing from weeks of data collection to minutes of demonstration. For an industry whose unit economics were supposed to die on data-collection costs, this is the most important technical news of the year.
Figure’s Helix 2.5: the zero-shot home experiment
The headline result is Helix 2.5, which Figure calls the most advanced neural network it has built. The experiment’s design is unusually clean for this industry. Figure pretrained Helix 2.5 on Index, its global-scale dataset of human behavior, then adapted that single foundation model to exactly three behaviors: tidying living rooms, folding towels, and making beds. Then the company took its humanoid into 30 Bay Area homes where it had collected zero data. No environmental familiarization, no fine-tuning, no adaptation to the specific homes or objects. One fixed checkpoint ran in all 30 houses, and Figure’s grading gave no partial credit: success meant every toy tidied, every towel folded, the whole bed made.
The numbers are the first credible measurement of whole-body generalization on a humanoid. A control policy trained from scratch on identical task data succeeded on 9% of zero-shot trials. The Index-pretrained policy succeeded on 56%, a better-than-six-fold improvement with pretraining as the only experimental variable. Figure verified that no evaluation home, toy, towel or bedding item appeared in the training data, and no single evaluation task made up more than 1.90% of the Index pretraining set. The company claims, credibly given the evaluation protocol, that this is the first demonstration of zero-shot whole-body generalization at this scope on a humanoid.
Two secondary results matter as much as the headline number. First, behavior specification got cheaper: Helix 2.5 used half as much task-specific data as a comparable Helix 02 behavior while generalizing across 30 unseen homes instead of the one environment where its data was collected. Second, and more consequential for everyone else in the field, Figure reports the first human-to-humanoid transfer scaling law: across an 8x increase in Index pretraining data, downstream robot-action prediction improved smoothly enough that the company forecast its largest run’s test loss to four decimal places before training began, with forecasting error of just 0.54% of the variation across the data range.
That scaling law deserves emphasis. Language model economics transformed when labs could predict the returns on the next 8x of compute before spending it. If human video reliably transfers to humanoid control in a predictable way, then data acquisition stops being a research gamble and becomes a budgeting exercise. Index, recall, is teleoperated robot data mixed with egocentric human video captured at civilization scale. The claim Figure is now equipped to defend to investors is that more cameras watching more humans doing chores translates, at a calculable rate, into more capable robots.
Skild S1: the video prompt as programming interface
Skild AI’s S1 takes the paradigm one step further and makes the boldest framing of the month. In its own words, robotics is stuck in the BERT era, and the point of pretraining is to enable robotics’ equivalent of the GPT-3 shift: in-context learning. The interface is disarmingly simple. Show S1 a video demonstration of a task, short or long, seen or unseen. The demonstration enters the model’s context window, the policy translates it to its own embodiment and current scene, and the robot executes. No fine-tuning, no post-training, no weight updates of any kind.
The demonstrated capabilities target the two axes Skild identifies as the real test of in-context learning: skill novelty and task horizon. On the novelty axis, S1 performed tasks absent from its pretraining distribution, including digging into soil to make room for a plant, pressing a coffee filter into a funnel, and flipping a pancake. On the horizon axis, unseen tasks like plant potting, pancake cooking, pour-over coffee and kit assembly ran up to ten minutes long, spanning dozens of manipulation steps, driven by a single demonstration each. The company’s plant-potting timeline is a quiet rebuke of the old deployment workflow: materials arrived at 8:54 PM, most of the elapsed time went to moving furniture, and the gap from demonstration to autonomous execution was minutes, not the weeks of teleoperation a task-specific policy would have required.
On unseen tasks, Skild reports 66% for the video-prompted S1 against 9% for a comparable language-prompted model, both trained on 100,000 hours of data. One caveat belongs in every analyst’s notebook: as Turing Post’s September robotics review notes, that metric measures cumulative per-step success with human intervention recovering failures, mostly for the baseline, so it is not a 66% fully autonomous completion rate. Even discounted, the 7x gap between showing and telling is the finding that will redirect research budgets: for long-horizon, nuanced manipulation, language is the wrong prompt. Humans do not learn to fold a fitted sheet from a sentence, and neither, it turns out, do robots.
Skild also published the clearest articulation this month of the data problem underlying everything. No single data source wins on the three axes that matter: hardware proximity, diversity and scalability. Teleoperation sits closest to the robot and scales worst. Egocentric video scales best and has the largest domain gap. Skild’s answer is to scale all sources in-house and combine them strategically, which is a capital-intensive answer, and one reason its S1 runs on NVIDIA’s AI infrastructure at a scale most labs cannot match.
Generalist GEN-1.5: the middle path
The third result calibrates the trend. Generalist’s GEN-1.5 accepts demonstrations of 3 to 12 seconds in its context and attempts tasks like opening a zipper or unscrewing a jar without weight updates, succeeding 59% of the time on average across ten tasks. After just ten fine-tuning steps using five minutes of data per task, success rises to 83%. Skild positions itself at the extreme end of the spectrum: unseen, ten-minute tasks from a single video. GEN-1.5 covers shorter, simpler tasks, but with an even lighter adaptation burden. Read together, the two results bracket a frontier that is moving in one direction: the amount of task-specific data needed to deploy a useful behavior is falling from hundreds of hours toward minutes, approaching zero for tasks near the training distribution.
Why this restructures the humanoid market
The business implications are larger than the benchmark improvements, because data collection has been the dominant deployment cost and the standard bear case against humanoids. Consider what the old model implied for Figure’s home-robotics ambitions or for warehouse deployments: a fleet of thousands of robots each requiring per-site, per-task data collection is a services business wearing a hardware costume, with margins to match. The new model implies something closer to software economics. Figure’s own numbers make the point: half the task data, 30 times the environmental scope, from one pretrained base model.
Three market consequences stand out.
First, the moat migrates from deployment logistics to pretraining data engines. If Index-style human video and Skild-style multi-source pipelines are what convert into generalization, the strategic asset is the data flywheel: Index for Figure, Skild’s in-house blend of teleoperation, UMI, egocentric video and simulation. Figure’s scaling law turns that asset into a forecastable return, which is precisely the story that attracts the hundreds of millions these programs consume. Tesla’s Optimus program, with its fleet learning inside its own factories, and Physical Intelligence’s π-family models are running versions of the same race.
Second, the addressable task list grows without re-engineering. A model that composes unseen skills from context, as S1 did across ten-minute horizons, does not need every customer use case enumerated and trained in advance. That is the difference between a product that must be sold to exactly the workflows it was trained for and one that adapts on site, and it is the prerequisite for the household market every consumer humanoid company is targeting. 1X has said the newest NEO should ship to consumers by the end of 2026 operating autonomously by default; results like these are what make such claims discussable rather than dismissible.
Third, the talent and capital destination shifts again. The same week these results landed, reporting consolidated around Sam Altman’s September 3 confirmation on the Sources podcast that OpenAI will build humanoids and other form factors, with infrastructure and manufacturing first and household robots a longer-term ambition, and Forbes counted 19 robotics openings on OpenAI’s careers page spanning actuators, firmware and prototyping. When the company with the deepest experience in scaling laws and in-context learning formally enters the robot business, it is not chasing the hardware. It is betting that the LLM playbook, pretrain broadly, prompt narrowly, scale predictably, now applies to bodies.
The hardware tier is internalizing the same lesson from the opposite direction. Agility’s Digit 5 launch, built on 65,000 operating hours of production data from its predecessor, is a hardware-scale argument: reliability compounds through fleet hours. Figure’s scaling law is the software-scale version: capability compounds through human-witness hours. The companies that win the next phase will hold both dials.
The honest caveats
Three qualifications keep this month in perspective. Helix 2.5’s 56% is a start, not a finish line: nearly half of zero-shot trials still failed somewhere in a full task, in homes selected for three well-defined behaviors. The tasks themselves, toys, towels, bedding, are structured and reset-friendly compared with the open-ended mess of real domestic work or the cycle-time discipline of paid warehouse labor. And evaluation standards differ across labs: Figure’s no-partial-credit, whole-task criterion is stricter than per-step metrics, which is exactly why cross-company comparisons remain hazardous and why independent evaluation of these claims, ideally on standardized task suites, is the field’s next pressing need.
None of which softens the conclusion. In language AI, the gap between BERT-era fine-tuning and GPT-3-era prompting was crossed in roughly two years, and the industry that existed on the far side of that crossing bore little resemblance to the one before it. September 2026 is when robotics produced its first credible evidence of the same crossing: Figure showed the generalization, Skild showed the interface, Generalist showed the adaptation curve, and Figure’s scaling law showed the forecastability. The robots are still slow and still fail too often. But the way we teach them changed this month, and everything downstream of teaching, deployment cost, task breadth, fleet margins, market size, is now renegotiable.
Sources: Figure: Helix 2.5 zero-shot 30-home generalization · Skild AI: Introducing S1, in-context learning for robotics · Generalist: GEN-1.5 · Turing Post: Robotics September 2026 roundup · Forbes: OpenAI is making a humanoid robot