The humanoid robotics industry has a public relations problem and a spreadsheet problem, and only one of them is improving quickly. The PR problem is the demo reel: choreographed videos that keep getting more impressive while paid deployments keep getting harder to find. The spreadsheet problem is quieter and more decisive: the cost of teaching a robot a new task. Not the cost of the actuators or the battery, but the cost of the data collection, the fine-tuning runs, the integration engineering, and the on-site hand-holding that turns a capable machine into a productive one.

Two announcements this August, from opposite ends of the sector, attack that line item from both sides. Generalist, the robot foundation model company, published results for GEN-1.5, a model that learns new physical tasks from a single demonstration in seconds, without a single gradient update. Persona AI, the humanoid startup founded by former NASA and IHMC roboticists, laid out in detail with IEEE Spectrum why it has abandoned the chase for general-purpose deployment volume and is instead building a machine that welds ship hulls in Korean shipyards. One story is about collapsing the marginal cost of a new skill. The other is about maximizing the value of a single skill. Together they sketch the most credible route to humanoid economics that does not depend on a 460 percent IPO pop.

GEN-1.5: the GPT-3 moment for physical skills

Start with the model, because the result is genuinely strange. GEN-1.5 is a large multimodal model that consumes video, proprioception, language, and other sensor streams, holds 30 seconds of memory in its context window, and outputs action trajectories at 100 Hz. None of that architecture is unusual in 2026. What is unusual is what emerged from pretraining without anyone asking for it.

The team at Generalist reports that across 10 diverse manipulation tasks, including twisting lids off glass jars, working zippers, and retrieving money from wallets, the pretrained model achieves 59 percent average success (plus or minus 10 percent standard deviation) when prompted with a single demonstration of 3 to 12 seconds, inserted directly into its context window. No fine-tuning. No gradient steps. No task-specific training of any kind. The demonstration, which Generalist calls a “physical prompt,” can be recorded by a human wearing handheld grippers or by the robot itself. With 10 gradient steps on 5 minutes of data, roughly 50 demonstrations, average success climbs to 83 percent (plus or minus 9 percent).

The deliberate framing is the GPT-3 analogy. When OpenAI reported GPT-3’s results in 2020, the headline capability was in-context learning: roughly 45 percent average accuracy on language tasks from a single example, rising to about 65 percent with around 100 examples, with no weight updates. Generalist explicitly positions GEN-1.5 as the embodied analogue: the first model, to their knowledge, in which one-shot and few-shot learning of closed-loop physical skills has emerged at scale. The lineage they trace goes back to teach-by-guiding in the 1954 Unimate patent and MIT’s 1970 Copy Demo, which is a useful reminder that robotics has chased this capability for seventy years.

Three details matter more than the headline number.

First, the capability was not engineered. Generalist states there were no architectural changes to promote in-context learning, no meta-learning loop pushing the model toward fast adaptation, no auxiliary objectives encouraging improvisation. The behavior emerged from pretraining on large amounts of physical interaction data captured in homes, warehouses, and factories, data the model has now been absorbing for more than eight months of continuous training. The company previously reported predictable scaling laws leading up to GEN-0 nine months ago, and GEN-1 five months ago reached 99 percent-plus success on tasks after post-training. The trend they describe is that each generation needs orders of magnitude less task-specific data: fine-tuning costs fell from hundreds of thousands of gradient steps, to tens, to 1 to 10 steps on one to five minutes of data. Ten steps alter the model’s weights by less than 0.15 percent, suggesting adaptation reconfigures existing knowledge rather than building new representations.

Second, physical prompts compose. Place two independent demonstrations in the context window, each recorded separately with no transition between them, and GEN-1.5 chains them into one continuous behavior, generating the intermediate repositioning, regrasping, and error recovery that appears in neither demonstration. Generalist draws the obvious implication: long-horizon behaviors could be assembled from a library of short, reusable physical prompts, the way language prompts are chained from instructions. They call this physical prompt engineering, and it comes with a drag-and-drop interface for selecting demonstrations.

Third, the prompts transfer across boundaries that were previously hard walls. A demonstration recorded entirely in simulation, from a scripted policy or a teleoperated simulated robot, works as a prompt for the real robot, even though GEN-1.5’s pretraining contains zero simulation data. In some cases a human can demonstrate a task with their own bare hands, in view of the robot’s cameras, and the model reproduces it with the robot’s grippers. The model also improvises: fine-tuned on brushing a block into a bowl, it used a banana as a makeshift brush and, when handed a dustpan, invented an entirely new contact sequence, lifting the block and dumping it into the bowl.

The honest caveats are stated by Generalist themselves. The tasks are simple and short-horizon. The one-shot success rate is modest. In-context skills are more brittle than fine-tuned ones. The company also concedes, in a note worth keeping whenever reading results of this kind, that claims of “no relevant pretraining data” can only ever be made to the best of their knowledge. This is a research milestone, not a product. But the direction is unambiguous: the marginal cost of teaching this class of model a new skill is collapsing toward the cost of showing it once.

Persona AI: economics as a design constraint

Now cross the Pacific and walk into a shipyard. Persona AI was founded in 2024 by Nicolaus Radford, who led NASA’s Valkyrie program at Johnson Space Center and founded Nauticus Robotics, and Jerry Pratt, who led IHMC’s DARPA Robotics Challenge team before serving as CTO of Figure. When IEEE Spectrum first spoke with them two years ago, the company had committed to building an economically viable humanoid but had not chosen its market. As Radford now admits: “We were all over the place. Warehousing, automotive, we probably even mentioned the home.”

Those are the same targets every humanoid developer is chasing, and the Spectrum interview is blunt about the scoreboard: despite an ever-growing catalog of demonstrations, none have succeeded at useful scale. Persona’s response was to reject the sector’s default assumption, that humanoids should be priced against unskilled labor, and to build instead around what Radford calls a skilled trades thesis. Their entry point is welding: handheld tool use, long linear welds, in shipyards. Radford calls the positioning “last-mover advantage”: they watched what everyone else was doing and chose a different way.

The choice of welding is more sophisticated than it looks, and the interview reveals the reasoning in layers.

The task fits robots unusually well. A weld can only travel as fast as metal melts, so the ceiling on task speed is roughly a centimeter per second, comfortably within the control bandwidth of a biped. Quality welding involves millimeter-scale repeated motions sustained over long durations, precisely where machines outperform humans, who get tired and bored. The first targets are not the hardest welds but the most voluminous ones: Radford notes a ship requires hundreds of kilometers of linear welds. Volume, not difficulty, is where the money is.

The customer has a scarcity problem, not a cost problem. Persona’s two public partnerships are HD Hyundai, the world’s largest shipbuilder, and POSCO, one of the world’s largest steel producers, both in Korea. Radford says these customers are running at significant backlog and are labor-constrained: the pitch is to grow their top line, not cut their bottom line. “Even if our robot was more expensive than a human, that would still be valuable to these companies, because it could unlock additional revenue,” he told Spectrum. That single sentence inverts the entire industry pricing debate. A robot that must undercut a warehouse wage has a hard ceiling on its bill of materials. A robot that unlocks an otherwise unbuildable ship has none.

The environment forgives what a home does not. Bipedal robots can fall, and falling near an untrained civilian is the nightmare scenario for consumer deployment. A shipyard is staffed by workers already trained around heavy industrial equipment, which relaxes the most expensive safety constraints. The form factor itself is justified by the terrain: yards are a couple hundred meters long, crisscrossed by horizontal and vertical spars that must be stepped over, with work on the ground, overhead, and through portholes. Persona considered four legs or more, and concluded two legs would be least disruptive to existing shipyard rhythms. Radford also notes the company is interested in customers who can support hundreds of robots per location, which is the demand density that makes service and deployment economics work.

And the niche is defensible precisely because it is unfashionable. With no other humanoid company seriously pursuing shipyard welding, Persona faces no price pressure and can build for quality: as Pratt puts it, a home robot would face brutal competition and cost pressure, while an industrial tool-user can be priced on the value it adds. The expansion path runs through adjacent skills, grinding, painting, and other fabrication, with Radford’s stated ambition to become “the largest repository of industrial skills.”

The caveats here are equally real. Radford admits investors have told them “we’re not thinking big enough,” and the company still needs to collect what Spectrum describes as tens of thousands of hours of expert demonstration data, build high-fidelity simulations, and prove the entire thesis in production. Nothing has been delivered at scale yet.

Two attacks on the same line item

Set the two stories side by side and the symmetry appears.

Persona’s problem, stated in their own roadmap, is data: tens of thousands of hours of expert demonstrations, at high cost, per skill family. That is the expensive half of the integration line item. GEN-1.5’s result, taken at face value, is that the data requirement per new skill is falling by orders of magnitude per model generation: from tens of thousands of gradient steps to 10, from full demonstration datasets to one recording, from robot-collected data to human hand gestures or synthetic rollouts from a simulator. Neither company references the other, but they are working the same equation from opposite ends. Persona maximizes the numerator: value per skill, by choosing tasks where a single competence is worth six figures a year. Generalist shrinks the denominator: cost per skill, until teaching a robot resembles teaching a person.

The uncomfortable implication is for the middle of the market: developers chasing unskilled warehouse and logistics labor with task-by-task integration models. They face the highest integration costs relative to value created. If one-shot and few-shot adaptation continue improving at the current rate, the moat of owning a proprietary data engine grows while the moat of owning a specific task pipeline shrinks. The volume players, Tesla, Figure, Agility, and the Chinese manufacturers now flush with public market capital, are betting they can power through with scale. Persona is betting the margins live in trades nobody else wants. Both can win. The companies in between, with general-purpose ambitions and bespoke integration cost structures, are the ones squeezed.

There is also a sober read of the numbers that both companies, to their credit, do not hide. Fifty-nine percent success is not a workforce. Eighty-three percent on short-horizon table tasks is not a shipyard. Persona has partnerships and a thesis, not yet a fleet. The synthesis of the two stories is a direction, not an arrival: the cost curve for robot skills is bending down at exactly the moment someone has finally found a place where even today’s cost curve clears the bar.

What to watch

First, whether Generalist, or any lab, demonstrates one-shot or few-shot learning on long-horizon industrial tasks rather than table-scale manipulation; that is the bridge between the two halves of this story. Second, whether Persona discloses robot counts, weld metrics, or throughput from the HD Hyundai and POSCO partnerships, the first hard numbers on skilled-trades humanoid economics. Third, whether the physical prompt engineering pattern, libraries of composable demonstrations, gets adopted by any of the major deployment players, which would signal that in-context skill acquisition has crossed from research into operations. And fourth, the pricing behavior of the volume manufacturers: if foundation-model adaptation keeps getting cheaper, hardware cost becomes the only remaining battlefield, and yesterday’s IPO valuations are a wager on exactly that endgame.

The humanoid sector spent 2024 and 2025 proving it could build the machines. In 2026 the interesting question has narrowed to unit economics: what does a robot-hour cost to produce, and what is a robot-hour worth? GEN-1.5 attacks the production side of that equation. Persona attacks the worth side. The demo reel era ends when the two curves meet. This month moved both.


Sources: Generalist AI: GEN-1.5, Embodied Foundation Models are One-Shot Learners; IEEE Spectrum: How Persona AI Makes Humanoids Pay Off In Welding; IEEE Spectrum Video Friday, August 21, 2026; Language Models are Few-Shot Learners (GPT-3 paper); Persona AI.