Don't Be Fooled by Demos: Embodied AI Profit and True Generalization Differ
Robot demos are now a fixture of tech forums: cooking oden, carrying water cups, tidying desks — all in smooth, fluid motion. That easily breeds a mainstream conclusion: keep stacking model capabilities, and the success of large language models will simply repeat, with demos paying for themselves as commercialization follows naturally.
Yet the reality is counterintuitive: at the same 50% task success rate, a large language model remains usable, while a physical robot can barely be delivered. This is not a matter of tuning details — it is a chasm between the underlying logic of the digital and physical worlds.
At the 2026 Inclusion Bund Conference, Xu Huazhe — founder of Poke Robot and assistant professor at Tsinghua's Institute for Interdisciplinary Information Sciences — split embodied AI into two entirely different commercialization paths, and picked a side: the scene post-training path can generate near-term cash flow, but it is not the real endpoint of this wave.
The first is the post-training path, now the route most projects use to earn revenue. Teams bind deeply with factories and brands, locking onto a single fixed scenario for closed-loop adaptation; Xu's example is an oden-cooking robot that performs the full set of stall cooking operations. The metrics are how many labor hours a deployment takes and how many sites it can be replicated to, ultimately landing on ROI.
Its advantage is concrete: with scenario boundaries fixed and the environment and process fully controlled, targeted fine-tuning can push localized success rates near 100%, producing direct cash flow. The cost is equally clear: every new scenario demands heavy debugging hours, and generalization stays very limited. These are bespoke projects, not a general-purpose robot brain.
The second path is the new yardstick for this wave: moving toward large-model scale and chasing general capabilities close to those of LLMs. This is 'frugal commercialization,' which generates almost no positive economic return in the short term.
One contrast is striking: even at 50% success, an LLM lets users reach a usable result through multi-turn correction and revision; but for a robot in the physical world, a 50% general-task success rate means every other attempt smashes a cup or knocks over materials — there is no undo or regeneration, so users simply judge the device unusable. That is why embodied AI's real-world experience trails LLMs so far.
A key misconception must be cleared up: Xu's 50% is not a single closed scenario but the success rate on open tasks facing everything in the physical world. Tuning a closed scenario to 99% is not hard; but in the open world, climbing from 50% to 60%, 70%, and finally the 99.9% threshold is what brings true general commercialization.
Hence the first counter-mainstream judgment: do not equate a commercially proven closed-scenario case with a commercial victory for general embodied AI. Many signable robot projects today are essentially upgraded traditional automation, buying stability with heavy manual post-training — not proof that a general robot brain has matured.
The second judgment: the industry widely hopes to 'build a general model while earning money from scenarios,' but the two paths have inherently conflicting resource demands. Post-training spends on scene engineering and on-site debugging; the model route must pour into massive physical data, simulation, and base-model iteration. Budget and engineer energy are finite — chasing short-term orders crowds out base-model work, while obsessing over a general foundation brings the strain of long periods without revenue.
None of this means rejecting scenario deployment outright. Xu is not arguing for abandoning cash flow: post-training projects are how companies survive, funding teams and accumulating some real physical data. But the means must be separated from the end — surviving on bespoke projects is not the same as reaching the destination of this embodied AI wave.
The real hope is that once base models cross the capability threshold, robots can take over the vast, non-predefined tasks of the physical world, rather than forever relying on manual tuning for every scenario.
The track is now in an awkward phase: capital wants to see general capabilities in demos while urgently demanding deployable orders. Many companies fall into a compromise trap, packaging heavily customized orders as general embodied AI results and treating single-scenario success as proof of general ability.
But the physical world does not compromise: an oden robot fixed on one stall will likely fail on another with a different set of pots. No launch-event demo video can smooth over that gap.
There is no standard answer. The future may not be a black-and-white trade-off: some companies will keep deepening scene post-training as cost-effective automation providers, while a few keep investing and slowly climb toward 99.9% general-task success. The real variable is whether capital's patience lasts until base models cross the threshold — and whether, when that day comes, today's scenario-order teams can still complete the leap.