How Are Humanoid Robots Trained? What 2026 Repriced in Robot-Free Data
Why it matters: the bar in robot training data moved from having enough to being able to prove it.
By Embodied AI Frontier
On September 2, 2026, X Square Robot released TwinDEX, a system pairing a wearable three-finger exoskeleton with a matching robot end-effector. Using a few hundred demonstrations collected by hand, the company's policy ran a chemistry experiment end to end, covering 24 sub-actions in a single continuous take, with no on-robot teleoperation data used in post-training1TwinDEX project pageView the entry below →2TwinDEX Introduces a Scalable Path from Robot-Free Data Collection to Real-World Dexterous ManipulationView the entry below →. The significance goes beyond one demo. It put a ten-year-old default on the table: why the most valuable training data sits at the apex of the pyramid. This article separates 2026's "robot-free data" into three collection sources, checks the cost claims against the validity claims, and argues that what was repriced this year was the scarcity premium on real-robot data, not its categorical status. The bar has moved from "do we have enough" to "is it usable".
Article structure
- 1. A Demo With No Teleoperation Data
- 2. The Demand Side: Task Complexity Is Moving Up
- 3. Why Real-Robot Data Sits at the Apex
- 4. Two Routes in 2026
- 5. One Company, Five Months, Four Framings
- 6. What Has Not Changed
- 7. Outlook: What to Watch
1. A Demo With No Teleoperation Data
On September 2, 2026, X Square Robot, based in Shenzhen, released a system called TwinDEX. It consists of a matched pair: a wearable three-finger exoskeleton for collecting demonstrations, and a robot end-effector built to pair with it. The two are co-designed, structurally isomorphic, and carry the same tactile sensors at the same locations, so finger states map directly into the robot's joint space without a human-to-robot retargeting step1TwinDEX project pageView the entry below →2TwinDEX Introduces a Scalable Path from Robot-Free Data Collection to Real-World Dexterous ManipulationView the entry below →.
The accompanying demo is a chemistry experiment: opening a bottle and taking a sample, drawing liquid with a pipette, guiding flow with a glass rod, and shaking the mixture to observe. The policy executes the whole sequence on its own, across 24 sub-actions, three tools, and repeated bimanual coordination and tool switching3自变量发布 TwinDEXView the entry below →1TwinDEX project pageView the entry below →. According to the company, the policy was trained on a few hundred demonstrations collected by the wearable device, and no on-robot teleoperation data was introduced during post-training. The same materials report collection throughput of roughly 5.3 times that of on-robot teleoperation1TwinDEX project pageView the entry below →2TwinDEX Introduces a Scalable Path from Robot-Free Data Collection to Real-World Dexterous ManipulationView the entry below →.
The caveats come first. This is a single 103-second recording with no timecode, no third-party camera, and no documentation of control signals, so "fully autonomous" can be neither confirmed nor ruled out from that file1TwinDEX project pageView the entry below →. The scene is a single desktop workspace with a limited set of objects, as the company states on its project page, which also notes that the ergonomics of long-duration wearable collection have not been systematically studied1TwinDEX project pageView the entry below →. As of October 3, 2026, the hardware documentation, the data, and the technical report are all unreleased: the BibTeX on the project page still reads "Coming soon", the paper link points back to the same page, the GitHub organization's repository list does not include it, and an arXiv search returns zero hits1TwinDEX project pageView the entry below →8[8] GitHub, arXiv and Hugging Face interface checks (Oct 3, 2026)View the entry below →.
With all of that laid out, the episode is still worth reading, and not because of this one company. The same problem surfaced across several players in 2026. Google DeepMind folded whole-body control and dexterous hands into a single narrative when it released Gemini Robotics 2 in August. 1X disclosed in January that its world model used 900 hours of egocentric human video and 70 hours of robot data19Company engineering blogView the entry below →. In February, NVIDIA, UC Berkeley and the University of Maryland released EgoScale, using more than 20,854 hours of action-labeled egocentric video and reporting a log-linear relationship between human data scale and downstream robot performance14EgoScale: Scaling Dexterous Manipulation with Diverse Egocentric Human DataView the entry below →. In July, a survey by 29 authors arranged the entire embodied data ecosystem into a single "data pyramid"10Data Pyramid for Embodied Manipulation: A SurveyView the entry below →.
A robot learning a chemistry experiment is not news. Not one of the few hundred demonstrations behind it came from a robot. That is the part worth counting.
2. The Demand Side: Task Complexity Is Moving Up
2.1 Work That Has Not Been Replaced
Before discussing data, it helps to look at the work itself. Over the past two years, robot bodies have cleared the mobility bar: running, jumping, flipping and climbing are standard capabilities for a group of vendors in 2026. What actually separates them are the tasks demanding millimeter-level positioning and stable force control: flexible assembly, cable mating, depalletizing and sorting, quality inspection, laboratory work, and care and rehabilitation assistance.
These tasks share one property. They are indifferent to how closely a hand resembles a human hand, and extremely sensitive to whether the job is done right every single time. One failed insertion on an assembly line stops the line; one burst of excess force from a pipette ruins the sample. The task does not care how many fingers the robot has. It cares about cycle time, first-pass yield, and whether the system is still stable after three months of continuous operation.
2.2 Who Does It Today
Sorting by motion precision and object uncertainty produces a clear order of substitution:
| Tier | Typical tasks | Who does it today | Why it has not switched |
|---|---|---|---|
| Low uncertainty, low precision | Transport, palletizing, machine loading | Two-finger grippers and dedicated automation | Already switched once; cost-optimal |
| Low uncertainty, high precision | Dispensing, screw fastening, inspection | Mostly dedicated equipment, some grippers | Cycle time and repeatability are already pushed to the limit |
| High uncertainty, low precision | Unpacking, sorting, clearing | Mostly manual, grippers in pilots | Object shapes vary too much for gripper generalization |
| High uncertainty, high precision | Cable mating, flexible assembly, lab work | Manual | Requires sustained force control after contact and bimanual coordination |
Table 1: The order of substitution in embodied manipulation tasks and who does the work today
The first two tiers belong to dedicated automation. The last two belong to people. Every discussion about dexterous manipulation in 2026 is really asking one question: whether the third and fourth tiers can begin to shift.
2.3 Five Bars a Buyer Applies
When evaluating a robotic hand or a manipulation policy, buyers apply five tests, and the number of degrees of freedom is not among them: reliability (does it run for three months), service life (how often does it need maintenance or replacement), maintainability (how fast can it be repaired), integration cost (how much rework does the line need), and cycle time (can it keep up with the line).
None of these five is about price, and none is about the price of training data. Chinese suppliers have pushed the price of a hand from RMB 50,000 (about $6,900) down to RMB 6,000 to 8,000 (about $830 to $1,100), yet by MIR Databank's estimate, less than 1% of 2025 demand for dexterous hands went into vehicle or component production lines. Lower prices did not awaken demand. Capability did.
So the question moves one layer down: how is that capability trained.
3. Why Real-Robot Data Sits at the Apex
3.1 A Default That Has Held for Ten Years
In March 2025, NVIDIA wrote a sentence in its official post introducing the Isaac GR00T N1 model: the apex of the pyramid is real-robot data collected through teleoperation17Accelerate Generalist Humanoid Robot Development with NVIDIA Isaac GR00T N1View the entry below →. That sentence has been cited repeatedly for the past year and a half and has become the default frame of reference for discussing data value.
The reasoning is not mysterious. Actions recorded through real-robot teleoperation are executed on the target robot itself: joint angles, end-effector poses, gripper states and force feedback all carry unambiguous executable semantics and can go straight into training. Human video, by contrast, records a human hand; handheld grippers record a different end effector; simulation records approximated physics. Only the actions in real-robot data can be executed directly on the machine that produced them.
That default is not without dissent. In 2026, Sergey Levine wrote in a public post that if the field is to build genuine robotic foundation models, real data remains indispensable, and proxy data can only supplement it22[22] Sergey LevineView the entry below →.
3.2 The Five Tiers and Two Axes
In July 2026, a survey titled "Data Pyramid for Embodied Manipulation" laid the map out in full10Data Pyramid for Embodied Manipulation: A SurveyView the entry below →. It organizes the ecosystem along two axes: scalability, meaning how efficiently data can be expanded relative to hardware dependence, human effort, environment resets and marginal cost; and alignment with the robot, meaning how directly observations, representations and supervision signals support learning and execution on a physical robot. The two axes are inherently opposed.
Along them, the paper sorts data into five tiers, from apex to base: real-robot data, handheld-gripper data in the UMI style, egocentric and third-person human video, simulation data, and general data10Data Pyramid for Embodied Manipulation: A SurveyView the entry below →. On the apex, the paper states that real-robot data sits at the top because its recorded actions are directly executable on the platform that produced them, fidelity bought at a cost in hardware, human effort and resets10Data Pyramid for Embodied Manipulation: A SurveyView the entry below →.

Fig 1 The embodied data pyramid: five tiers and two opposing axes
Source: Data Pyramid for Embodied Manipulation: A Survey (arXiv 2607.24744, Jul 27, 2026); NVIDIA Developer, Isaac GR00T N1 (Mar 2025)
The same paper is explicit about the UMI tier: it is best regarded as a scalable complement to embodiment-specific real-robot data rather than a complete replacement. It also notes that the optimal recipe remains unresolved, because the contribution of each data source has not been systematically isolated, and robot-only models can still achieve strong performance10Data Pyramid for Embodied Manipulation: A SurveyView the entry below →.
3.3 Why the Industry Accepted the Ordering
There is also an engineering reason the ordering stuck: the alternatives were previously not good enough. Handheld grippers depend on wrist-mounted visual SLAM for pose estimation, and trajectories drift when the view is blocked by a hand or an object. Relative poses between two hands are reconstructed from shared camera views, which introduces large errors in coordinated tasks. Cameras, gripper encoders and position sensors come from different devices and are aligned in software or wirelessly, so millisecond-level misalignment at the moment of contact directly degrades policy learning11HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data AloneView the entry below →. Egocentric human video offers scale and realism, but only images, with no executable actions and no tactile signal.
The industry's default practice took shape: robot-free data for large-scale pre-training, then a short stretch of real-robot data for alignment when moving into specific tasks and post-training. On September 24, 2026, Xinhua Finance still recorded that judgment in an industry survey: robot-free data and real-robot teleoperation data are more likely to divide the work, with the broader and cheaper robot-free data covering pre-training, and real-robot teleoperation data still needed for calibration when entering specific tasks, specific factories and post-training23从"遥操作"到"无本体" 具身智能数据迎有效性大考View the entry below →.
4. Two Routes in 2026
4.1 Route One: Raise Collection Fidelity
China's Simple AI released HiFi-UMI in July 2026, rebuilding both the capture hardware and the data pipeline around four dimensions: trajectory accuracy, bimanual pose, time synchronization, and field of view11HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data AloneView the entry below →.
Specifically, it moves the visual reference from the wrist to the head, reconstructing head trajectories offline with a head-mounted stereo camera and an inertial measurement unit. The relative poses of the two grippers are measured natively by hardware rather than reconstructed after the fact, and each hand carries two non-parallel wide-angle fisheye cameras covering roughly 200 degrees. For timing, a single GPIO hardware trigger replaces software alignment: six cameras, the IMU and the gripper encoders are driven by one signal, pushing cross-sensor synchronization error below 40 microseconds11HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data AloneView the entry below →.
The reported results: end-effector trajectory error of about 3 millimeters inside the workspace, with policies deployable to real robots after post-training on its demonstrations alone. Across representative models in both VLA and WAM architectures, success rates came close to in-domain real-robot teleoperation data. On the WAM model LingBot-VA specifically, the two routes reached 56.9% and 57.5%11HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data AloneView the entry below →. The team also open-sourced a 2,000-hour dataset12HiFi-UMI-2K datasetView the entry below →.
The paper also states its own boundary: it shows that high fidelity is sufficient, but does not answer how much of each property is needed, and does not isolate those factors through controlled degradation experiments11HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data AloneView the entry below →.
4.2 Route Two: Make Collection and Execution Structurally Identical
TwinDEX took the other route: bake the action structure the robot actually needs into the collection device up front. The exoskeleton and the end-effector are co-designed, structurally isomorphic, and integrate the same tactile sensors at the same locations, guaranteeing strict consistency at the contact-dynamics and tactile-sensing levels, so finger states map directly into the robot's joint space without retargeting1TwinDEX project pageView the entry below →2TwinDEX Introduces a Scalable Path from Robot-Free Data Collection to Real-World Dexterous ManipulationView the entry below →.
According to the company, the design was fixed after comparing configurations on a benchmark covering multiple manipulation primitives: a three-finger, nine-degree-of-freedom hand (seven active plus two passive) balances dexterity, mechanical complexity, wearer comfort and hardware cost, and was judged the minimum viable solution for the current task set1TwinDEX project pageView the entry below →2TwinDEX Introduces a Scalable Path from Robot-Free Data Collection to Real-World Dexterous ManipulationView the entry below →. The company also states the cost of that choice: tasks demanding higher dexterity, such as multi-finger precision assembly, may require exploring morphologies with additional degrees of freedom1TwinDEX project pageView the entry below →.
There is a telling detail here. X Square Robot's own product line already includes a five-fingered dexterous hand with 20 degrees of freedom, 15 of them active1TwinDEX project pageView the entry below →. A company that also sells a five-fingered hand chose a three-fingered hand to collect data, and wrote the reason down as "good enough". When data scale is the constraint, execution-side dexterity gets deliberately stepped down.