Yuankee Vision's Chen Pu: Robot Training Data Must Upgrade from 2D to 4D with Sub-Millimeter Precision
On September 23, a seminar on the high-quality development of embodied AI, themed "From Demonstration-Level Intelligence to Scaled Application," was held in Beijing. It was hosted by the Digital Economy Research Center of China Economic Information Service.
Chen Pu, Vice President of Product R&D at Beijing Yuankee Vision Technology, said that as world models gain traction, high-quality real data for robot training is drawing attention. Upgrading data from 2D to 4D and precision from centimeter to sub-millimeter is key to building a 4D data foundation for robots.
The embodied AI industry currently faces weak model generalization, poor real-world deployment, and inadequate scenario adaptability. Single training scenarios, insufficient dynamic data, and inconsistent underlying foundations are core bottlenecks to scaled deployment.
Chen Pu noted that manually built ideal scenarios differ greatly from real industrial scenarios. 2D video cannot fully represent 3D space and time changes, and data collection dimensions and modalities are often single, causing robots to only perform repetitive imitation and struggle to understand the world, predict the future, and plan.
To address this, Yuankee Vision has developed a new data collection and training paradigm based on spatial intelligence and 3D reconstruction. It involves three systematic upgrades. First, dimension upgrade: 2D data is upgraded to 4D spatiotemporally consistent data, capturing full data of continuous changes during robot-environment-object interactions. Second, modality enrichment: force and tactile interaction data are added to supplement physical contact information.
Third, precision improvement: high-precision motion capture technology raises collection precision from centimeter to sub-millimeter.