Embodied AI FRONTIER
- Your Embodied AI Intelligence Hub -
/ Topic: WRC 2026 World Robot Conference / Where Are World Models on the Road to Generalization?
Industry 📡 钛媒体

Where Are World Models on the Road to Generalization?

[TechPulse reports] The form of a general world model remains unsettled in the AI community: video generation companies see it as a simulator, robotics firms as a decision brain, and autonomous driving players are betting on it. At a session of the World Robot Conference on August 19, Shengshu Tech founder and chief scientist Zhu Jun offered his answer: a five-tier roadmap, with his company's research and deployment reaching the third tier.

Starting from first principles, Zhu drew an analogy to learning to ride a bike or drive a car, where the brain gradually builds an internal model through interaction with the external world. He argued that a general world model must possess three interconnected capabilities: understanding the world (perceiving environment, assessing states), predicting the future (anticipating what happens next and outcomes of different actions), and taking action (affecting the digital or physical environment and correcting understanding/prediction with real feedback). He stressed this is not a simple concatenation of generators, simulators, or policy models, but a closed-loop feedback system.

On data, Zhu proposed a multi-layer pyramid: vast internet videos at the base, then domain videos, first-person human videos, human demonstrations with action recordings, and real robot interaction data at the apex—each layer more scarce and costly but more directly linked to task actions and physical outcomes. Architecturally, Shengshu uses Mixture-of-Transformers (MoT), where each modality has dedicated parameters and shared attention enables cross-modal interaction, allowing environment understanding, state prediction, and action generation in one model. The world-action model Motubrain is built on this, integrating understanding, prediction, and action.

Zhu outlined five levels: L1 World Generation—model learns to generate coherent world evolution, typically via videos; L2 Interactive World—model responds to real-time inputs (language, speech, control signals) and iterates; L3 Actionable World—model enters the physical world, unifying environment understanding, future prediction, and action generation to output executable actions and refine with real feedback; L4 Autonomous World Agent—actively perceives environment, decomposes tasks, explores unseen states, and plans over long-term goals; L5 World Orchestrator—coordinates robots, digital agents, humans, and tools for multi-agent task allocation and collaboration. Shengshu's current practice covers the first three levels.

According to Shengshu's data, Motubrain's inference speed is about 10x faster than Motus, and it can adapt to a new robot body with just 50-100 human demonstration samples. It scored 96.1 on the RoboTwin 2.0 benchmark, ranking first, and has been validated on nearly ten platforms including Galaxy General, Xingchen Intelligence, and Qianxun Intelligence. However, high-level world intelligence remains distant: video model generation quality, controllability, and spatiotemporal consistency need improvement, and stability, generalization, and execution efficiency in more complex open environments require optimization.

For L4 and L5, Zhu highlighted six key challenges: establishing a joint evaluation system covering understanding, prediction, action, and transfer; learning real physical laws such as contact, friction, force, and intervention consequences; forming persistent and revisable memory; achieving online learning and self-evolution through real interaction; meeting real-time latency and resource constraints for efficient closed-loop deployment; and ensuring safety and controllability. He said that when understanding, prediction, and action form a true closed loop, the world model will become "an important foundation connecting the digital and physical worlds and moving toward general embodied intelligence."

✓ Verified Read Original → 2026-08-25
Recommended
Embodied AI FRONTIER - Your Embodied AI Intelligence Hub -