Can GPT-6 Really Become a Robot Brain? Experts Debate World Models
World models are becoming one of the hottest terms in embodied AI, discussed by those working on video prediction, simulation, and robot action generation. Meanwhile, GPT-6 has been connected directly to robots for real-machine tests, prompting a debate: since language models already have strong reasoning, memory, and planning abilities, can they directly take over a robot's brain?
Huxiu AI100 hosted a closed-door meeting with Chen Wenke, world model partner at Octopus Dynamics; Qiu Zhun, overseas partner at Huaying Capital; Wu Hao, co-founder of Domain Transformation; and Wu Tong, world model lead at Moqi Intelligent. They discussed from four perspectives: model R&D, capital, self-evolution, and academic frontiers.
Qiu Zhun proposed a way to judge a project's value: hide the words 'world model' and look at what it actually does. Can it understand the physical world and predict the result of an action? Can it make decisions and run them on a real robot? After a first failure, can it do better next time? Wu Hao distilled this into three questions: can it predict the future, can prediction help decision-making, and can decisions close the loop on a real robot? If it can only generate a beautiful video but cannot help a robot decide its next step, it is still far from a true brain.
Chen Wenke breaks down a robot brain into L, V, and T: language, vision, and touch. GPT-6's significance is that it again shows scaling a model to a certain point may naturally produce new abilities, but that does not mean Language Scaling can directly become Action capability. A person reaching for a cup does not first describe it in language; many physical actions do not use language as the most direct medium. L is better for understanding, planning, and long-horizon task decomposition, while V and T directly connect the physical world and robot actions.
Wu Tong's team offers another view from GPT-6 real-robot tests. GPT-6 showed stronger instruction generalization, in-context learning, and long-horizon understanding and interaction, but its limits are clear. First, real-time performance: a cloud model can reason slowly, while a robot must react quickly. Second, visual feedback is irreplaceable; when the camera is partly blocked, task performance drops noticeably. She suggests a realistic direction may be for a large language model to handle a higher-level, lower-frequency System 2, while a VLA or World Action Model handles a faster System 1 closer to action execution.
From demo to deployment, the brain must self-evolve. Wu Tong's technical standard is direct: for a World Action Model, see whether the generated policy can run stably, safely, and with high success in real environments; for a world simulator, see whether it actually improves subsequent policy training. Wu Hao argues the real world is long-tailed and open, so a brain that truly moves toward real-world deployment needs continuous evolution through physical interaction. He uses an analogy: if a model scores only 40 today, it does not have to be locked in a lab until it reaches 90 before deployment. A better state is to let it enter an environment and learn from success and failure through real interaction. Chen Wenke adds that Scaling can only hold if data and infrastructure are done right first.
The closer to the real world, the less there can be a suspended brain. Chen Wenke's answer is direct: doing only a brain absolutely will not work. The same robot model can differ across hardware batches; actuators have errors; vision has noise; contact moments require high-frequency feedback. These problems will not disappear just because the upper-level model becomes smarter. Wu Tong's judgment is that a general brain can be relatively independent, but a truly usable robot still needs software-hardware coupling. Her team, after getting a GPT interface, first tried it on a commercial robotic arm, then deployed the model onto their own body. The model had never seen that body's own training data, yet could still complete navigation, handshaking, and lifting a doll. This shows some abilities can cross embodiments. But if the goal is not barely working but natural, smooth operation for commercial scenarios, then the brain and hardware definitely need further adaptation.
Embodied AI's AI Coding has not appeared. Qiu Zhun uses the value logic of foundation model companies like OpenAI to understand the long-term possibility of a robot brain, but the problem is no one yet knows who will achieve it, or even which scenario will come first. He stressed an investment principle: high cognition, low judgment. One of embodied AI's most important business questions also has no answer: what exactly is embodied AI's AI Coding scenario? Large language models have gradually found Coding—a high-value, high-frequency scenario that can continuously absorb model capabilities. What about embodied AI? Factories, logistics, homes, elderly care? Qiu Zhun's judgment is that it is too early to answer. Wu Hao emphasizes closed-loop evolution in real environments; Chen Wenke emphasizes model capability itself. One is closer to Deployment first, the other to Scaling first. The real watershed may be neither who makes the biggest model first nor who sends robots into the most scenarios first, but who first truly connects these two lines: letting the brain keep getting smarter while the real world keeps giving it feedback.