Embodied AI Lags LLMs by at Least Five Years, Experts Say(Yicai) Aug. 24 -- Embodied intelligence, which enables artificial intelligence to interact with the physical world, remains at least five years behind large language models, with scarce data, limited model scale, and fragmented technical systems holding back its development, according to experts who spoke at the 2026 World Robot Conference in Beijing.
The gap is evident in model scale and computing requirements, while a shortage of physical-world training data poses an even greater challenge, according to Zhang Jianzhong, chairman and chief executive of graphics processing unit developer Moore Threads Technology.
The WRC brings together robotics companies, researchers, technology suppliers, and industry organizations. Participants this year also highlighted rising computing needs and the lack of unified robot bodies, software and hardware architectures, interfaces, and data systems.
The industry is pursuing different paths to overcome these hurdles, from scaling models and computing power and expanding synthetic data to improving robots' ability to understand goals, promoting open-source technologies, and developing common standards. Experts estimate large-scale deployment could still be three to eight years away, depending on the application.
Embodied Intelligence Is at Least Five Years Behind LLMs
According to Moore Threads' Zhang, foundation models for embodied intelligence are "at least five years behind" those for language models.
Most vision-language-action models, which enable robots to interpret visual and language inputs and translate them into physical actions, currently have around seven billion parameters, while earlier LLMs had already reached the hundred-billion-parameter level. Many teams can train VLA models using only a few dozen GPUs, or even just a few cards. "It is certainly still some way from its ‘GPT moment,’" Zhang said, referring to a breakthrough comparable to ChatGPT’s impact on generative AI.
An even greater challenge is data. Language models can draw on humanity's accumulated knowledge and information available on the internet, while the physical-world data needed to train AI agents remains scarce. Robots must learn to perform precise operations, adapt to changing lighting, weather, and other environmental disturbances, and handle flexible objects -- information that cannot be obtained as readily as online content.
If embodied intelligence follows the same scaling path as LLMs -- improving performance by using more data, larger models, and more computing power -- its computing needs will continue to grow. Major breakthroughs in the next generation of visual and world models will require clusters comprising thousands or even tens of thousands of GPUs, Zhang said.
Computing power, however, is only one requirement. Real-world data is expensive and difficult to collect exhaustively, making synthetic data and simulation another important avenue.
Some extreme scenarios are particularly difficult to capture, Zhang noted, such as the sudden change in lighting when a vehicle emerges from a tunnel. GPUs can instead generate synthetic data featuring different backgrounds, lighting, and weather conditions, as well as other extreme situations. Trained models must then be transferred from simulation to the real world and validated.
Do Robots Really Need an Infinite Amount of Data?
A team led by Gordon Cheng, professor and chair of cognitive systems at the Technical University of Munich, is taking a different approach to robot training. The team has taught robots dexterous manipulation by having them observe humans playing with exercise balls and has also had robots watch people slice bread and sort oranges.
The core idea is not for robots to simply replicate every action, but to gradually understand the goals behind those actions. Knowledge about the purpose of a task can then be transferred to other robots.
"If you scale up, this reproduction doesn't require a lot of data to learn purposeful tasks," Cheng said, adding that it may eventually even be possible to package a specific task capability into a "compressed file" and distribute it to different robots for execution.
Cheng's approach illustrates an alternative to simply scaling up models, data, and computing power: giving robots cognitive, imitation, and goal-understanding abilities closer to those of humans, reducing the need to train every sequence of actions separately.
From Data Silos to Open Source
Different companies are pursuing different technical routes, from hardware configurations to software architectures, according to Xie Shaofeng, head of a national humanoid robot and embodied intelligence standards committee under China's industry and information technology ministry.
As a result, many companies repeatedly carry out basic adaptation, system integration, and engineering validation, solving the same underlying problems multiple times. Sensor configurations, action spaces, task definitions, and data structures also vary among robots, creating "data silos," Xie said.
If this situation persists, divergent technical approaches will lead to duplicated research and development, while the absence of standards will further raise collaboration costs, ultimately hindering product maturity, cost reductions, and market expansion, Xie said.
Xiong Youjun, general manager of the Beijing Humanoid Robot Innovation Center, a state-backed robotics research and innovation platform, echoed those concerns, citing duplicated development, high verification costs, and disconnected systems. "Companies are not compatible with one another," he said.
Efforts are now underway to break down these technical and data barriers through greater open-source collaboration. In June, under the guidance of China's industry and information technology ministry, the OpenAtom Foundation and 88 other organizations jointly launched the Humanoid Robot Open Source Community. The initiative aims to connect the entire technology chain, including operating systems, algorithms, models, data, communications, and robot bodies.
The Beijing Humanoid Robot Innovation Center is also seeking to further open its technological foundation to the industry, including its robot body, brain, cerebellum, and data technologies. Its open-source technologies have been downloaded more than 16 million times, while it has cultivated more than 1,000 developers, Xiong said. The center has also deployed and validated its technologies in applications including guided tours, shopping assistance, and power inspections.
From Individual Capabilities to Integrated Systems
Progress in these areas may attract less attention than improvements in robots' running speeds or jumping heights, but integrating individual capabilities into a stable, operational system could determine whether robots can move from laboratories into factories.
Federico Pecora, global head of physical AI robotics research at Arm, said the next phase of large-scale robot deployment will revolve around four major system-level challenges: how capabilities are implemented, how different capabilities can continuously collaborate during operation, how those capabilities are mapped onto computing resources, and how the capabilities are managed.
Addressing these challenges will require collaboration and technological breakthroughs across the industry chain, Pecora said.
Large-Scale Deployment May Still Be Years Away
As for how long it will take robots to achieve truly large-scale deployment, Yang Yuxin, chief marketing officer of Chinese chipmaker Black Sesame Technologies, offered a relatively cautious estimate of around five to eight years.
Chips could achieve relatively acceptable performance and cost within three to five years, Yang said, but the software ecosystem will take longer to mature.
Li Zhaoshi, dean of the MetaX Research Institute, drew a distinction between factory and home environments. The researcher at the Chinese chip developer said that industrial and household applications may take five to eight years to reach scale, compared with three to five years in semi-structured factories.
Home environments involve numerous ethical, long-tail, and safety challenges. In semi-structured factories, however, robots performing jobs that require perception, decision-making, and memory could potentially reach an "eight-out-of-10" level of task performance within three years, Li said. They could then continue improving through reinforcement learning using samples collected during real-world operations, he added.
Editor: Emmi Laine
