China's National Data Administration said it will develop standards for the data used to train embodied AI systems and guide local authorities on the work, a step Bloomberg reported on September 13 that targets one of the biggest bottlenecks in the fast-growing robotics sector: access to high-quality, diverse and large-scale training data.
The announcement followed a meeting on September 10 chaired by Liu Liehong, who heads the agency. Research institutes, technology companies and humanoid robotics organizations attended, and the regulator said the industry is becoming increasingly data-driven. The administration also signaled it will support companies putting more money into data resources. For more context on this story, see our ongoing AI industry coverage.
Industry Asked, Beijing Answered
Notable is how quickly the regulator moved. Ten days earlier, the same agency had met with seven companies from the artificial intelligence, computing and data sectors, according to MLex. At that meeting, participants called for public data infrastructure for embodied AI, common data standards and token-measurement rules, along with cross-regional coordination of computing resources.
The formal response came less than two weeks later, underscoring how central data policy has become to China's robotics ambitions. Embodied AI — systems that combine machine intelligence with physical machines such as humanoid robots — is viewed in Beijing as a strategic industry, and data is its raw material.
The Data Bottleneck in Humanoid Robotics
The scale of the challenge is considerable. By the estimate of the China Academy of Information and Communications Technology, embodied AI foundation models need roughly 10 million hours of real-world training data. Somewhere between 100,000 and 1 million hours of high-quality data exist worldwide, according to figures cited by The Next Web.
The constraint in humanoid robotics, in other words, is no longer the model. It is hours of recorded reality — data capturing how machines perceive, reason and act in messy physical environments.
A Standards Pipeline Already in Motion
The new push builds on substantial groundwork. National standards published on August 27 cover the quality of real-world embodied-intelligence data and technical requirements for data-generation platforms, with several additional specifications still under development, according to an AP News report.
Among the projects overseen by the National Data Administration are proposed standards covering the sources and constituent elements of high-quality embodied-intelligence datasets, simulated synthetic-data generation and processing, and specifications for data collection and model training at training bases. The work is being developed under the National Data Standardisation Technical Committee.
One draft programme, registered in April, sets a 12-month timetable for a standard on data collection and model training at embodied-intelligence training bases. Its drafting group includes the China Electronics Standardization Institute, Beijing Institute of Technology, the Institute of Software at the Chinese Academy of Sciences and robotics companies. A separate draft addresses synthetic data, which is becoming increasingly important because collecting enough physical-world robot interactions can be expensive, slow and difficult to reproduce.
An implementation plan issued in June called for faster construction of datasets in strategic and emerging fields including embodied intelligence, intelligent driving and the low-altitude economy. It also called for data covering physical interaction, environmental perception and motion control in key scenarios, and directed authorities to improve data cleaning, enhancement, labelling, alignment and quality inspection.
China's Training-Ground Buildout
Beijing is also building the physical infrastructure to collect that data. More than 70 embodied AI training grounds are already operating across more than half of China's provincial regions, with 46 more planned or under construction, TechNode reported. About 86 percent of them are aimed at industrial manufacturing.
The facilities cluster in three regions: the Yangtze River Delta, the Beijing-Tianjin-Hebei region and the Pearl River Delta. Together they form a data-collection apparatus that few countries can currently match.
The Contrast With Europe
The announcement highlights a widening asymmetry with Western regulatory approaches. Europe has written extensive rules for this kind of data without building much of the supply. The EU's Data Act has applied since September 12, 2025 and governs who may access data generated by connected products, including industrial machinery, while design obligations bite on connected products placed on the EU market after September 12, 2026.
The European Commission published a Data Union Strategy on November 19, 2025 whose first priority is scaling up access to data for AI. Ten months on, most of those actions have not been carried out, including the data labs the strategy promised, The Next Web reported. The nearest European equivalent to a Chinese training ground is a private company: NEURA Robotics is building ten robotics training gyms, five of them meant to be running by the end of this year, split between Europe, the United States and China.
Demand-Side Reality Check
Standardizing data supply does not guarantee demand. Only 23 percent of Chinese enterprises surveyed this year said they were satisfied with the robots on offer, a figure that suggests the market — not the data pipeline — may ultimately determine the pace of adoption.
Still, the direction of travel is clear. China is treating embodied AI data as a strategic resource to be systematically produced, standardized and scaled, while its regulators coordinate closely with the industry they oversee. For robot makers worldwide, the gap between jurisdictions that generate training reality and those that mostly regulate its use is becoming one of the defining competitive questions of the field.
---
Stay Ahead of AIGet the latest AI news, analysis, and breakthroughs — all in one place.
Read more AI news →