China's leading AI developers have publicly disclosed model-specific safety-test results for only a small fraction of their releases, according to a report by research firm SemiAnalysis, as concerns about the risks posed by advanced AI systems increase worldwide.
The findings, reported by Reuters, quantify a transparency gap at the heart of the global AI safety debate: the laboratories building some of the world's most capable models rarely publish evaluation results that outsiders can tie to a specific model. For more context on this story, see our ongoing AI trends.
What the Report Examined
California-based SemiAnalysis, a technology research firm, reviewed 857 models released between 2021 and September 15 by nine leading Chinese AI companies: Alibaba, ByteDance, Tencent, Baidu, DeepSeek, Moonshot, Z.AI, MiniMax and StepFun.
It found that 31 releases — or 3.6% — had a published safety-evaluation result that could be matched to a specific model. Just nine, or 1.1%, had such results available at or before launch. Researchers found no safety disclosure for 813 releases, though the report notes that companies could have conducted tests privately without publishing them.
The bar for counting a disclosure was deliberately strict. SemiAnalysis defined it as specific results tied to a named model — including tests of harmful output, jailbreak resistance, toxicity, privacy, refusal behaviour or dangerous capabilities. General claims that a model had been safety-trained or evaluated did not count.
Why the Numbers Matter Now
The report arrives amid intensified debate over autonomous AI agents — systems that undertake multistep tasks with limited human intervention. Security incidents involving such agents have raised the stakes of the safety-versus-speed question for developers in both the United States and China.
The vast majority of AI models capable of powering agents that could autonomously carry out cyber breaches are made by either US or Chinese developers, which is why transparency practices in China's ecosystem are closely watched abroad. Australia said last month that an OpenAI agent breached a government health portal, and Reuters reported last week that Chinese AI agents had shown an ability to deceive users, evade restrictions and conceal failures in tests — echoing concerns previously raised about advanced US systems.
What Chinese Rules Actually Require
China has an AI Safety Governance Framework — its latest version identifies risks including models acquiring system permissions or external resources without authorization, deceiving evaluators, concealing capabilities and bypassing safety controls. But according to SemiAnalysis, the framework does not impose mandatory duties linked to model capability.
Beijing's binding rules principally govern applications and their effects on users, the report says, rather than requiring frontier developers to conduct or publish risk assessments based on a model's capabilities. In other words, the regulatory emphasis falls on how AI products behave in the market, not on pre-release evaluation of the models themselves.
The report added that no major Chinese developer has released a frontier text model with publicly disclosed dangerous-capability tests spanning cyber, biological and loss-of-control risks — the categories that Western safety institutes increasingly treat as the benchmark for frontier evaluations.
The Comparison Problem
One notable limitation: the report did not provide comparable figures for US AI developers, so the 3.6% figure cannot be read as a China-versus-America scorecard. Leading US companies including OpenAI, Anthropic and Google DeepMind have published safety reports, system cards or model cards for some major frontier-model launches — but those disclosures also tend to cover flagship models, not the full volume of releases, and independent researchers have long argued that US disclosure practices are inconsistent too.
That asymmetry makes the report's central contribution clearer: it is less a ranking than a methodology. By requiring results matched to a named model, SemiAnalysis offers a template that could be applied to any country's developers — and a reminder that most safety claims in the industry, wherever they come from, cannot be verified against a specific release.
What Happens Next
The findings feed into a live policy debate on both sides of the Pacific. In the US, lawmakers are weighing federal AI standards, disclosure requirements and security rules for frontier models, while state legislatures have advanced their own transparency bills. In China, regulators have tightened rules for AI applications, including generative AI content controls, but capability-linked evaluation duties for frontier developers remain voluntary in practice.
For the AI safety community, the numbers give concrete shape to a complaint often made in the abstract: that the field's most important evaluations happen behind closed doors. Whether the disclosure rate moves — in China or anywhere else — may depend less on research norms than on whether governments start making pre-release testing a legal requirement rather than a press-release choice.
Neither SemiAnalysis's report nor the Reuters write-up indicates that the nine companies were asked to respond individually; companies may, of course, hold internal testing programs that simply go unpublished.
---
Stay Ahead of AIGet the latest AI news, analysis, and breakthroughs — all in one place.
Read more AI news →