Chinese AI Labs Disclosed Safety Tests for 3.6% of Releases
Chinese AI labs disclosed safety tests for just 3.6% of 857 model releases reviewed by SemiAnalysis, exposing a large gap between the speed of model publication and the public evidence available to assess risk. Only 31 releases had a safety result matched to the specific model.
The study covers releases from nine leading developers between 2021 and September 15, 2026. It measures public disclosure, not whether companies tested models privately, an important distinction that makes the findings a transparency audit rather than proof that safety work never occurred.
The review produced four headline findings:
- 31 of 857 releases had a matched safety result.
- Only nine published results by launch day.
- 813 releases had no public safety disclosure.
- Post-launch results appeared after a median 42 days.
Chinese AI Labs Disclosed Safety Tests for 31 Releases
SemiAnalysis assembled an original dataset of 741 product models and 116 research models from Alibaba, ByteDance, Tencent, Baidu, DeepSeek, Moonshot, Z.ai, MiniMax and StepFun. Each entry included its first public date, weight availability, license and source.
The researchers then searched model cards, release notes and technical reports for safety findings tied to a named release. General statements that a model was safety-trained or evaluated did not qualify without a quantitative or substantive result that outsiders could examine.
That rule produced 31 matched disclosures, equal to 3.6% of the full dataset. Ten more releases carried a claim of evaluation without figures, while three were known only through press or investor accounts for which the researchers could not retrieve developer documentation.
The remaining 813 releases, or 94.9%, had no disclosure in the materials checked. SemiAnalysis stresses that an unavailable public result does not prove testing never occurred, because a company may run internal evaluations without publishing them or may store documentation outside the sources available to the review.
SemiAnalysis Used a Strict Model-by-Model Standard
The audit counted findings on harmful output, jailbreak resistance, toxicity, privacy, refusal behavior and dangerous capabilities. A result for one flagship model could not be applied to smaller sizes, later snapshots or related products simply because they shared a family name.
This model-level approach improves precision but also affects the denominator. Alibaba’s 238 releases include individual Qwen sizes and snapshots, so a single safety report covering a family may not count for every variant unless the document identifies them. The authors therefore describe company rates as indicative rather than a definitive ranking.
Publication practices varied substantially. Alibaba had seven releases with any matched result and three available at launch. Tencent had one among 133 releases, ByteDance two among 120 and Baidu one among 49, according to the dataset.
Startups collectively disclosed results for 20 of 317 releases, a 6.3% rate, compared with 11 of 540 releases, or 2%, for the four hyperscalers. Even the stronger group did not make model-specific safety reporting routine.
Related Research
At-Launch Safety Disclosure Fell to 1.1%
Timing produced an even narrower result. Just nine releases, or 1.1% of the dataset, had a matched safety finding available at or before launch. Sixteen were documented later, with a median delay of 42 days and a maximum delay of 349 days.
For six of the 31 matched releases, the researchers could not establish when the result appeared or confirm an exact timing relationship. That uncertainty prevents the published findings from being treated as a complete launch-day record.
The type of evaluation also matters. The report found 18 documents containing harmful-output or refusal results and seven addressing jailbreak resistance. Nine concerned code or cybersecurity, but seven of those measured secure code generation rather than a model’s offensive cyber capability.
Only three disclosures touched cyber-offence or biological risk, and none came from the four hyperscalers. SemiAnalysis also found that 93% of reasoning-model releases lacked a published safety result, even though reasoning systems represent one of the fastest-advancing product categories.
China’s Rules Focus on Applications More Than Frontier Capability
The report places the disclosure record beside China’s AI governance system. Chinese rules address generated content, service providers, labeling and application-level harm, while the latest safety framework identifies risks such as unauthorized resource access, evaluator deception and bypassed safeguards.
SemiAnalysis argues that those frameworks do not create binding duties triggered by a model reaching a defined capability threshold. That differs from governance proposals that require dangerous-capability testing before a frontier system is released or expanded.
Public results are not a guarantee of safe behavior, and disclosure quality can vary with test design. They do, however, let researchers, buyers and regulators compare claims, identify missing coverage and examine whether safeguards keep pace as models gain stronger coding, tool-use and autonomous capabilities.
The study’s clearest conclusion is therefore about evidence, not hidden company practice. Chinese developers are releasing models rapidly, but outsiders rarely receive model-specific safety findings at launch. Closing that gap would require routine documentation, consistent capability tests and publication schedules tied to releases rather than delayed or family-level assurances.