Summary: An analysis reported on October 9, 2026 found published model-specific safety evaluations for just 31 of 857 Chinese AI model releases reviewed between 2021 and September 2026, about 3.6%. The finding measures public disclosure of specified safety-test documents. It does not establish that all the other models were never tested privately, or that any one model is unsafe solely because a report was not published.

An analysis reported on October 9, 2026 found published model-specific safety evaluations for just 31 of 857 Chinese AI model releases reviewed between 2021 and September 2026, about 3.6%. The finding measures public disclosure of specified safety-test documents. It does not establish that all the other models were never tested privately, or that any one model is unsafe solely because a report was not published.

What researchers actually measured

The review examined hundreds of releases by major Chinese model developers and looked for public evaluations linked to identifiable model versions. Reuters reported 31 qualifying releases out of 857, with only nine evaluations available at or before the associated launch. General promises about responsible AI did not automatically count as model-specific evidence.

A missing public report is not proof that a developer ran no private tests. Conversely, a report does not establish that every possible failure has been found. The headline number concerns transparency under a defined methodology; it is not a statistical estimate that 96.4% of systems are dangerous.

Why test disclosure matters to US customers

Organizations need evidence before connecting a model to internal data, customer service, financial systems or software development workflows. Evaluation material can reveal what was tested, when the model version was assessed, which safety behaviors failed and which mitigations are recommended. Without it, a purchaser may have to rely on vendor marketing claims.

Timing is especially important. If a company publishes a safety summary months after a model launch, early adopters may already have exposed the system to data or tasks outside the evaluated scope. Buyers should request model identifiers, evaluation dates, known limitations and access to meaningful technical findings.

What the data cannot tell you

The figure does not establish that US-made models are automatically safe, nor does it provide a fair US-China comparison without applying the same disclosure definitions to both. Companies use different test frameworks and release practices, and a visible report may itself have significant gaps.

Real deployment risk depends on tools and permissions, data handling, application design, human oversight and the population affected. A standalone chatbot providing brainstorming assistance poses different risks from an autonomous agent authorized to send money, change account settings or execute code.

China also announced new AI policy guidance

In a separate October 9 development, Chinese authorities published guidance supporting advanced industry and artificial intelligence while calling for stronger risk controls and measures against speculative technology investment bubbles. That policy direction is distinct from the finding about historical model-level public disclosures.

Neither the government guidelines nor the report proves that a particular product complies with a binding global safety standard. The appropriate response is to inspect the specific model and use case, not to treat a broad government statement as an independent audit.

A practical checklist before selecting an AI service

Start by documenting the exact task and what could go wrong. Then request the model version, test methodology, incident response plan, data retention terms, training-data handling, and restrictions on third-party access. For sensitive deployments, demand clarity about encryption, audit logs, human review and vendor security certifications where applicable.

Run local tests using representative examples, difficult edge cases and potentially malicious inputs. Restrict access to critical tools, require human approval for irreversible actions, monitor errors after each model update and provide a fallback process if the system is unavailable or produces uncertain answers.

Avoid misleading national shortcuts

Country of origin may create legal, procurement or security considerations, but it is not a substitute for evidence. Certain government or regulated sectors may restrict particular providers for specific reasons; organizations should check applicable requirements and contracts. A US vendor should be evaluated just as rigorously on its own claims.

A procurement decision is strongest when it separates legal eligibility, the quality of published safety evidence, results of independent testing, operating costs, reliability and exit options. The transparency gap identified in this study is a reason to demand better documentation everywhere, not to conclude that every model from a region behaves identically.

Does 3.6% mean that just 3.6% of AI models are safe?

No. The statistic describes the availability of specific public safety-evaluation disclosures, not the share of models proven safe.

Were all the other Chinese models never tested?

That cannot be concluded. Some may have been tested privately or reported results outside the study's disclosure criteria.

What should an ordinary user do?

Avoid sharing confidential information without understanding privacy terms, verify consequential answers independently and keep human control of irreversible decisions.