The NMPA released the “Draft Guideline on Artificial Intelligence Medical Devices” and “Draft Guideline on Clinical Evaluation of Artificial Intelligence Multi-Disease Auxiliary Decision-Making Medical Device” on September 14 and 15 respectively. Together, they mark an important step in China’s AI medical device regulation. They show that China is moving from general AI software principles toward a more layered system: a horizontal lifecycle and algorithm framework, plus vertical clinical evaluation rules for complex AI products.
For foreign manufacturers, the message is clear: China’s market access path is becoming more structured, more evidence-driven, and less forgiving of unclear algorithms, uncontrolled updates, or weak clinical validation.
This guidelines form an integral part of China’s Digital Health regulatory framework and serves as a supplement to other software & AI related documents, such as those on medical device software, cybersecurity, AI-assisted software, AI-assisted diagnostics and human factors engineering.
Please click HERE for our technical review on AI-aided Software Guideline. The article was published on BioWorld, a Hong Kong-based biotech magazine.
Click HERE for our webinar on “SaMD Registration Requirements: China & US Perspectives”
Click HERE for our webinar on “Key Takeaways and Best Practices of China Human Factor/Usability”
AI Medical Device Guideline
The first draft is a general revision of the AI medical device registration review guidance. It applies to Class II and Class III AI independent software and medical devices containing AI software components, including in vitro diagnostic devices. It also applies to self-developed software, while off-the-shelf software components are referenced.
The guidance defines AI medical devices as devices based on medical device data and using AI to achieve medical purposes. It classifies AI devices from multiple angles: independent software versus software component, auxiliary decision versus non-auxiliary decision, processing/control/safety functions, and different algorithm types. It emphasizes three basic principles: evaluation based on algorithm characteristics, risk-oriented review, and full lifecycle quality control.
- The lifecycle process includes requirements analysis, data collection, algorithm design, verification and validation, and update control.
- Data collection covers acquisition, curation, annotation, and dataset construction. Training, tuning, and test sets must be separated without overlap. Annotation requires qualified personnel, clear rules, arbitration, and quality assessment.
- Algorithm design covers algorithm selection, training, performance evaluation, and influence factor analysis.
- Verification and validation include software verification, software confirmation, and clinical evaluation.
- Update control distinguishes algorithm-driven updates from data-driven updates. Algorithm-driven updates are usually major software updates requiring change registration; data-driven updates may be major if performance changes significantly. Continuous or adaptive learning should generally have the self-learning function turned off or not put into clinical use.
The guidance also addresses cybersecurity and data security, mobile and cloud computing, usability engineering, stress testing, adversarial testing, third-party databases, white-box algorithms, ensemble learning, transfer learning, reinforcement learning, federated learning, generative adversarial networks, cross-modal image synthesis, multimodal products, multi-disease products, large models, algorithm frameworks, and AI chips. For large models, API calls that are not necessary for the medical purpose may be treated as non-medical functions; if necessary, they are treated as off-the-shelf software components. Imported devices must also consider China-foreign differences such as ethnicity, epidemiology, and clinical practice.
Multi-Disease Clinical Evaluation Guideline
The second draft targets AI multi-disease auxiliary decision-making devices. These are products that use AI to provide auxiliary decision suggestions for two or more diseases with different imaging features.
- Covered types: auxiliary triage/referral, auxiliary detection, and auxiliary diagnosis.
- Regulatory status: classification code 21-04-02; management category Class III.
- Typical input models: one image input for multiple diseases, such as chest CT for lung nodules, pleural effusion, and rib fracture; or separate imaging data for independent disease modules, such as brain CTA/MRI for stroke triage and abdominal MRI for pancreatitis triage.
- Not applicable: software predicting disease probability, AI software used with in vitro diagnostic reagents, and AI for auxiliary treatment.
- The core principle is that auxiliary decision functions generally require clinical evaluation based on clinical trials. The core algorithm must have its own clinical trial evidence, usually a diagnostic clinical trial. Non-auxiliary functions, such as structured report generation, image comparison, normal anatomy segmentation, size measurement, and CT value measurement, may be evaluated as secondary endpoints or through comparative clinical evaluation.
A key contribution of this draft is its detailed logic for multi-disease trial design.
- Independent diseases: one case can be used for multiple disease evaluations, such as lung nodules and rib fractures on chest CT.
- Clinically similar diseases with clearly different imaging features: cases may be shared, but differential diagnosis must still be evaluated using a well-controlled testing database or clinical trial.
- Highly similar imaging features requiring distinction: the trial must assess both each disease versus negative cases and disease-to-disease differential diagnosis. In some cases, separate trials are required. For example, lung nodule AI and pneumonia/tuberculosis AI should generally be validated separately because severe infection may obscure the basis for nodule judgment.
The draft also sets strict requirements for study objects, reference standards, and endpoints. Imaging samples must be independent from the product development dataset, collected continuously under clear inclusion and exclusion criteria, and cover a reasonable disease spectrum. Positive and negative definitions should follow clinical guidelines and expert consensus. Primary endpoints should generally include sensitivity, specificity, ROC or derived curves for each target disease, and differential diagnosis endpoints are also required where applicable. Reference standards should be superior to the tested method, such as pathology, confirmed clinical diagnosis, expert panel review, or composite follow-up standards for negative cases. MRMC designs are encouraged, with multiple readers, blinding, training, and washout periods generally of four to six weeks. Testing databases used for differential diagnosis must be independent, authoritative, scientifically sampled, standardized, diverse, closed, and dynamic. They should include adequate rare subtypes and pairwise disease comparisons, prioritizing higher-risk diseases.
Implications for Foreign Manufacturers
Foreign manufacturers should plan early for China’s two-layer AI framework, where multi-disease products are likely Class III and require per-disease clinical evidence for auxiliary functions.
- Avoid one generic trial: trial design must reflect disease relationships, imaging similarity, and differential diagnosis needs, and some diseases require separate trials.
- Secure data governance: data must be independent, ethically collected, diverse, and cover the disease spectrum, with expert annotation and high-quality testing databases.
- Prepare core documents: algorithm research reports, data-quality records, cybersecurity materials, user training plans, and Chinese labeling/warnings are essential.
- Use testing databases carefully: they can support differential diagnosis evaluation but do not replace clinical trials for primary endpoints.
- Manage lifecycle and localization: algorithm updates, continuous learning, large models, and imported-device differences require change control, warnings, and a China-specific regulatory strategy, not a simple US/EU extension.
