Annex 22: what it requires of artificial intelligence in GMP
The draft of the new Annex 22 to the EU GMP Guide sets out how an AI model involved in a critical process is validated. This guide summarizes it section by section, from the official text, and shows where each requirement lands in the validation documentation.
It is still a draft
Checked against the official source.
As of October 8, 2026, Annex 22 is not among the annexes in force in EudraLex Volume 4, which lists Annexes 1 to 21. The reference text is the draft published by the European Commission for consultation, together with the revision of Annex 11 and Chapter 4.
A draft is not binding, but it shows what inspectors are going to ask. Designing to it today avoids redoing the validation when the final text is published. We will update this guide when that happens.
Where the text is heading
What has been discussed since the consultation.
The 2025 public consultation showed support for allowing LLMs and generative AI in GMP applications. The Annex 22 drafting group has therefore discussed widening the scope to dynamic or adaptive models, probabilistic models and generative AI, provided they meet the annex and rest on a fully documented, risk-based control strategy.
To work out that strategy, the European Medicines Agency brought together industry experts on June 30, 2026 and published the workshop report on October 2. Its conclusions point to a technology-neutral annex: what decides whether a use is acceptable is the intended use, the risk and the effectiveness of the controls, not the type of model.
It is a direction, not a new text: until the final annex is published, the draft remains the reference.
Which models it covers and which are left out
- Critical applications
- Computerized systems used in the manufacturing of medicinal products and active substances where an AI model has a direct impact on patient safety, product quality or data integrity, for example to predict or classify data.
- Machine learning
- Models that obtained their functionality through training with data rather than being explicitly programmed. A system may combine several, each automating a specific process step.
- Static models only
- Models that do not adapt their performance during use by incorporating new data. Dynamic models, which learn continuously, should not be used in critical applications.
- Deterministic output only
- Identical inputs give identical outputs. Models with a probabilistic output should not be used in critical applications.
- Generative AI and LLMs
- Excluded from critical applications in the published draft. In non-critical uses, qualified and trained personnel are responsible for making sure each output is suitable (human-in-the-loop), and the principles of the annex may be considered where applicable.
- Relationship with Annex 11
- It is additional guidance to Annex 11 for computerized systems with embedded AI models. It does not replace it.
What each section asks for and where it lands in the documentation
Summary of sections 2 to 10 of the draft. The third column shows the validation document where the evidence is kept.
| Section | What the draft asks for | Where the evidence goes |
|---|---|---|
| 2. Principles | Close cooperation between process subject matter experts (SMEs), QA, data scientists, IT and consultants. The regulated user reviews the documentation even when the model comes from a supplier, and the effort is based on risk. | Risk assessment (RA) |
| 3. Intended Use | A detailed description of the task and of the input data, with all common and rare variations, its limitations and possible biased inputs, divided into subgroups where applicable. If a person makes the decision, it includes the operator's responsibility. A process SME approves it before acceptance testing. | Model documentation: intended use |
| 4. Acceptance Criteria | Test metrics suited to the intended use, and acceptance criteria approved before testing. Never below the performance of the process the model replaces. | Model documentation: test plan |
| 5. Test Data | Representative and stratified, sufficient in size for adequate statistical confidence, and with verified labeling. Pre-processing and exclusions are justified. Generating them with generative AI is not recommended. | Model documentation: test data |
| 6. Test Data Independency | Test data are not used in development, training or validation. Access control with an audit trail, a record of which data were used, when and how many times, and staff independence or the four-eyes principle. | Model documentation: test data access log |
| 7. Test Execution | A test plan approved before starting. Any deviation or unmet criterion is documented, investigated and justified. The documentation is retained like any other GMP record. | Model documentation: test plan and report |
| 8. Explainability | During testing, the system records which features in the data contributed to each decision, using techniques such as SHAP, LIME or heat maps, and those features are reviewed when the results are approved. | Model documentation: test report |
| 9. Confidence | The confidence score of each prediction is logged and a threshold is set. Below it, the model can flag the outcome as “undecided” instead of risking an answer. | Technical specification (TS) |
| 10. Operation | Change control and configuration control before deployment, regular monitoring of performance and of drift in the input data, and records of the human review. | Model monitoring plan and plan for maintaining the validated state |
Checklist for a model in a critical GMP process
Before taking a model to production.
Checklist for a model in a critical GMP process
Your ticks are kept only in this browser.
Annex 22: frequently asked questions
About the draft Annex 22.
No. As of October 8, 2026 it is still a draft: it is not among the annexes of EudraLex Volume 4, which run up to Annex 21. In the meantime, it serves as a reference for what inspectors expect.
Under the published draft, not in critical applications: it excludes generative AI and LLMs, as well as dynamic models and models with a probabilistic output. In non-critical uses they can be used if qualified personnel are responsible for making sure each output is suitable for the intended use, with a human in the loop. Since the consultation, the drafting group has been considering allowing them in GMP applications with a risk-based control strategy; until the final text is published, the draft remains the reference.
No. The draft presents itself as additional guidance to Annex 11 for computerized systems with embedded AI models. The system is still validated under Annex 11, and the model adds the Annex 22 requirements.
Metrics suited to the intended use. For a model that classifies, it gives as examples a confusion matrix, sensitivity, specificity, accuracy, precision and F1 score. Acceptance criteria are approved before testing and must be at least as high as the performance of the process the model replaces.
That the data used to test the model have not been used to develop, train or validate it. The draft asks for access to those data to be controlled with an audit trail, for a record of when and how many times they are used, and for staff who have seen them not to train the same model, or to do so only in a pair with someone who hasn't (the four-eyes principle).
Do you have a model that should meet Annex 22?
We review it against the draft and tell you what evidence is missing before the next inspection.