Module 3 of the AI Act Provider certification: Art. 10 governance for training, validation and testing data, the bias examination duty, and the narrow Art. 10(5) permission to process special-category data.
Art. 10 is where the Regulation reaches furthest into how systems are actually built, and where retrofitting is most painful. Data governance decisions are made at the start of a project by people who are not thinking about conformity assessment, and they are close to irreversible by the time anyone does.
Scope: which data, and which systems
Art. 10 applies to high-risk systems which make use of techniques involving the training of AI models with data. For those systems it governs the training, validation and testing datasets.
For high-risk systems not developed by training models with data — knowledge-based systems, rules engines that nonetheless sit in Annex III — Art. 10 applies to the testing datasets only. That narrower duty is real and often overlooked, because teams read "data governance" as "training data governance" and conclude it does not apply to them.
The governance practices Art. 10(2) requires
Training, validation and testing datasets must be subject to data governance and management practices appropriate to the intended purpose, concerning in particular:
- the relevant design choices;
- data collection processes and the origin of the data, and in the case of personal data, the original purpose of collection;
- relevant data preparation operations, such as annotation, labelling, cleaning, updating, enrichment and aggregation;
- the formulation of assumptions, notably with respect to the information the data are supposed to measure and represent;
- an assessment of the availability, quantity and suitability of the datasets needed;
- examination in view of possible biases likely to affect health and safety of persons, have a negative impact on fundamental rights, or lead to discrimination prohibited under Union law, especially where data outputs influence inputs for future operations;
- appropriate measures to detect, prevent and mitigate those possible biases;
- identification of relevant data gaps or shortcomings that prevent compliance, and how they can be addressed.
Two of these deserve separate attention.
"The formulation of assumptions, notably with respect to the information the data are supposed to measure and represent." This is asking you to write down your construct validity argument: you are using this feature as a proxy for that quality, and here is why that holds. It is the single most useful item on the list and the least often produced, because it is usually an implicit belief rather than a recorded decision.
"Especially where data outputs influence inputs for future operations." This names feedback loops. A hiring model trained on past hiring decisions, retrained on the decisions it influenced, is the archetype. If your system's outputs re-enter its training data, Art. 10 requires you to have examined what that does.
Quality: relevant, representative, error-free, complete
Art. 10(3): datasets shall be relevant, sufficiently representative, and to the best extent possible free of errors and complete in view of the intended purpose. They shall have the appropriate statistical properties, including, where applicable, as regards the persons or groups on which the system is intended to be used.
The qualifiers are doing real work and are frequently misquoted. It is not "error-free" — it is "to the best extent possible free of errors". It is not "complete" — it is "complete in view of the intended purpose". The Regulation is asking for a defensible, documented judgement, not perfection.
Art. 10(4) adds that datasets shall take into account, to the extent required by the intended purpose, the characteristics or elements particular to the specific geographical, contextual, behavioural or functional setting within which the system is intended to be used.
That is the provision that catches a model validated on one national population and shipped into another. If your intended purpose covers the Union, your representativeness argument has to cover the Union.
Art. 10(5) — the special-category permission
This is the most misunderstood paragraph in the article, in both directions: some teams believe they may never touch special-category data, others treat the paragraph as a general licence.
To the extent strictly necessary for the purpose of ensuring bias detection and correction in relation to high-risk systems, providers may exceptionally process special categories of personal data, subject to appropriate safeguards including:
- bias detection and correction cannot be effectively fulfilled by processing other data, including synthetic or anonymised data;
- the special-category data are subject to technical limitations on re-use and state-of-the-art security and privacy-preserving measures, including pseudonymisation;
- measures to ensure the data are secured, protected, subject to suitable safeguards, including strict controls and documentation of the access, avoiding misuse and ensuring only authorised persons have access with appropriate confidentiality obligations;
- the data are not transmitted, transferred or otherwise accessed by other parties;
- the data are deleted once the bias has been corrected or the personal data has reached the end of its retention period, whichever comes first;
- records of processing activities include the reasons why processing was strictly necessary and why other data could not be used.
Read as a whole, this permits a controlled bias-testing programme with a documented necessity argument and a deletion trigger. It does not permit retaining a demographic-labelled dataset indefinitely because it might be useful for fairness work later.
The Digital Omnibus adjusted the framing of this permission in July 2026; if your compliance position was written before then, re-read the current text before relying on it.
Bias examination is not one number
The Regulation asks for examination of possible biases and appropriate measures to detect, prevent and mitigate them. It does not prescribe a metric, and there is no single metric that could serve — the common fairness criteria are mathematically incompatible with each other except in degenerate cases.
What survives review is: a statement of which groups you examined and why those; which measures of disparity you used and why; what you found; what you did; and what residual disparity remains and why it is judged acceptable. That last clause connects back to Art. 9 residual risk, and it is the honest position — a claim of no bias is neither achievable nor credible.
Where the file gets thin
In review, the recurring gaps are consistent:
- Provenance for data acquired years ago, before anyone anticipated needing it. This is unrecoverable, which is why acquiring data without recording provenance is now a compliance decision rather than an operational shortcut.
- The assumptions statement, because it was never explicit.
- The bias examination for the deployment population rather than the training population.
- Third-party datasets whose licence permits use but whose composition the provider cannot describe.
- Feedback loops, unexamined.
Each of these is cheap at the start of a project and close to impossible at the end. That is the practical argument for settling Module 1's four determinations before development begins.
Check yourself
- Our system is rules-based, not trained. Does Art. 10 apply? — Partly. For systems not developed by training models with data, the requirements apply to the testing datasets.
- We cannot check for ethnic bias because we do not hold ethnicity data. — Art. 10(5) permits processing special-category data exceptionally and strictly for bias detection and correction, subject to a heavy set of conditions including a necessity argument and deletion.
- Our dataset is large, so it is representative. — Representativeness is judged against the population the system is intended to be used on, and against the geographical, contextual, behavioural and functional setting. Size is not the test.
- Our model is retrained on the decisions it influenced. — Art. 10(2) singles this out: bias examination matters especially where data outputs influence inputs for future operations.
Previous: Module 2 — The risk management system (Art. 9) Next: Module 4 — Technical documentation (Art. 11 + Annex IV) →
AI Act meets DORA and NIS2
Is your organisation subject to both the AI Act and DORA? The two regulations intersect on the operational resilience of financial AI systems. Our sister site regulation-dora.eu covers DORA in depth — including what the AI Act adds on top of an existing DORA programme.
The AI Act for financial institutions ↗ Explore regulation-dora.eu ↗Frequently Asked Questions
Art. 10 addresses high-risk systems which make use of techniques involving the training of models with data. For systems not developed by training on data, the requirements apply to the testing datasets only. So a knowledge-based or rules-driven high-risk system does not escape Art. 10 entirely — its testing data is still in scope.
Art. 10(5) permits it exceptionally, to the extent strictly necessary for detecting and correcting bias, and subject to conditions: no other means suffice, technical limits on reuse, state-of-the-art security and pseudonymisation, strict access controls with documented access, and no transmission or transfer to third parties. It is a narrow permission with a heavy compliance envelope, not a general licence.
Datasets must be relevant, sufficiently representative, and to the best extent possible free of errors and complete in view of the intended purpose. Representativeness is judged against the persons or groups on which the system is intended to be used, taking account of the specific geographical, contextual, behavioural or functional setting. It is a question about your deployment population, not about dataset size.
Take compliance further with the AI Act Academy
A free course, a server-graded exam, a verifiable certificate — and the working templates.