GDPR-Compliant AI: The Evidence File Regulators Expect

GDPR-compliant AI is not a property of a model, and it is not a badge on a vendor’s website. It is a set of decisions, each with a record behind it, taken for every processing operation an AI system performs on personal data. In 2026 several of those decisions moved at once: the European Data Protection Board put two sets of guidelines out for consultation until 30 October 2026, the AI Act gained a legal basis for testing bias on sensitive data, a reform of the GDPR itself is still stuck in negotiation, and the first large fine against a generative AI provider was annulled in court. Most guides restate the principles. This one follows what a data protection authority asks to see.

GDPR-compliant AI evidence file: a brass key resting on a stack of paper folders tied with string, a small open padlock beside it

Key takeaways

  • The GDPR applies at four moments in an AI system’s life: collecting training data, training, deployment (prompts, retrieval, logs) and outputs. Each needs its own legal basis and its own record.
  • Legitimate interest can support AI training, but only through the three-step test of EDPB Opinion 28/2024, and a model trained on personal data is not presumed anonymous.
  • EDPB Guidelines 03/2026 on web scraping are open for comment until 30 October 2026. Ignoring robots.txt or ai.txt now counts against you in the balancing test.
  • Since 27 July 2026, Article 4a of the AI Act lets providers and deployers process special category data to detect and correct bias, under strict necessity.
  • The Digital Omnibus and its Article 88c are not law. GDPR-compliant AI is built on the regulation in force.

What GDPR-compliant AI means, and what it does not

The GDPR regulates the processing of personal data, not a technology. An AI system is therefore never compliant in the abstract. GDPR-compliant AI means that each processing operation the system performs, from the first dataset to the last logged prompt, has a controller, a purpose, a legal basis, a retention period and a way for people to exercise their rights. Three readings of GDPR-compliant AI are common, and all three are wrong:

  • “Our vendor is GDPR compliant.” A vendor’s statement covers the vendor’s own processing. It says nothing about your legal basis, your privacy notice or how long you keep prompts.
  • “The data stays in the EU.” Hosting location answers the transfer question in Chapter V and nothing else.
  • “The data was public.” Public availability is not one of the six legal bases in Article 6.

The stakes are those of any GDPR infringement: under Article 83(5), fines reach EUR 20 million or 4 percent of total worldwide annual turnover, whichever is higher. The starting point for GDPR-compliant AI is knowing which systems exist at all, which is why an AI inventory comes before any legal analysis, and why AI accountability starts with a named owner per system.

Two regulations, one system: the GDPR and the AI Act

Article 2(7) of the AI Act states that Union data protection law applies to personal data processed in connection with AI systems and that the AI Act does not affect the GDPR. The two apply in parallel, and they use different role tests. The AI Act speaks of providers and deployers. The GDPR speaks of controllers and processors. A deployer is usually a controller for the purposes it pursues with the system. A model provider is often a controller when it trains and a processor when it runs inference on a customer’s behalf. The two labels do not map one to one, so the allocation has to be written down per system. GDPR-compliant AI rests on that allocation: a contract signed under the wrong role protects nobody.

Four moments where the GDPR applies to an AI system

Most privacy programmes treat an AI tool as one processing activity. A regulator assessing GDPR-compliant AI does not. The questions change at each stage of the lifecycle, and so does the evidence. <table header-row=”true”> <tr> <td>Moment</td> <td>GDPR questions</td> <td>Record to keep</td> </tr> <tr> <td>Collecting training data (scraping, reuse of first-party data, purchased datasets)</td> <td>Legal basis, compatibility of purpose, Article 14 information, special categories</td> <td>Source list with dates, legitimate interest assessment, exclusion rules</td> </tr> <tr> <td>Training and fine-tuning</td> <td>Minimisation, retention, security, whether the resulting model is anonymous</td> <td>Anonymity assessment of the model, extraction and membership-inference test results</td> </tr> <tr> <td>Deployment (prompts, retrieval corpus, logs, memory)</td> <td>Controller and processor roles, Article 28 contract, retention of prompts and logs, transfers</td> <td>Data processing agreement, retention schedule, transfer assessment</td> </tr> <tr> <td>Outputs and decisions</td> <td>Accuracy, Article 22, access and explanation rights, erasure and objection</td> <td>Article 22 assessment, human review design, rights-handling procedure</td> </tr> </table> Most organisations document only the third row, because that is where the vendor contract sits and where procurement asks questions. Regulators start with the first. The Italian, Dutch and Irish cases of the last three years all began with how the data was collected, not with how the product was configured. For an organisation that buys rather than builds, rows one and two do not disappear. They become questions for the supplier, and the answers become part of your own GDPR-compliant AI file. A data governance framework that already tracks lineage and ownership for analytics data can carry these records with little extra structure.

Legal basis for GDPR-compliant AI: what the EDPB accepts in 2026

Consent is rarely the route to GDPR-compliant AI when a model is trained at scale: it cannot be collected from people the controller has no relationship with, and it can be withdrawn. The realistic basis is legitimate interest under Article 6(1)(f), and the European Data Protection Board has now said twice how it expects that basis to be argued.

Opinion 28/2024: legitimate interest and model anonymity

Opinion 28/2024, adopted on 17 December 2024 at the request of the Irish authority, settles three points.

  1. Legitimate interest is available for developing and deploying AI models, through a three-step test: the interest is lawful, clearly articulated and real; the processing is necessary, with no less intrusive way to reach the same result; and the interest is not overridden by the rights and reasonable expectations of the people concerned.
  2. A model trained on personal data is not anonymous by default. Anonymity must be demonstrated case by case: the likelihood of extracting personal data from the model, and of obtaining it through queries, has to be insignificant.
  3. Unlawful training travels downstream. A deployer is expected to have carried out an appropriate assessment, as part of its accountability, that the model it uses was not developed by unlawfully processing personal data.

The third point makes GDPR-compliant AI a concern for every buyer of a third-party model. It turns the supplier’s training practices into a question for vendor due diligence, with a written answer kept on file.

Guidelines 03/2026: web scraping for generative AI

In July 2026 the EDPB announced two new texts, both open for public consultation until 30 October 2026. Guidelines 03/2026 are the first GDPR framework written specifically for scraping the web to train generative models. According to Reed Smith’s reading, a final version is expected before the end of 2026. For GDPR-compliant AI built on scraped data, the draft says the following.

  • Technical signals count. Robots.txt, ai.txt, CAPTCHAs and login walls shape what people can reasonably expect, so ignoring them weighs against the controller in the balancing test.
  • Untargeted crawling is riskier than targeted scraping, because the controller knows less about what it collects.
  • Minimisation starts before collection: precise collection criteria, exclusion of sites that are sensitive by nature, then filters for identifiers and pseudonymisation afterwards.
  • Incidental special category data is not automatically unlawful, provided it is truly incidental and surrounded by safeguards across the lifecycle: filtering, prompt deletion, extraction-resistance testing and output monitoring. The Board draws on the Court of Justice’s GC and Others ruling (C-136/17).
  • Transparency survives. Individual notice may be excused under Article 14(5)(b), but a detailed public notice remains, ideally listing scraped sources with collection dates, with an opt-out available before collection.

Guidelines 02/2026 on anonymisation, announced the same day, set three criteria for calling data anonymous: no record isolation, no linkage and no inference. They refer to the Court’s judgment in EDPS v SRB of 4 September 2025. Any organisation that claims its model or its training set is anonymous should test that claim against the three criteria now, while the text can still be commented on.

Sensitive data and bias testing: Article 4a of the AI Act

Bias testing is the point where GDPR-compliant AI long sat on a contradiction. To check whether a system discriminates by ethnic origin, health or disability, you need data on those attributes, and Article 9 of the GDPR prohibits processing them by default. Until July 2026 the AI Act offered a narrow way out in Article 10(5), reserved for providers of high-risk systems. Regulation (EU) 2026/1744, the Digital Omnibus on AI, was published in the Official Journal on 24 July 2026 and entered into force on 27 July 2026. It repealed Article 10(5) and moved the rule into a new Article 4a, extended to providers and deployers of AI systems and models. The basis is exceptional and applies only where strictly necessary. The conditions are cumulative:

  • bias detection and correction cannot be achieved with synthetic or anonymised data;
  • the data is subject to technical limits on re-use, state-of-the-art security and pseudonymisation;
  • access is restricted and controlled;
  • the data is not transmitted to other parties;
  • it is deleted once the bias is corrected or the retention period ends;
  • the reasons are entered in the records of processing activities.

Two limits matter for GDPR-compliant AI. Article 4a creates no duty to test for bias; it removes an obstacle for those who do. And it does not replace Article 6 of the GDPR: a lawful basis is still required for the processing itself. Whether the balance is right is contested. In a June 2026 IAPP report, an EDPB legal officer listed bias detection among the topics the coming joint guidelines will address, while an industry privacy officer said the balance was not yet right. For the method itself, see our guide to AI bias.

Automated decisions: Article 22 after SCHUFA and Dun & Bradstreet

Article 22 gives people the right not to be subject to a decision based solely on automated processing that produces legal effects or similarly significant ones. Three exceptions exist (necessity for a contract, authorisation by law, explicit consent), each with safeguards: human intervention, the chance to express a view, and the ability to contest. Two Court of Justice rulings have widened the reach of that article. In SCHUFA (C-634/21, 7 December 2023), the Court held that producing a score is itself an automated decision when a third party draws strongly on it. A human who rubber-stamps a score does not take the decision out of Article 22. In Dun & Bradstreet Austria (C-203/22, 27 February 2025), the Court read Article 15(1)(h) as a right to an explanation of the procedure and principles actually applied to the person’s data to reach the result, in an intelligible form. Handing over an algorithm is not an explanation. Trade secrets do not justify a blanket refusal: the disputed information goes to the authority or the court, which strikes the balance. Three practical consequences follow for GDPR-compliant AI.

  1. Classify every AI use that scores or ranks people, including scores you pass to someone else.
  2. Design human review that can change the outcome, and log when it does. A review that never overturns anything is evidence against you.
  3. Prepare the explanation template before the first request arrives, not after.

The AI Act adds its own layer. Article 14 requires human oversight of high-risk systems, and Article 86 gives affected persons a right to an explanation of decisions taken with Annex III high-risk systems. One design, documented once, can serve both regimes if explainability is treated as a requirement from the start.

DPIA and FRIA for GDPR-compliant AI: one file, two duties

Article 35 of the GDPR requires a data protection impact assessment where processing is likely to result in a high risk to individuals. Most AI systems that profile, score, monitor or process sensitive data at scale qualify. The AI Act connects to that duty in two places. Article 26(9) tells deployers of high-risk systems to use the instructions the provider supplies under Article 13 when they carry out their DPIA. Article 27 requires certain deployers (public bodies, private entities providing public services, and those using AI for credit scoring or for life and health insurance pricing) to run a fundamental rights impact assessment. Under Article 27(4), where a DPIA already covers one of the required elements, the FRIA complements it. After Regulation (EU) 2026/1744, the obligations for Annex III high-risk systems, Article 27 included, apply from 2 December 2027. The DPIA duty has applied since 2018 and does not wait. At an IAPP conference in June 2026, an EDPB legal officer summarised the relation this way: if you need a FRIA you most likely need a DPIA, and the reverse is not necessarily true. Joint guidelines from the Commission and the EDPB on the interplay between the two regulations are announced, with a final text possible by the end of 2026. They were not published as of early October 2026. The practical answer for GDPR-compliant AI is one assessment file per system, with a DPIA core and a FRIA extension that share the same system description and the same risk register. Our guides to the privacy impact assessment and the AI impact assessment set out the method, and the high-risk classification guide tells you whether Article 27 applies to you at all.

What is still moving: the Digital Omnibus and Article 88c

On 19 November 2025 the Commission proposed the Digital Omnibus, which would amend the GDPR in two ways that matter here. A new Article 88c would state that processing personal data to develop and operate AI systems or models can rest on legitimate interest, with enhanced transparency and an unconditional right to object. And Article 4 would receive a contextual definition of personal data: information would not be personal data for an entity that cannot reasonably identify the person. In Joint Opinion 2/2026 of 10 February 2026, the EDPB and the EDPS supported the aim of simplification and firmly opposed the change to the definition of personal data. As of September 2026 the file had not advanced far. The Council had no general approach; a revised Presidency compromise was discussed in a preparatory group on 11 September 2026. A draft attributed to the Irish Presidency, published by the NGO noyb and reported on 21 September 2026, renumbers the provision as Article 88bis; noyb criticised it as making the use of personal data for AI lawful by default. The Parliament had no committee position. The conclusion for a team working on GDPR-compliant AI is short. Article 88c is a proposal. A legal basis analysis that relies on it has no legal basis. Even the Commission’s own text keeps the balancing test, transparency and the right to object. Enforcement is unsettled as well:

  • Italy’s Garante fined OpenAI EUR 15 million in December 2024. On 18 March 2026 the Court of Rome annulled the decision in its entirety.
  • The Dutch authority fined Clearview AI EUR 30.5 million in 2024 over its facial image database.
  • The Garante fined the company behind the Replika chatbot EUR 5 million in 2025.

None of this lowers the bar for GDPR-compliant AI. It means that what protects an organisation is the quality of its own reasoning on file, not a bet on how the next case will end. For a working method that applies beyond the United Kingdom, the ICO’s AI and data protection guidance remains the most complete regulator-written toolkit in English.

The GDPR-compliant AI evidence file: eight records a regulator can ask for

An authority’s first letter rarely asks whether you comply. It asks for documents. These are the eight that make up the file for GDPR-compliant AI.

  1. Inventory entry. Each AI system is linked to its entry in the records of processing activities (Article 30). An AI registry that holds both avoids two lists drifting apart.
  2. Role allocation and contract. Who is controller, who is processor, per system, with the matching Article 28 agreement or the joint controller arrangement under Article 26.
  3. Legal basis decision. One per purpose, with the legitimate interest assessment where that basis is relied on.
  4. Impact assessment. The DPIA, extended by the FRIA where Article 27 of the AI Act applies, plus the screening decision for systems that did not need one.
  5. Transparency notices. Articles 13 and 14 information, including the sources of training data where you trained or fine-tuned.
  6. Rights procedure that reaches the model. What happens on an erasure or objection request: output filters, suppression lists, the decision to retrain or not, who takes it and within what delay.
  7. Article 22 file. The assessment of each scoring or ranking use, the design of human review and the explanation template.
  8. Security and supplier evidence. Article 32 measures, retention of prompts and logs, the transfer assessment, and the supplier’s assurance on lawful training or the anonymity assessment of the model. Our guide to AI system documentation shows how these fit with the AI Act technical file.

In a GDPR-compliant AI file, each record has an owner and a date. Together they let you answer without a reconstruction exercise, and they are the same documents an AI audit will sample.

FAQ

What is GDPR-compliant AI in practice? It is an AI system for which every processing of personal data has a named controller, a documented purpose and legal basis, a retention period, appropriate security and a working way for people to exercise their rights. The test is applied operation by operation: collecting training data, training, running the system and using its outputs. A model is not compliant on its own, and a vendor cannot make your use compliant for you. In practice, GDPR-compliant AI is recognisable by its file: an inventory entry, a role allocation, a legal basis decision, an impact assessment, notices, a rights procedure and supplier evidence. Does GDPR-compliant AI require consent to train a model on personal data? No. The EDPB confirmed in Opinion 28/2024 that legitimate interest can be a valid basis for developing and deploying AI models, provided the three-step test is passed: a lawful and real interest, necessity, and a balancing against the rights and reasonable expectations of the people concerned. Consent is rarely workable at scale, because it cannot be collected from people you have no relationship with and can be withdrawn at any time. Special category data is a separate question: it needs an exception under Article 9, such as the new Article 4a of the AI Act for bias detection. Does a vendor’s certification make my use of an AI tool GDPR compliant? No. GDPR-compliant AI cannot be bought: a certification or a compliance statement covers the vendor’s own processing and security. You remain the controller for the purposes you pursue with the tool. That means an Article 28 contract, your own legal basis, your own privacy notice and a retention rule for prompts and outputs. Check two points in particular: whether the vendor uses your prompts to train its models, which would make it a controller for that purpose, and what the vendor can show about how the model was trained, since Opinion 28/2024 expects deployers to have assessed that. Does GDPR-compliant AI require a DPIA for every system? Not for every one, but for most systems that profile, score or monitor people, or that process sensitive data at scale. Article 35 applies where processing is likely to result in a high risk, and national authorities publish lists of processing types that always require an assessment. The safe practice is to screen every AI system and record the result either way. A written decision that no DPIA was needed, with reasons, is itself evidence. For high-risk systems under the AI Act, the provider’s instructions for use feed directly into the assessment. How does GDPR-compliant AI fit with the EU AI Act? The two regimes apply in parallel, so GDPR-compliant AI and AI Act conformity are separate questions. Article 2(7) of the AI Act leaves the GDPR untouched. The GDPR governs the personal data an AI system processes; the AI Act governs the system as a product and its use. Roles differ: provider and deployer on one side, controller and processor on the other. The two meet at defined points: the DPIA and the fundamental rights impact assessment, human oversight and Article 22, and Article 4a on sensitive data for bias testing. Joint guidelines from the Commission and the EDPB on this interplay have been announced. Will the Digital Omnibus make AI training lawful by default? Not as of October 2026. The Commission’s proposal of 19 November 2025 would add an Article 88c stating that AI development and operation can rest on legitimate interest. The Council has not agreed a position, the Parliament has no committee position, and the EDPB and EDPS have voiced reservations on parts of the package. Even the Commission’s text keeps the balancing test, enhanced transparency and an unconditional right to object. Until a regulation is adopted and published, GDPR-compliant AI rests on Article 6 as it stands and on the EDPB’s reading of it.

Conclusion

The GDPR did not change in 2026. Its reading for AI did: two sets of EDPB guidelines, two Court of Justice rulings on automated decisions, a new Article 4a in the AI Act and a reform still being negotiated. The organisations that come through an investigation well are those that can show, for each system, the four moments where personal data is processed, the decision taken at each and the record behind it. GDPR-compliant AI is that file, kept current. The consultation on Guidelines 02/2026 and 03/2026 closes on 30 October 2026, which leaves a few weeks to test your own position against the drafts. AI Sigil links each AI system in the registry to its legal basis, its assessments, its notices and its supplier evidence, so that the answer to an authority is a report, not a reconstruction.

GDPR-Compliant AI: The Evidence File Regulators Expect

GDPR-compliant AI in 2026: legal basis after EDPB Opinion 28/2024, web scraping guidelines, Article 4a bias data, DPIA vs FRIA, and the 8 records to keep.

Agentic AI Security: The OWASP Top 10 as Audit Evidence

Agentic AI security read as an auditor would: the ten OWASP agentic risks mapped to EU AI Act duties, the records that prove control, and the deadlines.

NIST AI 600-1: The Generative AI Profile as an Evidence Map

NIST AI 600-1 explained: the 12 generative AI risks, 211 suggested actions, its 2026 status, the Texas TRAIGA defense and the EU AI Act mapping.

Model Drift: When a Compliant Model Stops Complying

Model drift quietly turns a validated AI model into a non-compliant one. Learn the types, detection metrics and what EU AI Act Articles 15 and 72 require.

HIPAA Compliance Software: The 2026 AI-Era Buyer’s Guide

HIPAA compliance software was built for systems that store PHI, not for systems that infer from it. What a 2026 tool must cover, and what to demand.

AI Inventory: What Regulators Expect to Find in It

An AI inventory is the artefact every AI rule assumes. See which clauses compel one (EU AI Act, NIST, ISO 42001, OMB) and the fields each expects.