Prefer to listen? Click play for AI narration

For most of the past decade, GDPR compliance meant getting the fundamentals right: a documented legal basis, a tidy record of processing activities, a cookie banner that actually worked, a breach notification procedure nobody had to improvise under pressure. Businesses that invested in those fundamentals could reasonably call themselves compliant. In 2026, that's no longer enough. Artificial intelligence has become the single largest source of GDPR risk, and European regulators are showing — through guidance, investigations, and enforcement actions — that AI systems don't get a pass on the principles that have governed data protection since 2018. If anything, AI raises the stakes, because these systems can generate new, harmful outputs from personal data in ways the regulation's original drafters never had to imagine.

GDPR was written to be technology-neutral, and in principle its core obligations apply just as much to a machine learning model as to a spreadsheet of customer records. In practice, AI complicates nearly every one of them. Training data is often collected at scale, sometimes scraped from public sources, and rarely gathered with a specific, documented purpose in mind — which sits uneasily with purpose limitation and data minimization. Models can memorize and later reproduce personal data in ways nobody can fully predict or control. And once a model is deployed, its outputs — recommendations, classifications, generated images or text — can themselves constitute new processing of personal data, with new risks attached. Regulators aren't treating AI as a separate question; they're treating it as an extension of existing GDPR obligations, applied with particular rigor. Three requirements in particular have moved to the center of enforcement activity: legitimate interest assessments for training data, mandatory impact assessments for biometric processing, and proof that training data was lawfully obtained in the first place.

Legitimate Interest Is Not a Free Pass

Many organizations building or fine-tuning AI models have leaned on "legitimate interests" as the lawful basis for using personal data in training, particularly where getting individual consent from every data subject would be impractical. Regulators haven't rejected that approach outright — but they've made clear it isn't a blanket justification. It requires a documented, defensible assessment.

A proper legitimate interest assessment for AI training data needs to establish three things: that the interest being pursued is genuine and lawful, that the processing is necessary to achieve it (with no less intrusive alternative available), and that the interest isn't overridden by the rights and freedoms of the people whose data is being used. For AI specifically, that means being able to show why a particular dataset was necessary, why anonymization or synthetic data weren't viable substitutes, and what safeguards limit the impact on individuals. Organizations that can't produce this documentation when asked are increasingly finding themselves on the wrong side of an inquiry. We covered exactly how tight this test has gotten in our look at the EDPB's new web-scraping guidelines — old datasets don't get a pass just because the rules didn't exist when they were collected.

Biometric Processing and the DPIA Requirement

Biometric data — facial recognition, voiceprints, gait analysis, and similar identifiers — sits in a special category under GDPR as sensitive personal data, and AI has made biometric processing dramatically more common and more powerful. Facial recognition systems, emotion-detection tools, and voice-cloning technology all depend on processing this category of data at scale.

Because of the heightened risk, a Data Protection Impact Assessment isn't optional for these systems — it's mandatory before processing begins. A DPIA needs to describe the processing operations and their purposes, assess whether the processing is necessary and proportionate, identify risks to individuals, and set out the measures taken to address them. For biometric AI, that assessment needs to grapple honestly with questions that are easy to skip over in a rush to ship: How accurate is the system across different demographic groups? What happens if it misidentifies someone? Can people opt out, and what alternative is offered if they do? Regulators increasingly request these assessments as a first step in any inquiry into a biometric AI deployment — and the absence of one is treated as a significant compliance failure in its own right, independent of whatever the underlying system actually does.

Feb 17, 2026
Ireland's DPC opens a formal GDPR inquiry into X over Grok deepfakes
4%
Of global annual revenue — the maximum GDPR fine exposure in that inquiry
0
Grace period for biometric AI deployed without a DPIA
6
Practices that separate prepared businesses from exposed ones

Lawful Sourcing of Training Data

The third pillar of AI-specific GDPR risk is more basic, but arguably harder to fix after the fact: was the training data lawfully obtained in the first place? Large language models and image generators are frequently trained on data scraped from the open web, licensed datasets, or historical company records collected for entirely different purposes. Each source carries its own compliance questions. Scraped data may have been collected without any lawful basis for the scraping itself, let alone for repurposing it into AI training. Licensed datasets may not have been collected with consent informed enough to support downstream AI use. And internal records gathered for one business purpose years ago may not have a valid basis for repurposing into a training set today.

This is a difficult problem to fix retroactively — a model already trained on unlawfully sourced data can't simply have that data "removed" the way a database record can be deleted. That's precisely why regulators are pushing organizations to document data provenance before training begins, not after a complaint arrives: a clear paper trail showing where each dataset came from, what lawful basis applied to its collection, and whether that basis extends to AI training use.

💬
GDPRGard Product
An AI tool built to answer the questions regulators ask

GDPRChat runs on your own business's FAQs, policies, and product information — never on scraped third-party data or customer conversations. When someone asks what data trained it and under what legal basis, you have a documented answer, not a shrug.

See GDPRChat →

When AI Outputs Themselves Become the Violation

Perhaps the most significant shift in 2026 enforcement is the recognition that GDPR violations can arise not just from how personal data is collected or used to train a model, but from what a deployed model produces. This isn't hypothetical. On February 17, 2026, Ireland's Data Protection Commission — which leads GDPR enforcement against X given the company's European headquarters in Dublin — opened a formal statutory inquiry into the platform's Grok chatbot, after reports that users were prompting it to generate non-consensual, sexualized deepfake images of real people, including, in some reported cases, children.

The inquiry examines whether X handled EU and EEA users' data lawfully when it deployed Grok's image-generation tools, whether appropriate safeguards were in place, and whether the risks of misuse were properly assessed before launch. The legal theory underneath it is significant: using a real person's photograph as the basis for generating a sexualized or otherwise harmful synthetic image can itself constitute unlawful processing of that person's personal data — independent of any separate criminal or platform-policy violations involved. In other words, the harm doesn't only occur upstream, in how training data was gathered. It can occur downstream, at the moment of generation, in what the deployed system does with someone's likeness. X restricted Grok's image generation to paying subscribers after the reports emerged, but the inquiry — which carries potential fines of up to 4% of X's global annual revenue — is proceeding regardless.

For any organization building or deploying generative AI, that's a critical distinction. Safeguards can't stop at the training pipeline; they need to extend to what the model is capable of producing at inference time, and what happens when a user deliberately tries to misuse it.

Six Best Practices for Getting Ahead

Across regulatory guidance, enforcement actions, and industry analysis, six practices consistently separate the organizations managing AI-related GDPR risk well from those left exposed.

  1. Build privacy in from the start, not after the fact. Privacy-by-design is a GDPR obligation, not a suggestion — for AI, that means privacy reviews at the dataset-selection and model-architecture stages, not just before shipping.
  2. Map how personal data actually flows through the system. From ingestion, through training or fine-tuning, to inference and output. Without this map, most organizations can't answer a regulator's basic question: where has this person's data been, and what's been done with it?
  3. Document a specific purpose for every processing activity. "Improving our AI models" isn't specific enough. Regulators expect a stated purpose for a given dataset, with its use constrained to that purpose.
  4. Treat DPIAs as a gate, not a formality. Especially for biometric systems, large-scale profiling, and automated decisions with legal or similarly significant effects — complete the assessment before deployment decisions are locked in, not after.
  5. Be able to explain the decision, not just that one was made. GDPR gives people meaningful rights around automated decision-making. Organizations increasingly need to explain, in plain language, how a system arrived at a decision that affects someone.
  6. Monitor continuously, not annually. AI systems drift, get misused in ways their designers didn't anticipate, or produce harmful outputs that only surface once deployed at scale — as the Grok inquiry illustrates. A yearly compliance review is too infrequent to catch that before it becomes a regulatory incident.

AI isn't going to become less central to how organizations process personal data, and regulators aren't going to become less attentive to how it's used. The organizations that navigate this well are the ones treating GDPR compliance as an ongoing engineering and governance discipline built into AI systems from the outset — not a legal checklist applied after the fact. Regulators have now shown they'll act not only against how data is collected for training, but against what deployed systems actually generate. That shift in mindset isn't optional anymore.

Book a free consultation with GDPRGard →

Sources

Share this article

Also read:

⚠️ This article is for informational purposes only and does not constitute legal advice. For complex compliance situations, consult a qualified data protection professional. GDPR and EU AI Act requirements are subject to ongoing regulatory guidance — verify current obligations with your legal adviser.