Prefer to listen? Click play for AI narration

For years, the AI industry operated on a comfortable assumption: if data is sitting on the open web, it's fair game to scrape and feed into a training set. No consent needed, no notice required — it's "publicly available," so it's fine.

On July 8, 2026, the European Data Protection Board (EDPB) closed that loophole for good. Alongside new guidelines on anonymisation and blockchain, it adopted its Guidelines on Web Scraping in the Generative AI Context — and the message is unambiguous: GDPR has always applied to web-scraped personal data used for AI training, and regulators are now spelling out exactly what compliance requires.

What Actually Changed

The guidelines are open for public consultation until October 30, 2026, but the direction of travel is already clear, and EU data protection authorities are expected to enforce against it. Three requirements stand out.

Legitimate interest can't be assumed — it has to be assessed, every time. Developers can no longer treat "the data was public" as a legal basis on its own. Each scraping operation needs a genuine, case-by-case legitimate-interest assessment weighing the developer's interest in the data against the rights and reasonable expectations of the people it describes.

Data minimization has to happen before scraping starts, not after. Filtering out irrelevant or excessive data once it's already been vacuumed into a training set isn't good enough. The guidelines expect minimization built into the collection process itself — scraping only from reliable, identifiable sources, timestamping what's collected, and keeping special-category data (health, political views, sexual orientation, and similar) out from the start rather than scrubbing it later.

Old datasets don't get a pass. This is the detail catching the most people off guard. Companies can't point to data collected years ago and argue the rules didn't exist yet. The balancing test has to reflect the circumstances at the time of collection — which means developers without solid documentation of how and why they gathered older datasets are now sitting on a compliance gap they can't easily backfill.

Jul 8, 2026
EDPB adopts guidelines on web scraping for generative AI
3
New requirements for lawful AI training-data scraping
Oct 30, 2026
Public consultation deadline
0
Grace period for datasets collected before the guidelines

The Anonymization Trap

Some developers have argued that once personal data is baked into a trained model, it's effectively anonymous and GDPR no longer applies. The EDPB's companion guidelines on anonymisation close that door too, setting a strict three-part test: data only counts as anonymous if no individual can be singled out, no data can be linked to other information about them, and no individual can be identified through inference.

That's a high bar — arguably too high for how large language models actually work. Researchers have repeatedly shown that models can "regurgitate" verbatim snippets of their training data, including personal information, under the right prompting. Reporting on the guidelines has been blunt about the consequence: few AI systems, if any, can currently meet this anonymization standard. In practice, that means much of the data inside today's large models likely still counts as personal data under GDPR — with all the obligations that implies.

Who's Affected

Everyone building or fine-tuning models on public web data. That includes the largest labs — OpenAI, Meta, and Google are explicitly named in coverage of the guidelines — but the obligations don't scale down for smaller players. A startup fine-tuning an open-source model on scraped forum posts or a research team building a smaller domain-specific model faces the same legitimate-interest, minimization, and documentation requirements as a frontier lab. If anything, smaller organizations are more exposed, since they typically have fewer resources to build the documentation and impact-assessment processes regulators will expect to see.

💬
GDPRGard Product
Built on your content, not scraped third-party data

GDPRChat is an AI customer chat widget trained on your own business's FAQs, policies, and product information — not on scraped web data of unclear provenance. When a customer (or a regulator) asks where the training data came from, you'll have an answer.

See GDPRChat →

Why This Matters Even If You Don't Train Models Yourself

Most small and mid-sized businesses aren't building foundation models — but plenty are fine-tuning them, deploying third-party AI tools, or embedding AI features from vendors into their own products. That doesn't put you outside the blast radius. If a vendor's model was trained on improperly scraped data, the compliance risk doesn't stay contained to the vendor. Businesses that deploy AI tools are increasingly expected to do basic due diligence on how those tools were built — what data trained them, whether there's a documented legal basis, and whether the vendor can answer straightforward questions about data provenance.

That's a very different question from "does this AI tool work well," and it's one worth asking before signing a contract, not after a regulator asks it for you.

The Practical Takeaway

Three things worth doing now, whether you're training models or buying tools that were:

  1. Ask AI vendors directly what legal basis was used for their training data, and whether they can document it.
  2. Be skeptical of "anonymized" claims about training data — the EDPB's bar is high, and vague reassurance isn't a substitute for a real answer.
  3. Treat "trained before the guidelines existed" as no defense, because regulators won't either.

The AI Act gets more attention because it's new. But GDPR was already fully in force, and the EDPB has just confirmed, in detail, that it always covered how AI gets built — not just how it's used.

It's also why GDPRGard built GDPRChat the way it did: trained on your own business's content — your FAQs, policies, and product information — rather than on scraped third-party data of unclear provenance. If you're choosing an AI tool for your business, "where did the training data come from, and can you prove it?" is now a fair question to ask any vendor.

Book a free consultation with GDPRGard →

Sources

Share this article

Also read:

⚠️ This article is for informational purposes only and does not constitute legal advice. For complex compliance situations, consult a qualified data protection professional. GDPR and EU AI Act requirements are subject to ongoing regulatory guidance — verify current obligations with your legal adviser.