On July 8, 2026, the European Data Protection Board (“EDPB”), the body that coordinates the EU’s national data protection authorities, published its first draft of Guidelines 03/2026 on web scraping in the context of generative AI (the “Guidelines”). The Guidelines address practical compliance challenges for companies that develop AI models or systems and scrape personal data from internet sources to train generative AI, as well as companies that rely on third parties to carry out that scraping. They also matter for deployers and enterprise customers of downstream AI models and systems, including businesses using AI in EU-facing products, services, or operations. The EDPB has previously emphasized that deployers can face regulatory enforcement action if they fail to assess whether third-party AI models were developed through unlawful processing of personal data.
The Guidelines remain subject to public consultation before the final version is published, and comments can be submitted until October 30, 2026. Businesses with EU operations or exposure should consider whether the draft raises operational issues that merit feedback during the consultation process.
For international businesses, key takeaways from the Guidelines include:
- Controllers may be able to rely on an exemption from the requirement to directly inform individuals whose personal data is scraped, but only where stringent conditions are met. The EDPB underscores that controllers – organizations that determine why and how personal data is processed – must make certain information publicly available, typically through an online privacy notice. For international businesses, the practical point is that a generic AI privacy notice is unlikely to be sufficient if it does not explain the characteristics of crawlers used to scrape personal data and provide a “complete” list of sources of personal data to the greatest extent possible. Where that list cannot include every source, the notice should identify the types of sources omitted and explain why.
- Data minimization requirements apply even where training an AI model or system requires large amounts of personal data. In GDPR terms, data minimization means limiting personal data to what is necessary for the stated purpose; scale alone does not remove that obligation. Measures that controllers should implement include: (i) defining precise collection criteria; (ii) excluding categories of websites that are likely to contain personal data about vulnerable individuals or information that may be sensitive or private; and (iii) excluding collection from websites that clearly object to web scraping (e.g., through the use of robots.txt and ai.txt files, or CAPTCHAs).
- Legitimate interests will often be the most viable legal basis for web scraping, but it is not a shortcut and requires a documented legitimate interest assessment (“LIA”). Under the GDPR, the legitimate interests legal basis requires the controller to identify a legitimate purpose, show that the processing is necessary for that purpose, and balance the processing against the rights and interests of affected individuals. As with previous regulatory guidance in this space, the EDPB provides additional considerations that controllers should take into account when conducting LIAs for web scraping, and lists mitigating measures that controllers could implement to limit the impact of web scraping on individuals.
One further key topic covered by the Guidelines is the processing of special category personal data (“SCPD”), a GDPR concept covering more sensitive types of personal data (e.g., health-related information). The EDPB recalls that processing SCPD is prohibited under the EU GDPR unless an Article 9(2) condition is satisfied – even where the processing is incidental or residual, rather than intended by the controller. For global AI developers and deployers, this is an important point: accidental collection or output of SCPD can still create regulatory risk. The EDPB, drawing on the Court of Justice of the European Union’s ruling in GC & Others (C-136/17), indicates that a controller may be able to justify processing SCPD where the processing activity has relevant similarities with the processing activity of a search engine, and where the controller implements and regularly verifies measures within the framework of its “responsibilities, powers, and capabilities” to prevent the “dissemination” of SCPD. The EDPB sets out a list of such measures in the Guidelines, including, for example: (i) defining precise criteria and applying filters to prevent collection of SCPD; (ii) deleting SCPD from datasets immediately after collection or as soon as it is identified; and (iii) preventing extraction of SCPD from the model (e.g., through output filters).
After the AI model is developed, the controller using the model as part of an AI system should implement processes to continuously monitor the system’s outputs and take measures to prevent the AI system from generating SCPD (e.g., through updated or reinforced output filters, or by restricting prompts). In practice, GDPR compliance should extend beyond dataset creation or model training, and must also cover deployment, output testing, application of filters, and prompt governance.
The A&B Privacy, Cyber & Data Strategy Team regularly advises companies developing, providing and deploying AI on the applicability of EU and UK privacy and data protection requirements to their AI-powered products and services. We would be pleased to assist your organization in assessing its obligations and implementing practical measures to enhance compliance.