{"id":48181,"topic":"ai","source":"IAPP","title":"EDPB discusses focuses in draft anonymization, AI web scraping guidelines - IAPP","url":"https://iapp.org/news/a/edpb-discusses-focuses-in-draft-anonymization-ai-web-scraping-guidelines","url_hash":"806a87ba8655c1cbccd37049371b8b08ba4c4dd7","author":"","summary":"<a href=\"https://news.google.com/rss/articles/CBMinAFBVV95cUxNN1B0YnVhVFRpSXFCTkpBUGRFVzNyR2dRVUJ3alN5dl9oa0hRV3RiT2NYOWRHMlY0YnRXUUhrOXR5a3N6YlF0Sk0wVUZ4NHZwSUdNVmFJdi1vMlN3R3NxOEd5TDRjQlozbENCVEdGX294V3Z2eVVvc0xhdkJvSnZicVZLN2dMLU9tRnlHMHUydTI2YjY4Qkd0bGdDWmc?oc=5\" target=\"_blank\">EDPB discusses focuses in draft anonymization, AI web scraping guidelines</a>&nbsp;&nbsp;<font color=\"#6f6f6f\">IAPP</font>","content":"Contributors:\nLexie White\nStaff Writer\nIAPP\nThe European Data Protection Board's recent draft anonymization and artificial intelligence web scraping guidelines aim to clarify what is considered identifiable data under the EU General Data Protection Regulation while addressing the impacts of modern technology.\nThe guidelines, which remain under public consultation through 30 Oct., came as a follow-up to a joint opinion from the EDPB and the European Data Protection Supervisor on GDPR reforms proposed under the European Commission's Digital Omnibus. The proposal raises the potential for fundamental alterations to the landmark data protection regulation, with notable changes to the definition of personal data under consideration.\nDuring a LinkedIn Live with IAPP Research and Insights Director Joe Jones, EDPB Secretariat Deputy Head Gwendal Le Grand noted that while the EDPB supported the package, \"We advised quite strongly against any change to the definition of personal data.\" He stressed the necessity for the guidelines while touting the EDPB's ability to provide timely resources to help companies navigate the compliance landscape before and after the finalization of the omnibus, which is still being negotiated between EU institutions.\nAnonymization\nAnonymization was last addressed by the EDPB through guidance issued in 2014 when the board was still functioning as the Article 29 Working Party. The guidance predated the adoption of the GDPR, leaving more than a decade of developments around data use cases and privacy-enhancing technologies to consider.\nReidentification within a specific context was a key focus of the new draft guide, according to Le Grand.\n\"Rather than asking if an individual is identified or identifiable in an absolute sense, the question is rather about the likelihood that the individual will be identified or identifiable by some entity,\" Le Grand said. \"This may vary from one entity to another, and therefore anonymity has to be assessed from each relevant entity's perspective, meaning that any party for whom the data is intended to be anonymous.\"\nTo help organizations conduct assessments for anonymization, the EDPB outlined both contextual and simplified approaches for reviewing the legal standard for anonymized data.\nLe Grand indicated the simplified approach offers a means to \"voluntarily shift the risk from false positives to false negatives, where the anonymization controller treats anonymous data as personal data because they have, in effect, overestimated the likelihood that certain means will be used.\"\nThough simplified efforts may lead organizations to unnecessarily consider information to be personal data, Le Grand said, \"it can provide greater confidence, and it can be complemented with the contextualized approach to refine the finding.\"\nAs companies expand their abilities to process data, Le Grand said organizations should continue to reassess their compliance with the guidelines and the GDPR's obligations for processing sensitive data during their anonymization processes.\nAI web scraping\nThe web scraping guidance runs complementary to the EDPB's parallel work with the European AI Office on joint guidelines detailing compliance with the GDPR and the EU AI Act.\nThe draft guidelines state organizations should only collect data they consider necessary for AI training while stressing data controllers must comply with the GDPR's data minimization and purpose limitation obligations when collecting and storing consumers' personal data. Organizations must also provide transparency to individuals when required under the GDPR, including when personal data is collected from publicly available sources.\n\"Regardless of the safeguards you put in place, it's quite difficult to have 100% assurance that you're not collecting special categories of data,\" Le Grand said.\nThe CJEU's Case C-136/17 judgment ruled search engines may unintentionally process sensitive data. Le Grand said when using web scraping for AI, the same reasoning applies.\n\"The incidental and residual collection of sensitive data for AI training is not unlawful if the controller implements certain measures to prevent the dissemination of data,\" Le Grand said.\nTo avoid questionable personal data processing altogether, the EDPB recommended companies consider using synthetic data to train AI models. According to Le Grand, organizations could also \"apply syntax-based filtering, replace some or all of the real data with synthetic data if it's possible, or anonymization to anonymize.\"\n\"We've made a lot of efforts to engage more proactively and be more transparent about our stakeholder engagement and about how we use the input that we receive through the stakeholder consultations,\" he said. \"I can assure you that all the input you send us is considered, everything is analyzed and taken into account.\"","image_url":"https://images.contentstack.io/v3/assets/bltd4dd5b2d705252bc/blt349c6870a76cc3b2/6a68aae4b469f6daf23cd41b/silhouette-head-anonymization-data-072826.jpg","lang":"en","published_at":"2026-07-28T15:29:04+00:00","fetched_at":"2026-07-28T16:15:05+00:00","status":"read","starred":0,"extract_state":"ok","summary_auto":"Contributors:\nLexie White\nStaff Writer\nIAPP\nThe European Data Protection Board's recent draft anonymization and artificial intelligence web scraping guidelines aim to clarify what is considered identifiable data under the EU General Data Protection Regulation while addressing the impacts of modern technology. The guidelines, which remain under public consultation through 30 Oct., came as a follow-up to a joint opinion from the EDPB and the European Data Protection Supervisor on GDPR reforms proposed under the European Commission's Digital Omnibus.","cluster_id":null,"extract_retries":0,"extract_error":null,"contract_version":"news_item.v1","format_contract_version":"news_item_formats.v1","dedup_url":"https://iapp.org/news/a/edpb-discusses-focuses-in-draft-anonymization-ai-web-scraping-guidelines","quality_profile":{"profile_version":"extraction_quality.v2","bucket":"high","confidence":0.9,"failure_kind":"none","retryable":false,"retry_after_attempts":0,"reason":"High confidence: full text extraction produced 4859 characters.","operator_guidance":{"severity":"ok","recommended_action":"trust_full_text","next_step":"Use the extracted full text as the primary article source.","operator_label":"Ready","can_retry":false,"can_use_summary":false,"diagnostics_required":false},"content_depth":{"contract_version":"content_depth.v1","category":"full_text","label":"Full text","has_full_text":true,"has_summary":true,"content_length":4859,"summary_length":553,"usable_text_length":4859,"source_field":"content"},"legacy_collapsed":false,"signals":{"extract_state":"ok","extract_error":null,"extract_retries":0,"content_length":4859,"summary_length":553}},"news_item":{"id":48181,"canonical_url":"https://iapp.org/news/a/edpb-discusses-focuses-in-draft-anonymization-ai-web-scraping-guidelines","source_url":"https://iapp.org/news/a/edpb-discusses-focuses-in-draft-anonymization-ai-web-scraping-guidelines","title":"EDPB discusses focuses in draft anonymization, AI web scraping guidelines - IAPP","source_name":"IAPP","author":null,"published_at":"2026-07-28T15:29:04+00:00","locale":"en","topic":"ai","tags":[],"rss_summary":"<a href=\"https://news.google.com/rss/articles/CBMinAFBVV95cUxNN1B0YnVhVFRpSXFCTkpBUGRFVzNyR2dRVUJ3alN5dl9oa0hRV3RiT2NYOWRHMlY0YnRXUUhrOXR5a3N6YlF0Sk0wVUZ4NHZwSUdNVmFJdi1vMlN3R3NxOEd5TDRjQlozbENCVEdGX294V3Z2eVVvc0xhdkJvSnZicVZLN2dMLU9tRnlHMHUydTI2YjY4Qkd0bGdDWmc?oc=5\" target=\"_blank\">EDPB discusses focuses in draft anonymization, AI web scraping guidelines</a>&nbsp;&nbsp;<font color=\"#6f6f6f\">IAPP</font>","full_text":"Contributors:\nLexie White\nStaff Writer\nIAPP\nThe European Data Protection Board's recent draft anonymization and artificial intelligence web scraping guidelines aim to clarify what is considered identifiable data under the EU General Data Protection Regulation while addressing the impacts of modern technology.\nThe guidelines, which remain under public consultation through 30 Oct., came as a follow-up to a joint opinion from the EDPB and the European Data Protection Supervisor on GDPR reforms proposed under the European Commission's Digital Omnibus. The proposal raises the potential for fundamental alterations to the landmark data protection regulation, with notable changes to the definition of personal data under consideration.\nDuring a LinkedIn Live with IAPP Research and Insights Director Joe Jones, EDPB Secretariat Deputy Head Gwendal Le Grand noted that while the EDPB supported the package, \"We advised quite strongly against any change to the definition of personal data.\" He stressed the necessity for the guidelines while touting the EDPB's ability to provide timely resources to help companies navigate the compliance landscape before and after the finalization of the omnibus, which is still being negotiated between EU institutions.\nAnonymization\nAnonymization was last addressed by the EDPB through guidance issued in 2014 when the board was still functioning as the Article 29 Working Party. The guidance predated the adoption of the GDPR, leaving more than a decade of developments around data use cases and privacy-enhancing technologies to consider.\nReidentification within a specific context was a key focus of the new draft guide, according to Le Grand.\n\"Rather than asking if an individual is identified or identifiable in an absolute sense, the question is rather about the likelihood that the individual will be identified or identifiable by some entity,\" Le Grand said. \"This may vary from one entity to another, and therefore anonymity has to be assessed from each relevant entity's perspective, meaning that any party for whom the data is intended to be anonymous.\"\nTo help organizations conduct assessments for anonymization, the EDPB outlined both contextual and simplified approaches for reviewing the legal standard for anonymized data.\nLe Grand indicated the simplified approach offers a means to \"voluntarily shift the risk from false positives to false negatives, where the anonymization controller treats anonymous data as personal data because they have, in effect, overestimated the likelihood that certain means will be used.\"\nThough simplified efforts may lead organizations to unnecessarily consider information to be personal data, Le Grand said, \"it can provide greater confidence, and it can be complemented with the contextualized approach to refine the finding.\"\nAs companies expand their abilities to process data, Le Grand said organizations should continue to reassess their compliance with the guidelines and the GDPR's obligations for processing sensitive data during their anonymization processes.\nAI web scraping\nThe web scraping guidance runs complementary to the EDPB's parallel work with the European AI Office on joint guidelines detailing compliance with the GDPR and the EU AI Act.\nThe draft guidelines state organizations should only collect data they consider necessary for AI training while stressing data controllers must comply with the GDPR's data minimization and purpose limitation obligations when collecting and storing consumers' personal data. Organizations must also provide transparency to individuals when required under the GDPR, including when personal data is collected from publicly available sources.\n\"Regardless of the safeguards you put in place, it's quite difficult to have 100% assurance that you're not collecting special categories of data,\" Le Grand said.\nThe CJEU's Case C-136/17 judgment ruled search engines may unintentionally process sensitive data. Le Grand said when using web scraping for AI, the same reasoning applies.\n\"The incidental and residual collection of sensitive data for AI training is not unlawful if the controller implements certain measures to prevent the dissemination of data,\" Le Grand said.\nTo avoid questionable personal data processing altogether, the EDPB recommended companies consider using synthetic data to train AI models. According to Le Grand, organizations could also \"apply syntax-based filtering, replace some or all of the real data with synthetic data if it's possible, or anonymization to anonymize.\"\n\"We've made a lot of efforts to engage more proactively and be more transparent about our stakeholder engagement and about how we use the input that we receive through the stakeholder consultations,\" he said. \"I can assure you that all the input you send us is considered, everything is analyzed and taken into account.\"","excerpt":"Contributors:\nLexie White\nStaff Writer\nIAPP\nThe European Data Protection Board's recent draft anonymization and artificial intelligence web scraping guidelines aim to clarify what is considered identifiable data under the EU General Data Protection Regulation while addressing the impacts of modern technology. The guidelines, which remain under public consultation through 30 Oct., came as a follow-up to a joint opinion from the EDPB and the European Data Protection Supervisor on GDPR reforms proposed under the European Commission's Digital Omnibus.","extraction":{"state":"ok","confidence":0.9,"error":null,"explanation":"High confidence: full text extraction produced 4859 characters.","diagnostics_url":"/api/diagnose?url=https%3A//iapp.org/news/a/edpb-discusses-focuses-in-draft-anonymization-ai-web-scraping-guidelines","quality_profile":{"profile_version":"extraction_quality.v2","bucket":"high","confidence":0.9,"failure_kind":"none","retryable":false,"retry_after_attempts":0,"reason":"High confidence: full text extraction produced 4859 characters.","operator_guidance":{"severity":"ok","recommended_action":"trust_full_text","next_step":"Use the extracted full text as the primary article source.","operator_label":"Ready","can_retry":false,"can_use_summary":false,"diagnostics_required":false},"content_depth":{"contract_version":"content_depth.v1","category":"full_text","label":"Full text","has_full_text":true,"has_summary":true,"content_length":4859,"summary_length":553,"usable_text_length":4859,"source_field":"content"},"legacy_collapsed":false,"signals":{"extract_state":"ok","extract_error":null,"extract_retries":0,"content_length":4859,"summary_length":553}}},"display_formats":["compact","card","full","digest_section","json"]},"daily_stack_record":{"title":"EDPB discusses focuses in draft anonymization, AI web scraping guidelines - IAPP","url":"https://iapp.org/news/a/edpb-discusses-focuses-in-draft-anonymization-ai-web-scraping-guidelines","summary":"Contributors:\nLexie White\nStaff Writer\nIAPP\nThe European Data Protection Board's recent draft anonymization and artificial intelligence web scraping guidelines aim to clarify what is considered identifiable data under the EU General Data Protection Regulation while addressing the impacts of modern technology. The guidelines, which remain under public consultation through 30 Oct., came as a follow-up to a joint opinion from the EDPB and the European Data Protection Supervisor on GDPR reforms proposed under the European Commission's Digital Omnibus.","source":"IAPP","date":"2026-07-28T15:29:04+00:00","content":"Contributors:\nLexie White\nStaff Writer\nIAPP\nThe European Data Protection Board's recent draft anonymization and artificial intelligence web scraping guidelines aim to clarify what is considered identifiable data under the EU General Data Protection Regulation while addressing the impacts of modern technology.\nThe guidelines, which remain under public consultation through 30 Oct., came as a follow-up to a joint opinion from the EDPB and the European Data Protection Supervisor on GDPR reforms proposed under the European Commission's Digital Omnibus. The proposal raises the potential for fundamental alterations to the landmark data protection regulation, with notable changes to the definition of personal data under consideration.\nDuring a LinkedIn Live with IAPP Research and Insights Director Joe Jones, EDPB Secretariat Deputy Head Gwendal Le Grand noted that while the EDPB supported the package, \"We advised quite strongly against any change to the definition of personal data.\" He stressed the necessity for the guidelines while touting the EDPB's ability to provide timely resources to help companies navigate the compliance landscape before and after the finalization of the omnibus, which is still being negotiated between EU institutions.\nAnonymization\nAnonymization was last addressed by the EDPB through guidance issued in 2014 when the board was still functioning as the Article 29 Working Party. The guidance predated the adoption of the GDPR, leaving more than a decade of developments around data use cases and privacy-enhancing technologies to consider.\nReidentification within a specific context was a key focus of the new draft guide, according to Le Grand.\n\"Rather than asking if an individual is identified or identifiable in an absolute sense, the question is rather about the likelihood that the individual will be identified or identifiable by some entity,\" Le Grand said. \"This may vary from one entity to another, and therefore anonymity has to be assessed from each relevant entity's perspective, meaning that any party for whom the data is intended to be anonymous.\"\nTo help organizations conduct assessments for anonymization, the EDPB outlined both contextual and simplified approaches for reviewing the legal standard for anonymized data.\nLe Grand indicated the simplified approach offers a means to \"voluntarily shift the risk from false positives to false negatives, where the anonymization controller treats anonymous data as personal data because they have, in effect, overestimated the likelihood that certain means will be used.\"\nThough simplified efforts may lead organizations to unnecessarily consider information to be personal data, Le Grand said, \"it can provide greater confidence, and it can be complemented with the contextualized approach to refine the finding.\"\nAs companies expand their abilities to process data, Le Grand said organizations should continue to reassess their compliance with the guidelines and the GDPR's obligations for processing sensitive data during their anonymization processes.\nAI web scraping\nThe web scraping guidance runs complementary to the EDPB's parallel work with the European AI Office on joint guidelines detailing compliance with the GDPR and the EU AI Act.\nThe draft guidelines state organizations should only collect data they consider necessary for AI training while stressing data controllers must comply with the GDPR's data minimization and purpose limitation obligations when collecting and storing consumers' personal data. Organizations must also provide transparency to individuals when required under the GDPR, including when personal data is collected from publicly available sources.\n\"Regardless of the safeguards you put in place, it's quite difficult to have 100% assurance that you're not collecting special categories of data,\" Le Grand said.\nThe CJEU's Case C-136/17 judgment ruled search engines may unintentionally process sensitive data. Le Grand said when using web scraping for AI, the same reasoning applies.\n\"The incidental and residual collection of sensitive data for AI training is not unlawful if the controller implements certain measures to prevent the dissemination of data,\" Le Grand said.\nTo avoid questionable personal data processing altogether, the EDPB recommended companies consider using synthetic data to train AI models. According to Le Grand, organizations could also \"apply syntax-based filtering, replace some or all of the real data with synthetic data if it's possible, or anonymization to anonymize.\"\n\"We've made a lot of efforts to engage more proactively and be more transparent about our stakeholder engagement and about how we use the input that we receive through the stakeholder consultations,\" he said. \"I can assure you that all the input you send us is considered, everything is analyzed and taken into account.\"","confidence":0.9,"diagnostics_url":"/api/diagnose?url=https%3A//iapp.org/news/a/edpb-discusses-focuses-in-draft-anonymization-ai-web-scraping-guidelines","quality_bucket":"high","failure_kind":"none","retryable":false,"quality_reason":"High confidence: full text extraction produced 4859 characters.","quality_profile":{"profile_version":"extraction_quality.v2","bucket":"high","confidence":0.9,"failure_kind":"none","retryable":false,"retry_after_attempts":0,"reason":"High confidence: full text extraction produced 4859 characters.","operator_guidance":{"severity":"ok","recommended_action":"trust_full_text","next_step":"Use the extracted full text as the primary article source.","operator_label":"Ready","can_retry":false,"can_use_summary":false,"diagnostics_required":false},"content_depth":{"contract_version":"content_depth.v1","category":"full_text","label":"Full text","has_full_text":true,"has_summary":true,"content_length":4859,"summary_length":553,"usable_text_length":4859,"source_field":"content"},"legacy_collapsed":false,"signals":{"extract_state":"ok","extract_error":null,"extract_retries":0,"content_length":4859,"summary_length":553}},"tags":[]},"fallback_formats":["markdown","json","html"],"actions":{"read":"/item/48181","export_markdown":"/api/items/48181/export?format=markdown","export_json":"/api/items/48181/export?format=json","diagnose":"/api/diagnose?url=https%3A//iapp.org/news/a/edpb-discusses-focuses-in-draft-anonymization-ai-web-scraping-guidelines"},"formats":{"full":{"id":48181,"title":"EDPB discusses focuses in draft anonymization, AI web scraping guidelines - IAPP","url":"https://iapp.org/news/a/edpb-discusses-focuses-in-draft-anonymization-ai-web-scraping-guidelines","source":"IAPP","author":null,"published_at":"2026-07-28T15:29:04+00:00","locale":"en","topic":"ai","tags":[],"excerpt":"Contributors:\nLexie White\nStaff Writer\nIAPP\nThe European Data Protection Board's recent draft anonymization and artificial intelligence web scraping guidelines aim to clarify what is considered identifiable data under the EU General Data Protection Regulation while addressing the impacts of modern technology. The guidelines, which remain under public consultation through 30 Oct., came as a follow-up to a joint opinion from the EDPB and the European Data Protection Supervisor on GDPR reforms proposed under the European Commission's Digital Omnibus.","full_text":"Contributors:\nLexie White\nStaff Writer\nIAPP\nThe European Data Protection Board's recent draft anonymization and artificial intelligence web scraping guidelines aim to clarify what is considered identifiable data under the EU General Data Protection Regulation while addressing the impacts of modern technology.\nThe guidelines, which remain under public consultation through 30 Oct., came as a follow-up to a joint opinion from the EDPB and the European Data Protection Supervisor on GDPR reforms proposed under the European Commission's Digital Omnibus. The proposal raises the potential for fundamental alterations to the landmark data protection regulation, with notable changes to the definition of personal data under consideration.\nDuring a LinkedIn Live with IAPP Research and Insights Director Joe Jones, EDPB Secretariat Deputy Head Gwendal Le Grand noted that while the EDPB supported the package, \"We advised quite strongly against any change to the definition of personal data.\" He stressed the necessity for the guidelines while touting the EDPB's ability to provide timely resources to help companies navigate the compliance landscape before and after the finalization of the omnibus, which is still being negotiated between EU institutions.\nAnonymization\nAnonymization was last addressed by the EDPB through guidance issued in 2014 when the board was still functioning as the Article 29 Working Party. The guidance predated the adoption of the GDPR, leaving more than a decade of developments around data use cases and privacy-enhancing technologies to consider.\nReidentification within a specific context was a key focus of the new draft guide, according to Le Grand.\n\"Rather than asking if an individual is identified or identifiable in an absolute sense, the question is rather about the likelihood that the individual will be identified or identifiable by some entity,\" Le Grand said. \"This may vary from one entity to another, and therefore anonymity has to be assessed from each relevant entity's perspective, meaning that any party for whom the data is intended to be anonymous.\"\nTo help organizations conduct assessments for anonymization, the EDPB outlined both contextual and simplified approaches for reviewing the legal standard for anonymized data.\nLe Grand indicated the simplified approach offers a means to \"voluntarily shift the risk from false positives to false negatives, where the anonymization controller treats anonymous data as personal data because they have, in effect, overestimated the likelihood that certain means will be used.\"\nThough simplified efforts may lead organizations to unnecessarily consider information to be personal data, Le Grand said, \"it can provide greater confidence, and it can be complemented with the contextualized approach to refine the finding.\"\nAs companies expand their abilities to process data, Le Grand said organizations should continue to reassess their compliance with the guidelines and the GDPR's obligations for processing sensitive data during their anonymization processes.\nAI web scraping\nThe web scraping guidance runs complementary to the EDPB's parallel work with the European AI Office on joint guidelines detailing compliance with the GDPR and the EU AI Act.\nThe draft guidelines state organizations should only collect data they consider necessary for AI training while stressing data controllers must comply with the GDPR's data minimization and purpose limitation obligations when collecting and storing consumers' personal data. Organizations must also provide transparency to individuals when required under the GDPR, including when personal data is collected from publicly available sources.\n\"Regardless of the safeguards you put in place, it's quite difficult to have 100% assurance that you're not collecting special categories of data,\" Le Grand said.\nThe CJEU's Case C-136/17 judgment ruled search engines may unintentionally process sensitive data. Le Grand said when using web scraping for AI, the same reasoning applies.\n\"The incidental and residual collection of sensitive data for AI training is not unlawful if the controller implements certain measures to prevent the dissemination of data,\" Le Grand said.\nTo avoid questionable personal data processing altogether, the EDPB recommended companies consider using synthetic data to train AI models. According to Le Grand, organizations could also \"apply syntax-based filtering, replace some or all of the real data with synthetic data if it's possible, or anonymization to anonymize.\"\n\"We've made a lot of efforts to engage more proactively and be more transparent about our stakeholder engagement and about how we use the input that we receive through the stakeholder consultations,\" he said. \"I can assure you that all the input you send us is considered, everything is analyzed and taken into account.\"","reading_time_min":4,"extraction":{"state":"ok","confidence":0.9,"error":null,"explanation":"High confidence: full text extraction produced 4859 characters.","diagnostics_url":"/api/diagnose?url=https%3A//iapp.org/news/a/edpb-discusses-focuses-in-draft-anonymization-ai-web-scraping-guidelines","quality_profile":{"profile_version":"extraction_quality.v2","bucket":"high","confidence":0.9,"failure_kind":"none","retryable":false,"retry_after_attempts":0,"reason":"High confidence: full text extraction produced 4859 characters.","operator_guidance":{"severity":"ok","recommended_action":"trust_full_text","next_step":"Use the extracted full text as the primary article source.","operator_label":"Ready","can_retry":false,"can_use_summary":false,"diagnostics_required":false},"content_depth":{"contract_version":"content_depth.v1","category":"full_text","label":"Full text","has_full_text":true,"has_summary":true,"content_length":4859,"summary_length":553,"usable_text_length":4859,"source_field":"content"},"legacy_collapsed":false,"signals":{"extract_state":"ok","extract_error":null,"extract_retries":0,"content_length":4859,"summary_length":553}}},"quality_profile":{"profile_version":"extraction_quality.v2","bucket":"high","confidence":0.9,"failure_kind":"none","retryable":false,"retry_after_attempts":0,"reason":"High confidence: full text extraction produced 4859 characters.","operator_guidance":{"severity":"ok","recommended_action":"trust_full_text","next_step":"Use the extracted full text as the primary article source.","operator_label":"Ready","can_retry":false,"can_use_summary":false,"diagnostics_required":false},"content_depth":{"contract_version":"content_depth.v1","category":"full_text","label":"Full text","has_full_text":true,"has_summary":true,"content_length":4859,"summary_length":553,"usable_text_length":4859,"source_field":"content"},"legacy_collapsed":false,"signals":{"extract_state":"ok","extract_error":null,"extract_retries":0,"content_length":4859,"summary_length":553}},"actions":{"read":"/item/48181","export_markdown":"/api/items/48181/export?format=markdown","export_json":"/api/items/48181/export?format=json","diagnose":"/api/diagnose?url=https%3A//iapp.org/news/a/edpb-discusses-focuses-in-draft-anonymization-ai-web-scraping-guidelines"}},"digest":{"id":48181,"title":"EDPB discusses focuses in draft anonymization, AI web scraping guidelines - IAPP","url":"https://iapp.org/news/a/edpb-discusses-focuses-in-draft-anonymization-ai-web-scraping-guidelines","source":"IAPP","topic":"ai","published_at":"2026-07-28T15:29:04+00:00","excerpt":"Contributors: Lexie White Staff Writer IAPP The European Data Protection Board's recent draft anonymization and artificial intelligence web scraping guidelines aim to clarify what is considered identifiable data under the EU General Data Protection Regulation while addressing…","quality_bucket":"high","quality_reason":"High confidence: full text extraction produced 4859 characters.","reading_time_min":4,"cluster_id":null},"card":{"display_title":"EDPB discusses focuses in draft anonymization, AI web scraping guidelines - IAPP","subtitle":"IAPP · 2026-07-28","summary":"Contributors: Lexie White Staff Writer IAPP The European Data Protection Board's recent draft anonymization and artificial intelligence web scraping guidelines aim to clarify what is considered identifiable data under…","badges":["quality:high"],"links":{"read":"/item/48181","original":"https://iapp.org/news/a/edpb-discusses-focuses-in-draft-anonymization-ai-web-scraping-guidelines","diagnose":"/api/diagnose?url=https%3A//iapp.org/news/a/edpb-discusses-focuses-in-draft-anonymization-ai-web-scraping-guidelines"},"quality_warning":null},"export":{"title":"EDPB discusses focuses in draft anonymization, AI web scraping guidelines - IAPP","url":"https://iapp.org/news/a/edpb-discusses-focuses-in-draft-anonymization-ai-web-scraping-guidelines","summary":"Contributors:\nLexie White\nStaff Writer\nIAPP\nThe European Data Protection Board's recent draft anonymization and artificial intelligence web scraping guidelines aim to clarify what is considered identifiable data under the EU General Data Protection Regulation while addressing the impacts of modern technology. The guidelines, which remain under public consultation through 30 Oct., came as a follow-up to a joint opinion from the EDPB and the European Data Protection Supervisor on GDPR reforms proposed under the European Commission's Digital Omnibus.","source":"IAPP","date":"2026-07-28T15:29:04+00:00","content":"Contributors:\nLexie White\nStaff Writer\nIAPP\nThe European Data Protection Board's recent draft anonymization and artificial intelligence web scraping guidelines aim to clarify what is considered identifiable data under the EU General Data Protection Regulation while addressing the impacts of modern technology.\nThe guidelines, which remain under public consultation through 30 Oct., came as a follow-up to a joint opinion from the EDPB and the European Data Protection Supervisor on GDPR reforms proposed under the European Commission's Digital Omnibus. The proposal raises the potential for fundamental alterations to the landmark data protection regulation, with notable changes to the definition of personal data under consideration.\nDuring a LinkedIn Live with IAPP Research and Insights Director Joe Jones, EDPB Secretariat Deputy Head Gwendal Le Grand noted that while the EDPB supported the package, \"We advised quite strongly against any change to the definition of personal data.\" He stressed the necessity for the guidelines while touting the EDPB's ability to provide timely resources to help companies navigate the compliance landscape before and after the finalization of the omnibus, which is still being negotiated between EU institutions.\nAnonymization\nAnonymization was last addressed by the EDPB through guidance issued in 2014 when the board was still functioning as the Article 29 Working Party. The guidance predated the adoption of the GDPR, leaving more than a decade of developments around data use cases and privacy-enhancing technologies to consider.\nReidentification within a specific context was a key focus of the new draft guide, according to Le Grand.\n\"Rather than asking if an individual is identified or identifiable in an absolute sense, the question is rather about the likelihood that the individual will be identified or identifiable by some entity,\" Le Grand said. \"This may vary from one entity to another, and therefore anonymity has to be assessed from each relevant entity's perspective, meaning that any party for whom the data is intended to be anonymous.\"\nTo help organizations conduct assessments for anonymization, the EDPB outlined both contextual and simplified approaches for reviewing the legal standard for anonymized data.\nLe Grand indicated the simplified approach offers a means to \"voluntarily shift the risk from false positives to false negatives, where the anonymization controller treats anonymous data as personal data because they have, in effect, overestimated the likelihood that certain means will be used.\"\nThough simplified efforts may lead organizations to unnecessarily consider information to be personal data, Le Grand said, \"it can provide greater confidence, and it can be complemented with the contextualized approach to refine the finding.\"\nAs companies expand their abilities to process data, Le Grand said organizations should continue to reassess their compliance with the guidelines and the GDPR's obligations for processing sensitive data during their anonymization processes.\nAI web scraping\nThe web scraping guidance runs complementary to the EDPB's parallel work with the European AI Office on joint guidelines detailing compliance with the GDPR and the EU AI Act.\nThe draft guidelines state organizations should only collect data they consider necessary for AI training while stressing data controllers must comply with the GDPR's data minimization and purpose limitation obligations when collecting and storing consumers' personal data. Organizations must also provide transparency to individuals when required under the GDPR, including when personal data is collected from publicly available sources.\n\"Regardless of the safeguards you put in place, it's quite difficult to have 100% assurance that you're not collecting special categories of data,\" Le Grand said.\nThe CJEU's Case C-136/17 judgment ruled search engines may unintentionally process sensitive data. Le Grand said when using web scraping for AI, the same reasoning applies.\n\"The incidental and residual collection of sensitive data for AI training is not unlawful if the controller implements certain measures to prevent the dissemination of data,\" Le Grand said.\nTo avoid questionable personal data processing altogether, the EDPB recommended companies consider using synthetic data to train AI models. According to Le Grand, organizations could also \"apply syntax-based filtering, replace some or all of the real data with synthetic data if it's possible, or anonymization to anonymize.\"\n\"We've made a lot of efforts to engage more proactively and be more transparent about our stakeholder engagement and about how we use the input that we receive through the stakeholder consultations,\" he said. \"I can assure you that all the input you send us is considered, everything is analyzed and taken into account.\"","confidence":0.9,"diagnostics_url":"/api/diagnose?url=https%3A//iapp.org/news/a/edpb-discusses-focuses-in-draft-anonymization-ai-web-scraping-guidelines","quality_bucket":"high","failure_kind":"none","retryable":false,"quality_reason":"High confidence: full text extraction produced 4859 characters.","quality_profile":{"profile_version":"extraction_quality.v2","bucket":"high","confidence":0.9,"failure_kind":"none","retryable":false,"retry_after_attempts":0,"reason":"High confidence: full text extraction produced 4859 characters.","operator_guidance":{"severity":"ok","recommended_action":"trust_full_text","next_step":"Use the extracted full text as the primary article source.","operator_label":"Ready","can_retry":false,"can_use_summary":false,"diagnostics_required":false},"content_depth":{"contract_version":"content_depth.v1","category":"full_text","label":"Full text","has_full_text":true,"has_summary":true,"content_length":4859,"summary_length":553,"usable_text_length":4859,"source_field":"content"},"legacy_collapsed":false,"signals":{"extract_state":"ok","extract_error":null,"extract_retries":0,"content_length":4859,"summary_length":553}},"tags":[],"format_contract_version":"news_item_formats.v1"}}}