GDPR dan big data: landasan hukum untuk pelatihan AI di Belanda

Sebuah gembok digital bercahaya di atas pola papan sirkuit di kantor berkonsep terbuka.

The GDPR and big data are not incompatible, but they force a choice that many organisations postpone: before you collect or reuse a large dataset, you must be able to name the lawful basis for doing so and show that the processing is necessary for that purpose. In practice this means documenting a ground under Article 6 of the General Data Protection Regulation, testing whether a later use for training an algorithm is compatible with the purpose for which the data was originally collected, and assessing the risk before the first model is built. This article sets out how those rules apply to large-scale analytics and AI training in the Netherlands, and where the Autoriteit Persoonsgegevens draws the line.

What counts as personal data in a big data set

Gambar abstrak yang mewakili persimpangan data, AI, dan kerangka hukum, dengan roda gigi dan sirkuit yang saling terkait dengan palu.

The GDPR applies to information relating to an identified or identifiable natural person. In a big data context that threshold is crossed far more often than organisations expect, because identifiability is assessed by reference to all the means reasonably likely to be used, by the controller or by anyone else, to single someone out. A dataset of device identifiers, location traces, transaction records or clickstreams is therefore almost always personal data, even where no name appears in a single column.

Two concepts are regularly confused, and the difference decides whether the regulation applies at all. Pseudonymised data, where the identifiers have been replaced but the key still exists somewhere, remains personal data and remains fully subject to the GDPR. Anonymised data, where re-identification is no longer reasonably possible for anyone, falls outside it. The bar for genuine anonymisation is high: aggregation, hashing or the removal of direct identifiers usually leaves enough of a fingerprint for a determined party to re-identify individuals, particularly when the dataset is rich and covers a long period.

This matters because the claim that a dataset is anonymous is often the load-bearing assumption in an entire AI programme. If it fails, every downstream step, from the training run to the retention schedule, has been carried out without a lawful basis. The safer approach is to treat the dataset as personal data unless a documented re-identification analysis says otherwise, and to revisit that analysis when the dataset is enriched or combined with another source. The European Data Protection Board has also addressed when an AI model itself can be regarded as anonymous, and its answer is that this cannot be assumed: a model trained on personal data may still contain that data in extractable form and has to be assessed case by case.

Choosing a lawful basis for big data and AI training

Gambar yang menunjukkan kontras mencolok antara kisi-kisi terstruktur seperti cetak biru dan nebula berwarna-warni yang cair, melambangkan konflik antara GDPR dan AI.

Every processing operation needs one of the six grounds in Article 6, and the ground has to be chosen before the processing starts rather than reconstructed afterwards. For large-scale analytics, only three are realistic candidates.

Consent is the cleanest in theory and the most fragile in practice. It has to be freely given, specific, informed and unambiguous, and it must be as easy to withdraw as it was to give. Consent that is bundled into general terms, obtained as a condition of a service that does not need the data, or framed so broadly that the individual cannot see what is being agreed to, will not hold. Where consent is withdrawn, you need a technical answer to the question of what happens to a model that has already been trained.

Necessity for the performance of a contract is narrower than it looks. The test is objective necessity for the service the individual actually asked for, not commercial usefulness. Profiling a customer to improve a recommendation engine is generally not necessary for the delivery of the product they bought, and the Court of Justice has repeatedly rejected attempts to stretch this ground to cover behavioural advertising and personalisation.

That leaves legitimate interests, which is the ground most large-scale processing has to rely on. It requires three steps, documented in that order: identifying a real and lawful interest, showing that the processing is necessary because there is no less intrusive way to achieve it, and balancing the interest against the rights and reasonable expectations of the people concerned. In a reference from a Dutch court concerning the national tennis association, the Court of Justice confirmed in 2024 that a purely commercial interest can qualify as a legitimate interest, but only if the necessity and balancing tests are genuinely satisfied. Commercial value is therefore an admissible interest, not a licence.

Special categories and inferred characteristics

Data revealing racial or ethnic origin, political opinions, religious beliefs, trade union membership, health, sex life or sexual orientation, together with genetic and biometric data used for identification, may not be processed at all unless one of the narrow exceptions in Article 9 applies. Explicit consent is the usual route; the others rarely fit a commercial dataset.

The trap in big data is that special category data does not have to be collected deliberately. Where a dataset allows a protected characteristic to be inferred, for example through purchase patterns, location history or the contents of free-text fields, the stricter regime applies to that processing even though nobody ever asked the question directly. Model features that act as proxies for a protected characteristic should be identified during development, not after a complaint.

Purpose limitation: can you reuse data to train a model

You can, but not automatically. The GDPR requires personal data to be collected for specified, explicit and legitimate purposes and not to be further processed in a way that is incompatible with those purposes. Training an algorithm on data that was gathered for order fulfilment, customer support or fraud prevention is a further processing operation, and it is lawful only if it passes the compatibility test the regulation sets out.

That test is not a formality. It asks about the link between the original purpose and the new one, the context in which the data was collected and what the individual could reasonably expect on that basis, the nature of the data and whether special categories are involved, the possible consequences of the new processing, and the existence of safeguards such as pseudonymisation or encryption. A dataset of medical or financial records will fail the test where a dataset of anonymised operational logs would pass it.

There is one shortcut, and it is narrower than it is usually assumed to be. Further processing for archiving in the public interest, for scientific or historical research or for statistical purposes is treated as compatible, provided the safeguards for research processing are in place. Commercial model development dressed up as research does not qualify; the exception is aimed at genuine research subject to methodological and ethical standards, not at product improvement. Where the compatibility test is not met, you need a fresh legal basis and, in most cases, fresh information to the people concerned.

The practical consequence is that purpose limitation has to be handled at the point of collection. Privacy statements that describe purposes in terms broad enough to be meaningless will not save a later training run, because the test looks at what the individual could reasonably expect, not at what the drafting technically permits. Our guide to writing a privacy policy in the Netherlands sets out how to describe purposes in a way that is both honest and usable.

Minimisation, retention and accuracy when the model wants everything

Data minimisation requires personal data to be adequate, relevant and limited to what is necessary. That principle is genuinely in tension with a development method whose logic is that more data produces a better model, and the tension cannot be resolved by ignoring it. What it can be resolved by is a documented argument: which fields are needed for the stated purpose, what was tested without them, and why the remainder was discarded. A controller who can show that analysis has a defensible position even where the dataset is large. A controller who kept everything because storage is cheap does not.

Storage limitation raises the same question over time. Training data, feature stores, model checkpoints and inference logs all fall within the retention obligation, and each needs its own period tied to a purpose. Inference logs are frequently forgotten and frequently contain more personal data than the training set ever did.

Accuracy is the principle most often overlooked in this context, and it has a legal edge that is easy to miss. Personal data must be accurate and, where necessary, kept up to date, and individuals have a right to rectification. Where a model produces an output about an identifiable person, that output is itself personal data. An inference that someone is a poor credit risk, a likely fraudster or an unsuitable candidate is data about them, it can be inaccurate, and it can be challenged. Building a system in which such outputs cannot be corrected creates a compliance problem that no amount of documentation will fix afterwards.

When a data protection impact assessment is compulsory

Sebuah gedung pemerintahan Belanda yang tampak tegas dengan kaca pembesar di atasnya, melambangkan pengawasan regulasi.

A data protection impact assessment is mandatory where a type of processing is likely to result in a high risk to the rights and freedoms of individuals, and the regulation names three cases in particular: systematic and extensive evaluation of personal aspects based on automated processing, including profiling, on which decisions with legal or similarly significant effects are based; large-scale processing of special categories of data or of criminal conviction data; and systematic monitoring of a publicly accessible area on a large scale. Most serious big data projects fall into the first or the second.

The Autoriteit Persoonsgegevens has also published a list of processing operations for which an assessment is always required in the Netherlands, covering areas such as large-scale profiling, the systematic evaluation of employees, health data platforms and the use of camera systems. That list should be the first document consulted at the start of a project, because it removes the argument about whether an assessment is needed.

Two points about timing are decisive. The assessment has to be carried out before the processing begins, which in an AI project means before the training data is assembled, not before deployment. And where the assessment shows a high residual risk that you cannot mitigate, you must consult the supervisory authority before proceeding. Skipping that consultation is itself a breach, independent of whether the processing turns out to be lawful.

A useful assessment for an algorithmic system goes beyond the standard template. It records which data sources feed the model and on what basis, what the model is allowed to decide on its own and what a human decides, how the output is explained to the person affected, how the training data was tested for bias against protected groups, and what happens if the model is wrong. Those are the questions a regulator asks first.

How the Autoriteit Persoonsgegevens enforces this

Gambar yang menunjukkan perisai digital retak dengan aliran data bocor keluar, mencerminkan pelanggaran data dalam sistem yang digerakkan oleh AI.

The Dutch supervisory authority is the Autoriteit Persoonsgegevens, which enforces the GDPR together with the Dutch implementing act, the Uitvoeringswet AVG. Its powers run from a warning and a reprimand through to an order subject to a penalty payment, a temporary or definitive ban on the processing and an administrative fine. The regulation sets the ceilings: up to ten million euros or two per cent of total worldwide annual turnover for the lower tier, and up to twenty million euros or four per cent for breaches of the basic principles, the lawful bases, the rights of data subjects and the rules on international transfers, whichever amount is higher.

Two enforcement themes are visible in the Dutch practice and both bear directly on big data. The first is transparency: a substantial share of recent action concerns privacy statements that do not explain in intelligible language which data is used, for which purpose and for how long. Vague purpose descriptions are treated as a breach in their own right, not as a drafting flaw. The second is algorithmic oversight. The authority has a dedicated unit for the supervision of algorithms and publishes periodic reports on algorithmic risks, which are worth reading as a statement of what it expects before it becomes the subject of an investigation.

Alongside enforcement sits the breach notification duty. A personal data breach must be reported to the authority without undue delay and, where feasible, within seventy-two hours of becoming aware of it, unless the breach is unlikely to result in a risk to individuals; where the risk to individuals is high, they must be informed as well. In an AI environment the assessment is harder than it used to be, because a compromised training set or feature store affects everyone whose data it contains and the consequences extend to every decision the model has made. Note that this is a separate obligation from the incident reporting duties under the Cyberbeveiligingswet, the Dutch implementation of NIS2, which has applied since 15 August 2026 and imposes its own notification deadlines of twenty-four and seventy-two hours on entities within its scope. Our overview of NIS2 dan Undang-Undang Keamanan Siber Belanda explains how the two regimes sit alongside each other.

The AI Act does not replace the GDPR

The EU Artificial Intelligence Act regulates AI systems as products: it classifies them by risk and imposes obligations on providers and deployers. It does not provide a legal basis for processing personal data, and compliance with it says nothing about compliance with the GDPR. Where an AI system processes personal data, both regimes apply in full and in parallel.

The timetable matters for planning. The prohibitions on unacceptable practices, the obligations for general purpose AI models and the transparency duties for systems that interact with people or generate synthetic content are already in force. The high-risk regime was postponed by the digital omnibus package: the obligations for high-risk systems listed in Annex III now apply from 2 December 2027, and those for systems that are safety components of regulated products under Annex I from 2 August 2028. The separate proposal for an AI Liability Directive has been withdrawn, so liability for damage caused by an AI system continues to be governed by ordinary Dutch rules on contract and tort and by the European product liability regime.

For a data-driven organisation the practical consequence is that the GDPR analysis remains the binding constraint today, while the AI Act determines what documentation, testing and human oversight the same system will need before the end of the decade. Building the two exercises separately duplicates work; building them together does not. Our guides to the UU AI UE dan untuk sistem AI berisiko tinggi set out the classification and the obligations in detail.

Group claims and damages: the civil side of the risk

Regulatory fines are not the only exposure, and in the Netherlands they may not be the largest. The Wet afwikkeling massaschade in collectieve actie (WAMCA) allows a foundation or association that meets strict governance and funding requirements to bring a collective action for damages on behalf of a defined group, with an opt-out regime for people domiciled in the Netherlands. Data-driven processing is an obvious target: a single design decision affects every user in the same way, which is precisely the homogeneity a collective action needs.

The GDPR reinforces this from its own side. It gives anyone who has suffered material or non-material damage as a result of an infringement a right to compensation from the controller or processor, and it allows a not-for-profit body active in the field of data protection to bring proceedings on behalf of data subjects, in some cases without a mandate from them. The Court of Justice has confirmed that a consumer-protection association may act on that basis.

There is a limit, and it is a useful one for defendants. The Court of Justice has held that an infringement of the GDPR does not by itself give rise to a right to compensation: the claimant must show actual damage and a causal link, although no threshold of seriousness applies to non-material damage once it is established. Individual awards in Dutch practice have therefore been modest, but the arithmetic of a class of hundreds of thousands changes the picture entirely. Our article on klaim kolektif jika terjadi kerusakan massal describes how these proceedings run.

Who is the controller when the data comes from everywhere

Big data projects rarely stay inside one organisation. Data is enriched by a broker, hosted by a cloud provider, cleaned by an analytics agency and fed into a model supplied by a vendor, and each of those relationships has to be characterised correctly, because the allocation of roles determines who carries which obligation and who answers to the regulator.

A controller determines the purposes and means of the processing; a processor acts only on the controller's documented instructions. Where two or more parties jointly determine the purposes and means, they are joint controllers and must set out their respective responsibilities in an arrangement, in particular for providing information and for handling requests from individuals, who may in any event exercise their rights against either of them. The label used in the contract is not decisive: what counts is who actually decides why and how the data is processed. A supplier that reserves the right to use your data to improve its own product has, to that extent, become a controller for its own purposes, whatever the agreement calls it.

Two clauses in supplier contracts deserve particular attention in an AI context. The first is the one permitting the provider to use customer content for training or for service improvement; if it is present, you are disclosing personal data to another controller and you need a basis and a notice for that disclosure. The second is the sub-processor clause, because model providers frequently rely on further suppliers and on infrastructure outside the European Economic Area, which brings the transfer rules into play. A processing agreement that covers the required subject matter but does not describe what actually happens to the data is of little use in an investigation.

Finally, remember that the accountability principle puts the burden of proof on the controller. It is not enough to be compliant; you must be able to demonstrate it, with records, assessments and agreements that match the systems as they are actually built.

What to put in place before the next project starts

Compliance for big data is built at the design stage, because the regulation requires data protection by design and by default: appropriate technical and organisational measures have to be integrated into the processing itself, and only the personal data necessary for each specific purpose may be processed by default. Retrofitting is expensive and usually incomplete.

Five things make the difference in practice. Keep a record of processing activities that actually describes the data flows behind each model rather than the departments that own them. Decide and write down the lawful basis for each processing operation, including the further processing involved in training, and keep the legitimate interests assessment with it. Carry out the impact assessment before the data is assembled, and consult the supervisory authority where the residual risk stays high. Put a data processing agreement in place with every supplier that touches the data, including the providers of the model or the platform, and check what they are permitted to do with your data for their own purposes; the requirements are set out in our guide to the perjanjian pemrosesan data. And check where the data goes: transfers outside the European Economic Area need a transfer mechanism under Chapter V of the regulation and a transfer impact assessment, and the position of individual third countries can change.

Above all, make the people who build the models and the people who assess the risk work on the same document. The foundations of the regulation itself are summarised in our overview of the Peraturan Perlindungan Data Umum, and the specific questions raised by putting personal data into an AI system, from automated decision-making to the rights of the people affected, are dealt with in our companion article on GDPR and AI in the Netherlands.

Beberapa pertanyaan umum

Which lawful basis should we use for training an AI model

In most commercial settings the answer is legitimate interests, supported by a written assessment covering the interest, the necessity and the balancing exercise. Consent is preferable where the data is sensitive or the use would surprise the people concerned, but it has to be genuinely free and revocable, which is hard to organise for a training set. Necessity for a contract almost never covers model development, because the training is not what the customer asked for. Whichever ground you pick, record it before the processing starts: choosing a basis retrospectively is treated as no basis at all.

Does the GDPR still apply if we anonymise the data first

Only if the anonymisation genuinely works. Data is anonymous when re-identification is no longer reasonably possible for anyone, taking account of all the means likely to be used and of other datasets that could be combined with it. Removing names, replacing identifiers with hashes or aggregating to small groups usually falls short of that standard, and the result is pseudonymised data, which remains subject to the regulation in full. Where anonymisation is the basis for treating a project as out of scope, that conclusion needs to be tested and documented, and revisited whenever the dataset changes.

Can we rely on the research exception to reuse customer data

Rarely. Further processing for scientific or statistical purposes is treated as compatible with the original purpose, but the exception is aimed at research carried out to recognised methodological standards and subject to safeguards such as pseudonymisation and access restrictions. Product development or model improvement carried out by a commercial party for its own benefit does not become research because it involves data science. If the compatibility test is not met on its own merits, you need a new lawful basis and, in most cases, new information to the individuals concerned.

What happens if someone withdraws consent after a model has been trained

Withdrawal takes effect for the future and does not retrospectively make earlier processing unlawful, but it does mean the data may no longer be used on that basis. Practically, that requires the ability to remove the individual from the training set and from feature stores and logs, and to retrain or otherwise ensure the model no longer reflects that person's data where it can be shown to do so. This is one of the reasons consent is a difficult basis for model training, and it is a question worth answering in the design phase rather than after the first request arrives.

Seterpercayaapakah Olymp Trade? Kesimpulan Law and More dapat membantu

Law and More advises Dutch and international organisations on the data protection side of analytics and AI: selecting and documenting a lawful basis, carrying out legitimate interests assessments and impact assessments, drafting privacy statements and processing agreements that survive scrutiny, responding to the Autoriteit Persoonsgegevens, and defending claims brought by individuals or by representative foundations. If you are planning a data-driven project or a regulator has already been in touch, we are happy to look at the file with you.

Butuh Bantuan Hukum?

Kontak Law & More Untuk panduan ahli mengenai masalah hukum Anda. Tim multibahasa kami siap membantu.

Terkait artikel

Hak cipta timbul secara otomatis pada sebuah foto, pada saat foto tersebut diambil, dengan syarat foto tersebut

Perjanjian layanan TI adalah kontrak di mana penyedia memberikan layanan teknologi kepada

Menggunakan logo tanpa izin adalah melanggar hukum di Belanda jika logo tersebut

Tunjangan nafkah pasangan di Belanda berakhir secara otomatis berdasarkan hukum ketika penerima menikah lagi, memasuki

Chatbot yang diterapkan di Belanda berada di persimpangan tiga badan hukum:

Mempublikasikan gambar seksual seseorang tanpa persetujuan mereka merupakan tindak pidana tersendiri.

Tetaplah mengikuti perkembangan hukum Belanda.

Berlangganan buletin kami untuk mendapatkan wawasan hukum terbaru, pembaruan peraturan, dan saran praktis.