AI and data protection: what companies have to watch out for when using AI-supported systems

Jens Bimberg
AI and data protection: what companies have to watch out for when using AI-supported systems

In a world increasingly shaped by technological progress, artificial intelligence (AI) opens up fascinating possibilities as well as potentially unsettling challenges, especially in the context of processing personal data. The advancing integration of AI algorithms into various areas of our lives promises gains in efficiency, personalised services and innovative approaches to problems. At the same time, though, the growing volume of collected data and the complexity of AI systems raise legitimate questions about data protection, privacy and ethical responsibility.

In this article I take a look at some data protection questions and set out ways in which you can handle data worth protecting safely in AI projects.

Areas of AI use that are not relevant to data protection

There are numerous AI applications that are not directly relevant to data protection. AI is used to gain insights into how stars form, for example, or to optimise telecommunications networks. In robotics, AI comes into play when it comes to navigating safely in unknown environments. A manufacturing company can use AI algorithms to monitor wear on production machinery: based on the data recorded, the AI can predict when certain parts have to be replaced or serviced in order to prevent unforeseen downtime. AI can also be used in art, generating images, music or videos. In our digital publishing software Purple, we use AI to free journalists from unloved tasks such as linking or SEO, so that they can concentrate on their actual journalistic work.

None of these applications process personal data, and to that extent they are not relevant to data protection.

Processing personal data

Unlike the cases of AI use described above, any processing of personal data is subject to strict data protection rules. The GDPR expressly prohibits any processing of personal data unless one of 6 possible conditions applies:

  • The informed consent of the data subject
  • Performance of a contract
  • Compliance with a legal obligation
  • Protection of vital interests
  • Public interest
  • Legitimate interest of the controller

Most providers rely on informed consent. Here the person agrees to their data being processed after being informed about the purpose of the processing "in an intelligible and easily accessible form, using clear and plain language". The data subject can request information about the data processed about them, and its correction, at any time. Beyond that, the person can withdraw consent to data processing at any time once it has been given. That does not affect processing that has already taken place, but it prohibits further use of the data and forces its deletion.

Processing personal data using AI

Given the above, an unsolvable problem arises for the controller if they have used personal data to train an AI, for instance. Because it is practically impossible to change individual pieces of data fed into an AI model after the fact, or to remove them from it again. The controller's only option is to throw the model away and retrain it, at considerable cost in time, leaving out the data to be deleted, provided they still have the remaining training data at all. (That would not be the case, for example, if users communicate with the AI directly.)

Anonymisation

One solution to this problem can be anonymising the data. Anonymous data is by definition not personal data, which means its use is not subject to the provisions of the GDPR or of other data protection laws.

The anonymisation itself is of course processing, which the data subject has to consent to. If they later withdraw their consent, that no longer affects the processing that has already taken place (the anonymisation), as set out above, and the anonymous data can continue to be used without consent.

The anonymisation obviously has to be carried out in such a way that no conclusions about individuals are possible any more (otherwise it would not be anonymisation). If only the height and eye colour of 100,000 people are stored in a data pool, for example, that data pool can safely be regarded as anonymous, since the records no longer allow you to identify a person. If an AI model is trained only on anonymised data, there should be no data protection concerns about using it.

Wondering what potential AI could unlock in your company? Request an AI workshop now.

Other sufficient conditions for data processing

Consent from the data subject is not strictly required for processing personal data, as I showed above. Processing is also permitted where it is in the public interest, for example, or in the legitimate interest of the controller (and where the interest of the data subject does not override it). In these cases personal data can also be processed by AI.

Article 22 of the GDPR must not be overlooked here, however: people may not be subjected to decisions based solely on automated processing. In essence this article anticipates the EU Commission's AI regulation, which is currently still being agreed and which will govern such cases in detail.

If the reason for the processing falls away, the data has to be deleted. That is not a problem if the reason falls away for everyone at the same time, since you would delete the whole AI model anyway. With processing that is set up so that it is only necessary for a certain period for individual people, in order to perform a contract for example, the problem mentioned above comes into play again: nobody can remove one person's data from an AI model. Processing of that kind therefore cannot be implemented with legal certainty.

Publicly available or private models

Various public AI models such as Google Bard are used by millions of people worldwide. All the input data, meaning the user prompts, flows to the operator. The operator uses the input data to develop its models further, and in doing so the data is also read by humans (reviewers), among other things. It is evident that public models of this kind cannot be used in a way that complies with data protection law.

Fortunately, though, it is not difficult to run an AI of your own, whether in your own data centre or at a cloud provider. Here you can draw on pre-trained models, meaning AI models that have already been trained on a large volume of data for particular purposes, and adapt or develop them further for your own purposes.

Conclusion

There are numerous ways in which you can use AI sensibly without touching data protection at all. But AI-supported processing of personal data can also be set up so that it complies with existing data protection laws, as I have shown in this article.

The biggest risks of AI are compliance, data security, controllability and ethical conflicts. How do we deal with them? Proactively. We document the development of our AI systems transparently and in a way that can be traced. We also work closely with you at a technical and a process level in order to spot possible risks early. That puts us in a position to integrate the right safeguards into your workflows and systems. Find out more about our AI services here.

Tech Newsletter

Join our 2,000+ subscribers and receive monthly updates on our latest articles, case studies, webinars, events, and industry news.

Fünf Menschen sitzen an einem Konferenztisch, konzentriert und mit Laptops in einem modernen Büro.