Guide

GDPR for researchers using AI and cloud tools: transcription, surveys and AI assistants

Researchers now routinely send interview recordings to transcription services, run surveys on cloud platforms and paste text into AI assistants. Each of those steps is processing of personal data. This guide explains, with sources, what the GDPR expects from you and your institution, and gives a checklist to run before you choose a tool.

Published 7 October 2026 · Sources checked 7 October 2026

The short answer

Prefer a tool built in the EU? Kahubi, from Avidemic AB in Sweden, covers an AI assistant for research work with EU hosting and, for institutions, only European subprocessors. See how Kahubi handles research data

Who is responsible: controller and processor

The GDPR assigns responsibility by role. The controller decides the purposes and means of processing; a processor acts only on the controller's instructions [4]. In most university research the controller is the university (sometimes jointly with partner institutions), and the researcher acts on its behalf. That has two practical consequences:

The EU AI Act adds a separate layer. Universities and researchers who use AI tools are deployers, and the AI literacy duty in Article 4 has applied since 2 February 2025. The research exclusion in the AI Act covers AI systems developed solely for scientific research, not general tools used to do research [14][15]. Our AI Act guide for universities covers this in detail.

The EDPB's Guidelines 1/2026 on processing personal data for scientific research were adopted on 15 April 2026 as a version for public consultation; the consultation closed on 25 June 2026 and we found no final version when we checked [1][2]. They are the most detailed EU-level guidance available and the basis for this section.

Public interest: Article 6(1)(e)

The EDPB describes public interest as a legal basis "of particular importance for scientific research" [1]. Under Article 6(3) GDPR, it must rest on Union or Member State law to which the controller is subject, but that law does not need to spell out how the research is done. The guidelines give national examples, including the Swedish Higher Education Act and the Finnish Data Protection Act. Public authorities cannot rely on legitimate interests (Article 6(1)(f)) for processing in the performance of their tasks [1].

Consent: Article 6(1)(a)

Consent is possible, and the EDPB accepts "broad consent" for an area of research where purposes are not fully known, and "dynamic consent" project by project, both with extra safeguards [1]. But consent must be freely given. The EDPB says that where there is an imbalance of power, for example research involving pupils, convicted persons or unemployed people, and there is doubt, controllers should not rely on consent [1].

Validemic's analysis A researcher interviewing their own students, or a PhD student surveying colleagues in their own department, is in a position where consent as a legal basis is hard to defend. Many universities therefore use public interest as the legal basis and keep consent as an ethical safeguard.

Ethical consent is not the same as GDPR consent

Ethics rules often require consent to take part in research. The EDPB stresses that this should be distinguished from consent as a GDPR legal basis. Ethical consent can be an additional safeguard under Article 89, even where the legal basis is something else. If you do want GDPR consent, the request must be clearly distinguishable from the request to participate [1].

Special category data: Article 9(2)(j) and explicit consent

Health, ethnicity, political opinions, religion, sexual orientation, trade union membership, genetic and biometric data need an Article 9(2) condition as well. The EDPB explains that Union or Member State law can provide a research derogation under Article 9(2)(j), with suitable and specific safeguards, and that controllers must be able to show that the law applies and that its safeguards are in place. The guidelines list national examples such as section 10 of the Danish Data Protection Act and the Swedish Ethical Review Act. Where no such law applies, explicit consent is the alternative [1].

Interviews often produce special category data even when the topic seems neutral: a participant mentions an illness or a religious practice. Plan for it in the legal basis and the tool choice.

Further processing and retention

The EDPB says further processing for scientific research is presumed compatible with the original purpose, so no compatibility test is needed, but the controller must still check lawfulness. Data may be stored for longer for research purposes, for example for follow-up projects [1].

Article 89 safeguards, pseudonymisation and anonymisation

Research benefits from flexibilities in the GDPR, but Article 89(1) requires appropriate safeguards, in particular data minimisation. The EDPB summarises the order of preference [1]:

  1. Anonymised data first, where the research purpose can be met with it. Truly anonymous data falls outside the GDPR, but the anonymisation process itself is processing that must comply.
  2. Pseudonymised data where anonymisation is not possible, for example where you need to follow people over time.
  3. Directly identifying data only where strictly necessary and proportionate.

Pseudonymisation is defined in Article 4(5) GDPR as processing so that data can no longer be attributed to a specific person without additional information that is kept separately and protected [1]. Pseudonymised data is still personal data. The EDPB also lists other safeguards: independent or ethical oversight, secure processing environments, privacy-enhancing technologies, protective measures when publishing results and confidentiality arrangements [1].

On anonymisation specifically, the EDPB adopted draft Guidelines 02/2026 on 8 July 2026, open for consultation until 30 October 2026. They test anonymisation against three risks: singling out a record, linking records, and inferring information [3].

Data stateExample in qualitative researchGDPR applies?
IdentifyingAudio recording; transcript with namesYes
PseudonymisedTranscript with names replaced by codes, key held separatelyYes
Possibly anonymisedAggregated themes with no quotes that could identify anyoneOnly if re-identification is not reasonably likely

Validemic's analysis Removing names from a rich interview transcript rarely makes it anonymous. Job titles, places, events and turns of phrase can identify someone in a small field. Treat de-identified transcripts as pseudonymised unless your DPO has assessed otherwise.

Interview recordings and transcription

An interview recording is personal data, and a voice can identify a person. Transcription is often the first point where research data leaves the university's own systems. The options, from most to least contained, are:

  1. Transcription inside a secure institutional environment. Some universities run AI services within their own secure research infrastructure. The University of Oslo, for example, announced on 28 September 2026 a large language model inside its TSD secure environment, stating that sensitive data can be processed "without leaving the TSD environment" [11].
  2. An institutionally licensed service with a processor agreement, known data location and contractual terms on AI training and sub-processors.
  3. A consumer or free account with a cloud transcription service, accepted by the researcher personally. This is the option most likely to lack a processor agreement and to allow the vendor wider use of the content.

Questions to answer for any transcription tool:

Prefer a tool built in the EU? Kahubi, from Avidemic AB in Sweden, covers an AI assistant for research work with EU hosting and, for institutions, only European subprocessors. See how Kahubi handles research data

Our fact sheets on Otter.ai, Trint, Rev and Zoom set out what each vendor publicly documents about hosting, agreements and training, plan by plan.

Surveys

Survey platforms are processors too, and the same contract and transfer questions apply. Points specific to surveys:

See our fact sheets on Qualtrics, SurveyMonkey and Google Forms.

Uploading data to AI assistants

Pasting a transcript into a chatbot to summarise it, asking an assistant to code qualitative data, or uploading a dataset for analysis are all processing operations. Before doing any of them with personal data, the same questions apply as for any processor: is there an agreement, where does the data go, and can the vendor reuse it?

Supervisory authorities have started to treat this as a security issue. In its report on data breaches in 2024, published in July 2025, the Dutch Autoriteit Persoonsgegevens said it had received several breach notifications from organisations whose staff entered personal data of patients or customers into AI chatbots. It noted that staff often use chatbots on their own initiative and against agreements with their employer, and that entering personal data in that situation is a (reportable) data breach [9].

University guidance points the same way. The University of Helsinki's guidelines on generative AI in research tell researchers not to use sensitive or confidential data or unpublished results as input, because of the risk of leaks, and to report substantial AI use, including the tool, version, date and purpose [10].

Practical rules that follow:

For specific tools, see ChatGPT, Microsoft Copilot, Gemini, NotebookLM and Elicit.

Consumer versus institutional accounts

The same brand can have very different terms depending on the plan. OpenAI's documentation is a clear example (checked 7 October 2026):

Services for individuals (e.g. ChatGPT personal accounts)Business services (ChatGPT Business, Enterprise, Edu, API)
Use of content for model trainingOpenAI "may use your content to train our models"; users can opt out in data controls [7]Not used to improve models by default [7][8]
Retention controlSet by the user's settingsCustomer controls retention for Enterprise, Healthcare and Edu [8]
Access controlIndividualOrganisation controls access, with SAML SSO [8]

Note one detail from OpenAI's help page: even after opting out on a personal account, giving feedback on a response (for example a thumbs up or down) may allow the whole conversation to be used for training [7].

Validemic's analysis An institutional plan does not by itself make a use lawful. It gives your university the contractual controls it needs as controller. The legal basis, minimisation, transfers and participant information still have to be right. Conversely, a consumer account with training switched off still lacks a processor agreement between your university and the vendor.

Reviewing a vendor right now? Validemic checks the vendor's documents against GDPR and the EU AI Act and cites every finding. Try the demo workspace

Transfers outside the EEA

If a tool transfers personal data outside the EEA, for example to a hosting location or sub-processor abroad, the transfer rules in Chapter V GDPR apply. Ask vendors where data is stored and from where staff and sub-processors can access it. The EDPB's guide for small organisations summarises the tools [6]:

For US vendors, check whether the specific company is certified under the Data Privacy Framework. On 3 September 2025 the EU General Court dismissed an action to annul the framework's adequacy decision (Case T-553/23, Latombe v Commission). The Court's press release notes that an appeal on points of law may be brought before the Court of Justice [13]. Check the current position before relying on the framework for a long project.

Our transfer mechanism tool walks through which tool applies to a given vendor and data flow.

Participant information sheets

The GDPR requires transparent information to participants, and the EDPB says that for long-running research the controller should keep participants informed throughout, including updates if processing changes. The duty also applies where a processor holds the data [1]. For tool use, the information sheet should cover:

Validemic's analysis Example wording, to be adapted with your DPO and ethics committee:

"We will record the interview and use [service], a transcription service contracted by [University], to produce a written transcript. [Service] processes the recording on our behalf, in [location], and may not use it for its own purposes, including training AI models. We delete the recording [timing]. In the transcript we replace your name and other details that could identify you with a code. The key linking codes to names is stored separately at [University]. This means your data is pseudonymised, not anonymous, until we delete the key on [date]."

Illustrative wording drafted by Validemic

Avoid promising things the tool cannot guarantee, such as "your data will never leave the EU", unless the vendor's contract and sub-processor list support it.

Ethics review, DPIAs and data management plans

Ethics review

Research ethics review and the GDPR overlap but are separate. The EDPB treats adherence to ethical standards as one of six indicative factors of scientific research, and names ethical oversight as an Article 89 safeguard [1]. If you change tools after approval, for example switching from manual transcription to an AI service, check whether your ethics committee and DPO need to be told.

DPIA

A DPIA is mandatory where processing is likely to result in high risk. The EDPB lists criteria including sensitive or highly personal data, vulnerable data subjects, large-scale processing, matching datasets, systematic monitoring and innovative technology, and says that in most cases two criteria should lead to a DPIA [5]. A qualitative study on health experiences that uses an AI transcription service may well meet two. The EDPB research guidelines say the choice of data format (anonymised, pseudonymised, identifying) should be made when the research plan is drafted, as part of a risk analysis or, where relevant, a DPIA [1]. Use our DPIA screening tool for a first check.

Data management plans

Some funders require a data management plan (DMP). For example, ERC grantees under Horizon Europe must submit a DMP by the end of month 6 of the project, and the ERC applies the principle "as open as possible, as closed as necessary" [17]. A DMP is a good place to record which tools touch personal data, where, under which agreement, and when data will be pseudonymised, anonymised or deleted. Keep it consistent with the information sheet and the DPIA.

Checklist before you use a tool

Work through these questions before any personal data goes into a transcription service, survey platform or AI assistant. Share the answers with your DPO.

  1. Is it personal data? Voices, transcripts, free text and pseudonymised data all count. If you can genuinely work with anonymised data, do.
  2. Is the tool approved? Is it on your university's approved list for this data class, under an institutional licence rather than a personal account?
  3. Is there a processor agreement? Between your university and the vendor, covering instructions, security, sub-processors and deletion [4].
  4. Does the vendor use your data for its own purposes? Check training and product improvement terms for the exact plan you will use.
  5. Where is the data processed? Including sub-processors and support access. If outside the EEA, which transfer tool applies?
  6. What is the legal basis? Article 6, plus Article 9 for sensitive data. Does your ethics consent form also need to be a GDPR consent?
  7. Which Article 89 safeguards apply? Pseudonymisation timing, access control, separate key storage, secure environment.
  8. Do you need a DPIA? Screen against the EDPB criteria [5].
  9. Do participants know? Does the information sheet describe this tool category, location, retention and pseudonymisation accurately?
  10. Are new AI features switched on? Summaries, sentiment or emotion analysis and speaker identification add processing; some may be prohibited in education or workplace settings under the AI Act [16].
  11. Is it in the DMP? Record the tool, data flow and retention.
  12. Can you delete it? Confirm how to delete recordings and outputs at the end, and do it.

Rules still moving

Sources

All sources retrieved 7 October 2026.

  1. European Data Protection Board, Guidelines 1/2026 on processing of personal data for scientific research purposes, adopted 15 April 2026, version for public consultation (executive summary; sections 4.1 to 4.4, 5 and 8.3).
  2. European Data Protection Board, Public consultation page for Guidelines 1/2026 (consultation 16 April to 25 June 2026) and news release, 16 April 2026.
  3. European Data Protection Board, EDPB sheds light on anonymisation and web scraping for generative AI, 8 July 2026.
  4. European Data Protection Board, Data controller or data processor (SME data protection guide).
  5. European Data Protection Board, Be compliant: how to conduct a DPIA (SME data protection guide).
  6. European Data Protection Board, International data transfers (SME data protection guide).
  7. OpenAI Help Center, How your data is used to improve model performance.
  8. OpenAI, Enterprise privacy at OpenAI, updated 8 January 2026.
  9. Autoriteit Persoonsgegevens, Datalekkenrapportage 2024, July 2025 (in Dutch).
  10. University of Helsinki, Use of generative artificial intelligence in research.
  11. University of Oslo, TSD LLM: large language models inside TSD, 28 September 2026.
  12. European Commission, Adequacy decisions.
  13. Court of Justice of the European Union, Press release No 106/25, Judgment of the General Court in Case T-553/23, Latombe v Commission, 3 September 2025.
  14. AI Act (as amended), Article 2: Scope, and Recital 25, AI Act Service Desk explorer.
  15. European Commission, AI Literacy: Questions & Answers.
  16. AI Act (as amended), Article 5: Prohibited AI practices, point (1)(f).
  17. European Research Council, Open Science in ERC projects.
  18. European Parliament Legislative Observatory, Procedure 2025/0360(COD), Digital Omnibus, status: awaiting committee decision.

About this page

Sources checked on 7 October 2026. We read the EDPB's research guidelines in full where cited, the EDPB's guidance pages, the vendor pages named, and the university and supervisory authority documents listed. EUR-Lex was only partly available on that date, so GDPR provisions are described as set out in EDPB guidance. Vendor terms are summarised for the plans named and can change; check the current terms for the plan your institution uses. This page is general information for researchers and DPOs, not legal advice, and national research and ethics rules vary. If you spot an error, please contact us and we will correct it.

Frequently asked questions

Can I upload interview transcripts to ChatGPT for my research?

Only if your institution has approved the tool and plan for that data. Consumer accounts and institutional plans differ: OpenAI says it may use content from its services for individuals to train models unless you opt out, while by default it does not use inputs from ChatGPT Edu, Enterprise, Business or the API. Your university also needs a processor agreement and a legal basis covering the step.

What is the legal basis for processing personal data in university research?

For public universities it is often performance of a task in the public interest (Article 6(1)(e) GDPR), which must rest on Union or national law. Consent is possible but the EDPB warns against relying on it where there is an imbalance of power. Special category data also needs an Article 9(2) condition, such as research under Union or national law (Article 9(2)(j)) or explicit consent.

Is pseudonymised research data still personal data?

Yes. Pseudonymised data can still be attributed to a person with additional information, so the GDPR still applies. Only anonymised data, where people are no longer identifiable, falls outside the GDPR. The EDPB says researchers should use anonymised data where the purpose allows, otherwise pseudonymised data.

Is consent on a research consent form the same as GDPR consent?

Not necessarily. The EDPB distinguishes consent to take part in research, required by ethics rules, from consent as a GDPR legal basis. Ethical consent can be a safeguard under Article 89 even when the legal basis is public interest.

Do I need a DPIA for a research project using AI tools?

A DPIA is mandatory where processing is likely to result in high risk. The EDPB's criteria include sensitive data, vulnerable data subjects, systematic monitoring and innovative technology; in most cases two criteria should trigger one. Qualitative projects that use AI tools on sensitive interviews may meet that threshold, so ask your DPO.

Can research data be transferred to US cloud services?

Transfers to US organisations certified under the EU-US Data Privacy Framework can rely on the Commission's adequacy decision. Otherwise you need another transfer tool, such as standard contractual clauses with a transfer impact assessment. Check where each of the vendor's sub-processors is located.