GDPR for researchers using AI and cloud tools: transcription, surveys and AI assistants
Researchers now routinely send interview recordings to transcription services, run surveys on cloud platforms and paste text into AI assistants. Each of those steps is processing of personal data. This guide explains, with sources, what the GDPR expects from you and your institution, and gives a checklist to run before you choose a tool.
The short answer
- Your university, not the vendor, is usually the controller of research data. Any transcription, survey or AI service that handles the data for you is a processor and needs a contract that meets Article 28 GDPR. A personal consumer account does not give you one.
- Pick the legal basis before the tool. At public universities this is often public interest (Article 6(1)(e)), with Article 9(2)(j) or explicit consent for sensitive data, plus Article 89 safeguards such as pseudonymisation and ethics review.
- Pseudonymised is not anonymous. Do not tell participants their data is anonymous if a key, a voice or a rich transcript can still identify them.
- Consumer AI accounts are the main risk. The Dutch DPA has treated staff pasting personal data into AI chatbots against their employer's rules as a reportable data breach.
- Before you choose a tool, run the checklist below and talk to your data protection officer (DPO).
Prefer a tool built in the EU? Kahubi, from Avidemic AB in Sweden, covers an AI assistant for research work with EU hosting and, for institutions, only European subprocessors. See how Kahubi handles research data
Who is responsible: controller and processor
The GDPR assigns responsibility by role. The controller decides the purposes and means of processing; a processor acts only on the controller's instructions [4]. In most university research the controller is the university (sometimes jointly with partner institutions), and the researcher acts on its behalf. That has two practical consequences:
- A tool vendor that transcribes, hosts or analyses data for you is normally a processor. The controller-processor relationship must be governed by a contract that, among other things, limits processing to the controller's instructions, requires security and confidentiality, requires approval before sub-processors are engaged, and requires deletion or return of data at the end [4].
- Accepting a vendor's consumer terms with a personal account is not the same as your university concluding that contract. Where several institutions share a project, the EDPB says roles should be assessed and documented, as this determines accountability and where participants exercise their rights [1].
The EU AI Act adds a separate layer. Universities and researchers who use AI tools are deployers, and the AI literacy duty in Article 4 has applied since 2 February 2025. The research exclusion in the AI Act covers AI systems developed solely for scientific research, not general tools used to do research [14][15]. Our AI Act guide for universities covers this in detail.
Legal bases in research
The EDPB's Guidelines 1/2026 on processing personal data for scientific research were adopted on 15 April 2026 as a version for public consultation; the consultation closed on 25 June 2026 and we found no final version when we checked [1][2]. They are the most detailed EU-level guidance available and the basis for this section.
Public interest: Article 6(1)(e)
The EDPB describes public interest as a legal basis "of particular importance for scientific research" [1]. Under Article 6(3) GDPR, it must rest on Union or Member State law to which the controller is subject, but that law does not need to spell out how the research is done. The guidelines give national examples, including the Swedish Higher Education Act and the Finnish Data Protection Act. Public authorities cannot rely on legitimate interests (Article 6(1)(f)) for processing in the performance of their tasks [1].
Consent: Article 6(1)(a)
Consent is possible, and the EDPB accepts "broad consent" for an area of research where purposes are not fully known, and "dynamic consent" project by project, both with extra safeguards [1]. But consent must be freely given. The EDPB says that where there is an imbalance of power, for example research involving pupils, convicted persons or unemployed people, and there is doubt, controllers should not rely on consent [1].
Validemic's analysis A researcher interviewing their own students, or a PhD student surveying colleagues in their own department, is in a position where consent as a legal basis is hard to defend. Many universities therefore use public interest as the legal basis and keep consent as an ethical safeguard.
Ethical consent is not the same as GDPR consent
Ethics rules often require consent to take part in research. The EDPB stresses that this should be distinguished from consent as a GDPR legal basis. Ethical consent can be an additional safeguard under Article 89, even where the legal basis is something else. If you do want GDPR consent, the request must be clearly distinguishable from the request to participate [1].
Special category data: Article 9(2)(j) and explicit consent
Health, ethnicity, political opinions, religion, sexual orientation, trade union membership, genetic and biometric data need an Article 9(2) condition as well. The EDPB explains that Union or Member State law can provide a research derogation under Article 9(2)(j), with suitable and specific safeguards, and that controllers must be able to show that the law applies and that its safeguards are in place. The guidelines list national examples such as section 10 of the Danish Data Protection Act and the Swedish Ethical Review Act. Where no such law applies, explicit consent is the alternative [1].
Interviews often produce special category data even when the topic seems neutral: a participant mentions an illness or a religious practice. Plan for it in the legal basis and the tool choice.
Further processing and retention
The EDPB says further processing for scientific research is presumed compatible with the original purpose, so no compatibility test is needed, but the controller must still check lawfulness. Data may be stored for longer for research purposes, for example for follow-up projects [1].
Article 89 safeguards, pseudonymisation and anonymisation
Research benefits from flexibilities in the GDPR, but Article 89(1) requires appropriate safeguards, in particular data minimisation. The EDPB summarises the order of preference [1]:
- Anonymised data first, where the research purpose can be met with it. Truly anonymous data falls outside the GDPR, but the anonymisation process itself is processing that must comply.
- Pseudonymised data where anonymisation is not possible, for example where you need to follow people over time.
- Directly identifying data only where strictly necessary and proportionate.
Pseudonymisation is defined in Article 4(5) GDPR as processing so that data can no longer be attributed to a specific person without additional information that is kept separately and protected [1]. Pseudonymised data is still personal data. The EDPB also lists other safeguards: independent or ethical oversight, secure processing environments, privacy-enhancing technologies, protective measures when publishing results and confidentiality arrangements [1].
On anonymisation specifically, the EDPB adopted draft Guidelines 02/2026 on 8 July 2026, open for consultation until 30 October 2026. They test anonymisation against three risks: singling out a record, linking records, and inferring information [3].
| Data state | Example in qualitative research | GDPR applies? |
|---|---|---|
| Identifying | Audio recording; transcript with names | Yes |
| Pseudonymised | Transcript with names replaced by codes, key held separately | Yes |
| Possibly anonymised | Aggregated themes with no quotes that could identify anyone | Only if re-identification is not reasonably likely |
Validemic's analysis Removing names from a rich interview transcript rarely makes it anonymous. Job titles, places, events and turns of phrase can identify someone in a small field. Treat de-identified transcripts as pseudonymised unless your DPO has assessed otherwise.
Interview recordings and transcription
An interview recording is personal data, and a voice can identify a person. Transcription is often the first point where research data leaves the university's own systems. The options, from most to least contained, are:
- Transcription inside a secure institutional environment. Some universities run AI services within their own secure research infrastructure. The University of Oslo, for example, announced on 28 September 2026 a large language model inside its TSD secure environment, stating that sensitive data can be processed "without leaving the TSD environment" [11].
- An institutionally licensed service with a processor agreement, known data location and contractual terms on AI training and sub-processors.
- A consumer or free account with a cloud transcription service, accepted by the researcher personally. This is the option most likely to lack a processor agreement and to allow the vendor wider use of the content.
Questions to answer for any transcription tool:
- Where are recordings and transcripts stored and processed, and by which sub-processors?
- Does the vendor use customer audio or text to train or improve models, and can the institution switch that off contractually?
- How long are recordings kept after transcription, and can you delete them on demand?
- Does the tool add features such as summaries, speaker identification or sentiment, which create new processing you may not have told participants about? (Under the EU AI Act, inferring emotions is prohibited in workplaces and education institutions, outside narrow medical or safety reasons. That matters if the participants are staff or students being studied in that setting [16].)
Prefer a tool built in the EU? Kahubi, from Avidemic AB in Sweden, covers an AI assistant for research work with EU hosting and, for institutions, only European subprocessors. See how Kahubi handles research data
Our fact sheets on Otter.ai, Trint, Rev and Zoom set out what each vendor publicly documents about hosting, agreements and training, plan by plan.
Surveys
Survey platforms are processors too, and the same contract and transfer questions apply. Points specific to surveys:
- "Anonymous" surveys often are not. Platforms may log IP addresses, device data or email invitations, and free-text answers can identify people. The EDPB says participants must not be given the impression that their data will be anonymised if it will in fact be pseudonymised and still identifiable [1].
- Free text attracts special category data. If you ask about wellbeing, workplace experiences or identity, expect health and other sensitive data and plan the Article 9 condition.
- AI features inside survey platforms, such as automatic text analysis, may involve further sub-processors. Check whether they are on by default.
- Panels and recruitment through third parties raise questions about who is controller for which step. Document the roles [1].
See our fact sheets on Qualtrics, SurveyMonkey and Google Forms.
Uploading data to AI assistants
Pasting a transcript into a chatbot to summarise it, asking an assistant to code qualitative data, or uploading a dataset for analysis are all processing operations. Before doing any of them with personal data, the same questions apply as for any processor: is there an agreement, where does the data go, and can the vendor reuse it?
Supervisory authorities have started to treat this as a security issue. In its report on data breaches in 2024, published in July 2025, the Dutch Autoriteit Persoonsgegevens said it had received several breach notifications from organisations whose staff entered personal data of patients or customers into AI chatbots. It noted that staff often use chatbots on their own initiative and against agreements with their employer, and that entering personal data in that situation is a (reportable) data breach [9].
University guidance points the same way. The University of Helsinki's guidelines on generative AI in research tell researchers not to use sensitive or confidential data or unpublished results as input, because of the risk of leaks, and to report substantial AI use, including the tool, version, date and purpose [10].
Practical rules that follow:
- Use only tools and plans your institution has approved for personal data, and only for the data classes they are approved for.
- Minimise first: remove names and direct identifiers before upload, even with an approved tool.
- Never upload special category data or data from vulnerable participants to a consumer account.
- Check whether the assistant keeps files, memories or chat history, and how long.
- Record AI use in your methods section and data management plan.
For specific tools, see ChatGPT, Microsoft Copilot, Gemini, NotebookLM and Elicit.
Consumer versus institutional accounts
The same brand can have very different terms depending on the plan. OpenAI's documentation is a clear example (checked 7 October 2026):
| Services for individuals (e.g. ChatGPT personal accounts) | Business services (ChatGPT Business, Enterprise, Edu, API) | |
|---|---|---|
| Use of content for model training | OpenAI "may use your content to train our models"; users can opt out in data controls [7] | Not used to improve models by default [7][8] |
| Retention control | Set by the user's settings | Customer controls retention for Enterprise, Healthcare and Edu [8] |
| Access control | Individual | Organisation controls access, with SAML SSO [8] |
Note one detail from OpenAI's help page: even after opting out on a personal account, giving feedback on a response (for example a thumbs up or down) may allow the whole conversation to be used for training [7].
Validemic's analysis An institutional plan does not by itself make a use lawful. It gives your university the contractual controls it needs as controller. The legal basis, minimisation, transfers and participant information still have to be right. Conversely, a consumer account with training switched off still lacks a processor agreement between your university and the vendor.
Reviewing a vendor right now? Validemic checks the vendor's documents against GDPR and the EU AI Act and cites every finding. Try the demo workspace
Transfers outside the EEA
If a tool transfers personal data outside the EEA, for example to a hosting location or sub-processor abroad, the transfer rules in Chapter V GDPR apply. Ask vendors where data is stored and from where staff and sub-processors can access it. The EDPB's guide for small organisations summarises the tools [6]:
- Adequacy decisions: transfers can proceed without further safeguards. The Commission's list includes, among others, Japan, the Republic of Korea, Switzerland, the United Kingdom and the United States for commercial organisations participating in the EU-US Data Privacy Framework [12].
- Standard contractual clauses, with a transfer impact assessment and supplementary measures where needed.
- Binding corporate rules, codes of conduct and certification.
- Article 49 derogations, such as explicit consent, which the EDPB describes as a last resort.
For US vendors, check whether the specific company is certified under the Data Privacy Framework. On 3 September 2025 the EU General Court dismissed an action to annul the framework's adequacy decision (Case T-553/23, Latombe v Commission). The Court's press release notes that an appeal on points of law may be brought before the Court of Justice [13]. Check the current position before relying on the framework for a long project.
Our transfer mechanism tool walks through which tool applies to a given vendor and data flow.
Participant information sheets
The GDPR requires transparent information to participants, and the EDPB says that for long-running research the controller should keep participants informed throughout, including updates if processing changes. The duty also applies where a processor holds the data [1]. For tool use, the information sheet should cover:
- the controller (your university) and the DPO's contact details;
- the legal basis and, for sensitive data, the Article 9 condition;
- which categories of service providers will process the data (for example "a transcription service" or "a survey platform"), and whether any are outside the EEA, with the safeguard used;
- whether data will be pseudonymised or anonymised, and when; participants should also be told that once data is anonymised it cannot be retrieved or withdrawn [1];
- retention periods and what happens to recordings after transcription;
- participants' rights and their limits in research, which the EDPB discusses for erasure and objection [1].
Validemic's analysis Example wording, to be adapted with your DPO and ethics committee:
"We will record the interview and use [service], a transcription service contracted by [University], to produce a written transcript. [Service] processes the recording on our behalf, in [location], and may not use it for its own purposes, including training AI models. We delete the recording [timing]. In the transcript we replace your name and other details that could identify you with a code. The key linking codes to names is stored separately at [University]. This means your data is pseudonymised, not anonymous, until we delete the key on [date]."
Avoid promising things the tool cannot guarantee, such as "your data will never leave the EU", unless the vendor's contract and sub-processor list support it.
Ethics review, DPIAs and data management plans
Ethics review
Research ethics review and the GDPR overlap but are separate. The EDPB treats adherence to ethical standards as one of six indicative factors of scientific research, and names ethical oversight as an Article 89 safeguard [1]. If you change tools after approval, for example switching from manual transcription to an AI service, check whether your ethics committee and DPO need to be told.
DPIA
A DPIA is mandatory where processing is likely to result in high risk. The EDPB lists criteria including sensitive or highly personal data, vulnerable data subjects, large-scale processing, matching datasets, systematic monitoring and innovative technology, and says that in most cases two criteria should lead to a DPIA [5]. A qualitative study on health experiences that uses an AI transcription service may well meet two. The EDPB research guidelines say the choice of data format (anonymised, pseudonymised, identifying) should be made when the research plan is drafted, as part of a risk analysis or, where relevant, a DPIA [1]. Use our DPIA screening tool for a first check.
Data management plans
Some funders require a data management plan (DMP). For example, ERC grantees under Horizon Europe must submit a DMP by the end of month 6 of the project, and the ERC applies the principle "as open as possible, as closed as necessary" [17]. A DMP is a good place to record which tools touch personal data, where, under which agreement, and when data will be pseudonymised, anonymised or deleted. Keep it consistent with the information sheet and the DPIA.
Checklist before you use a tool
Work through these questions before any personal data goes into a transcription service, survey platform or AI assistant. Share the answers with your DPO.
- Is it personal data? Voices, transcripts, free text and pseudonymised data all count. If you can genuinely work with anonymised data, do.
- Is the tool approved? Is it on your university's approved list for this data class, under an institutional licence rather than a personal account?
- Is there a processor agreement? Between your university and the vendor, covering instructions, security, sub-processors and deletion [4].
- Does the vendor use your data for its own purposes? Check training and product improvement terms for the exact plan you will use.
- Where is the data processed? Including sub-processors and support access. If outside the EEA, which transfer tool applies?
- What is the legal basis? Article 6, plus Article 9 for sensitive data. Does your ethics consent form also need to be a GDPR consent?
- Which Article 89 safeguards apply? Pseudonymisation timing, access control, separate key storage, secure environment.
- Do you need a DPIA? Screen against the EDPB criteria [5].
- Do participants know? Does the information sheet describe this tool category, location, retention and pseudonymisation accurately?
- Are new AI features switched on? Summaries, sentiment or emotion analysis and speaker identification add processing; some may be prohibited in education or workplace settings under the AI Act [16].
- Is it in the DMP? Record the tool, data flow and retention.
- Can you delete it? Confirm how to delete recordings and outputs at the end, and do it.
Rules still moving
- The EDPB research guidelines (1/2026) and anonymisation guidelines (02/2026) are consultation versions; final texts may change [2][3].
- The Commission's Digital Omnibus proposal of 19 November 2025, which would amend the GDPR, is still at committee stage in the European Parliament (procedure 2025/0360(COD)) and is not law [18].
- The AI Act's high-risk rules for education apply from 2 December 2027 after the AI Omnibus; AI literacy and the prohibitions already apply. See our AI Act guide.
Sources
All sources retrieved 7 October 2026.
- European Data Protection Board, Guidelines 1/2026 on processing of personal data for scientific research purposes, adopted 15 April 2026, version for public consultation (executive summary; sections 4.1 to 4.4, 5 and 8.3).
- European Data Protection Board, Public consultation page for Guidelines 1/2026 (consultation 16 April to 25 June 2026) and news release, 16 April 2026.
- European Data Protection Board, EDPB sheds light on anonymisation and web scraping for generative AI, 8 July 2026.
- European Data Protection Board, Data controller or data processor (SME data protection guide).
- European Data Protection Board, Be compliant: how to conduct a DPIA (SME data protection guide).
- European Data Protection Board, International data transfers (SME data protection guide).
- OpenAI Help Center, How your data is used to improve model performance.
- OpenAI, Enterprise privacy at OpenAI, updated 8 January 2026.
- Autoriteit Persoonsgegevens, Datalekkenrapportage 2024, July 2025 (in Dutch).
- University of Helsinki, Use of generative artificial intelligence in research.
- University of Oslo, TSD LLM: large language models inside TSD, 28 September 2026.
- European Commission, Adequacy decisions.
- Court of Justice of the European Union, Press release No 106/25, Judgment of the General Court in Case T-553/23, Latombe v Commission, 3 September 2025.
- AI Act (as amended), Article 2: Scope, and Recital 25, AI Act Service Desk explorer.
- European Commission, AI Literacy: Questions & Answers.
- AI Act (as amended), Article 5: Prohibited AI practices, point (1)(f).
- European Research Council, Open Science in ERC projects.
- European Parliament Legislative Observatory, Procedure 2025/0360(COD), Digital Omnibus, status: awaiting committee decision.
About this page
Sources checked on 7 October 2026. We read the EDPB's research guidelines in full where cited, the EDPB's guidance pages, the vendor pages named, and the university and supervisory authority documents listed. EUR-Lex was only partly available on that date, so GDPR provisions are described as set out in EDPB guidance. Vendor terms are summarised for the plans named and can change; check the current terms for the plan your institution uses. This page is general information for researchers and DPOs, not legal advice, and national research and ethics rules vary. If you spot an error, please contact us and we will correct it.
Frequently asked questions
Can I upload interview transcripts to ChatGPT for my research?
Only if your institution has approved the tool and plan for that data. Consumer accounts and institutional plans differ: OpenAI says it may use content from its services for individuals to train models unless you opt out, while by default it does not use inputs from ChatGPT Edu, Enterprise, Business or the API. Your university also needs a processor agreement and a legal basis covering the step.
What is the legal basis for processing personal data in university research?
For public universities it is often performance of a task in the public interest (Article 6(1)(e) GDPR), which must rest on Union or national law. Consent is possible but the EDPB warns against relying on it where there is an imbalance of power. Special category data also needs an Article 9(2) condition, such as research under Union or national law (Article 9(2)(j)) or explicit consent.
Is pseudonymised research data still personal data?
Yes. Pseudonymised data can still be attributed to a person with additional information, so the GDPR still applies. Only anonymised data, where people are no longer identifiable, falls outside the GDPR. The EDPB says researchers should use anonymised data where the purpose allows, otherwise pseudonymised data.
Is consent on a research consent form the same as GDPR consent?
Not necessarily. The EDPB distinguishes consent to take part in research, required by ethics rules, from consent as a GDPR legal basis. Ethical consent can be a safeguard under Article 89 even when the legal basis is public interest.
Do I need a DPIA for a research project using AI tools?
A DPIA is mandatory where processing is likely to result in high risk. The EDPB's criteria include sensitive data, vulnerable data subjects, systematic monitoring and innovative technology; in most cases two criteria should trigger one. Qualitative projects that use AI tools on sensitive interviews may meet that threshold, so ask your DPO.
Can research data be transferred to US cloud services?
Transfers to US organisations certified under the EU-US Data Privacy Framework can rely on the Commission's adequacy decision. Otherwise you need another transfer tool, such as standard contractual clauses with a transfer impact assessment. Check where each of the vendor's sub-processors is located.