DEV Community

Micky Irons
Micky Irons

Posted on

Research Interview Transcripts in AI Tools: UK GDPR

Only if your lawful basis, your participant information and your security all hold up. Interview transcripts are usually personal data, often special category data, and participants were told how their words would be handled. Public AI tools rarely match that promise. Mickai indexes transcripts into private knowledge bases on hardware the institution owns, so nothing leaves the estate.

Are interview transcripts personal data?

Almost always, yes. A transcript records what an identifiable person said, in their own words, usually next to a participant code that links back to a consent form and a contact record. That makes it personal data under the UK GDPR, and the transcript on its own is not the whole picture.

Researchers often assume that swapping names for codes settles it. It does not. The ICO's position is that pseudonymised data is still personal data, because the key that re-identifies the participant exists somewhere. Then there are the indirect identifiers you cannot code out: a job title, an employer, a town, a committee, an unusual anecdote. Picture a study of forty local-authority planning officers: one participant describing a particular dispute is recognisable to any colleague who reads it. Audio carries more again, since a voice is distinctive even when the words are not. Treat the recording, the transcript, the coding frame and the linkage file as one governed set, not four separate files with four different rules.

What did we promise participants, and does AI use break it?

Read the participant information sheet before you read a supplier's terms. Most sheets say the recording will be heard by the named research team, held on institutional systems, and destroyed after a stated period. If that is what you promised and you paste a transcript into a hosted tool, you have broken the promise regardless of what your lawful basis says.

The failure is rarely dramatic. It is a mismatch between the document a participant signed and where their words actually went. That matters legally, because transparency under Article 13 turns on telling people what will really happen, and practically, because your committee will read both documents side by side. A hosted tool that retains content for service improvement, allows supplier staff to review inputs, or processes outside the UK introduces a processor and possibly a purpose that nobody disclosed. The workable test is blunt: could you tell the participant afterwards, in one sentence, exactly where their interview went, and would they shrug? If not, do not do it.

Which lawful basis applies to research processing?

You still need one. The research provisions do not create a lawful basis, and using a different tool does not change the basis you already recorded. In UK higher education, public task under Article 6(1)(e) is commonly used for research processing, while legitimate interests or contract is more usual in commercial settings.

The recurring confusion is between ethics consent and UK GDPR consent. They are different instruments. Ethics consent discharges research governance and the common-law duty of confidence. UK GDPR consent is a specific lawful basis carrying a withdrawal right that is genuinely hard to honour once a transcript is coded into a longitudinal dataset. That is precisely why many institutions require informed consent for ethics while relying on public task for data protection. Both can be true at once, and the participant sheet should say so plainly rather than implying that withdrawing consent erases the data. Switching analysis tools mid-study does not usually need a new basis, but it will often need a fresh DPIA and a new transparency step.

When do transcripts contain special category data?

More often than the study design predicted. The ICO lists nine special categories: racial or ethnic origin, political opinions, religious or philosophical beliefs, trade union membership, genetic data, biometric data used for identification, data concerning health, sex life, and sexual orientation. A qualitative interview about workplace culture, teaching practice or engineering safety wanders into several of those without anyone intending it.

It arrives as an aside. A participant mentions their union representative, their faith, how they voted on a site committee. Nothing on your data collection form flagged it, and the transcript now needs an Article 9 condition alongside your Article 6 basis. For research, the route is usually Article 9(2)(j), covering research and statistical purposes subject to the Article 89(1) safeguards, with the domestic conditions set out in Schedule 1 of the Data Protection Act 2018. Some of those conditions require an appropriate policy document, so check which one you are relying on. The practical move is to assume special category data will appear in qualitative work and design the storage and access controls for that from the start.

Do the UK GDPR research provisions help?

They help with retention and reuse. They do not help with security. The provisions relax the purpose limitation and storage limitation principles and give exemptions from several individual rights, including access, rectification, erasure and objection. None of that is permission to send transcripts somewhere new.

Two conditions cut the other way. The ICO's safeguards guidance sets out that you cannot rely on the provisions where the processing is likely to cause substantial damage or substantial distress to a participant, or where it is carried out for the purposes of measures or decisions about a particular participant, subject to a narrow statutory exception. Article 89(1) then requires appropriate safeguards in the form of technical and organisational measures, with specific weight on data minimisation. The ICO's own order of questions is instructive: first ask whether the research can be done without personal data at all, then whether anonymous information would serve, then pseudonymisation, and only then the rest of your controls.

One live caveat. The ICO's research provisions guidance carries a notice that it is under review following the Data (Use and Access) Act and may change. Treat the detail as moving, and check the page on the day you write the DPIA rather than quoting a slide from last year.

What is different about commercial and contract research?

The basis usually changes and the contract usually bites harder. Legitimate interests or contractual necessity tends to replace public task, and the funder's agreement often carries confidentiality and data-location clauses stricter than anything data protection law would demand on its own.

In industrially funded work, market research, or a multi-party engineering consortium, a transcript can hold a third party's commercially sensitive material as well as personal data. Sponsor agreements frequently prohibit offshore processing or unapproved subprocessors outright. Breaching that is a contract problem before it is a regulatory one, and it ends collaborations quietly rather than publicly. Settle two further questions in writing before fieldwork starts: which party is controller at which stage, and who answers a subject access request when three institutions hold copies of the same interview.

How can we use AI on transcripts without sending them out?

Run the analysis where the data already sits. That is what we build. The Mickai Sovereign Intelligence Operating System runs on hardware the institution owns, with no data egress, and it is offline capable on supported hardware.

Documents become private knowledge bases, which we call brains: fifty specialised models, indexed and queried on your own machines. SIOS carries 63 studios in total, 14 production-ready at launch and 49 in development. Every consequential action is sealed in the Open Audit Record under ML-DSA-65, the post-quantum signature scheme NIST published as FIPS 204 in 2024. That record is tamper-evident, not tamper-proof, and the distinction is the point. We cannot stop someone altering a stored record. Altering it makes verification fail, and your data protection officer or an external auditor can verify an exported record offline with a public key, using tools that are not ours. Consequential actions wait for a named person to approve them, rather than proceeding because a model judged them reasonable.

One honest limit, since transcripts usually arrive as files. A local OCR runtime has read scanned PDFs in controlled tests, and the extraction and ingestion integration into SIOS is still being completed. If your transcripts are scanned images rather than text, ask us where that work stands before you build a plan around it. We are in closed beta, with one regulated company onboarding as a design partner.

None of this is an argument against the organisations building the compute or the cloud layer. Cloud remains useful for plenty of research work that carries no confidentiality burden. The argument is with the assumption that a regulated institution must rent its intelligence, ship participant data offsite, and take a supplier's word for what happened to it.

What should the ethics application say about AI tools?

Name the tool, state where the processing physically happens, and say who can read the content. Wording such as "AI-assisted analysis may be used" will not survive a competent committee, and it will not match your participant information sheet either.

Cover eight things. What the tool does: transcription, coding, summarising, retrieval. Where it runs, naming the hardware, the region and any supplier. Whether any human at that supplier can access inputs, including for support or moderation. Retention and deletion, including logs and any reuse of content for training. Your Article 6 basis and, where relevant, your Article 9 condition. The Article 89(1) safeguards you are applying, with the DPIA reference. What participants are told, in the words they will actually read. And what status the AI output holds in your analysis, since a summary no researcher checks against the transcript is a finding nobody verified. If the analysis will inform any measure or decision about an individual participant, flag it explicitly, because the research provisions are not available for that.

Frequently asked questions

Can I upload an interview transcript to an AI tool?

Not without checking three things: your recorded lawful basis, what the participant information sheet promised, and where the tool actually processes and retains content. A transcript is personal data even when names are coded out. If the sheet said data stays on institutional systems, a hosted tool breaks that promise whatever your lawful basis permits.

Does anonymising a transcript make it safe to use with AI?

Only if it is genuinely anonymous, which is a high bar in qualitative work. The ICO is clear that pseudonymous data remains personal data, because a key still links back to the participant. Job titles, employers, committees and distinctive anecdotes re-identify people after names are removed. Assess re-identification risk before assuming the data is out of scope.

Do participants need to consent to AI analysis?

Consent for research ethics and consent as a UK GDPR lawful basis are different instruments. Many institutions rely on public task for data protection while still requiring informed ethics consent. Either way, participants must be told accurately what happens to their words, so if you add an AI tool mid-study, update the information sheet and the DPIA.

Do the research provisions let us skip consent?

No. The research provisions do not create a lawful basis, so you still need one and you still need it recorded. They relax purpose limitation and storage limitation and exempt you from some individual rights. They cannot be relied on where processing is likely to cause substantial damage or distress, or where it drives decisions about particular people.

Can AI work on interview transcripts without an internet connection?

Yes. SIOS is offline capable on supported hardware, so indexing, retrieval and analysis run on machines the institution owns, with no data egress. Every consequential action is sealed in the Open Audit Record under ML-DSA-65, and an auditor can verify an exported record offline using a public key and tools that are not ours.

What should an ethics application say about AI tools?

Name the tool, state where the processing physically happens, and say who at any supplier can read the content. Then give retention and deletion, your Article 6 basis, any Article 9 condition, the Article 89(1) safeguards you apply, the DPIA reference, and the exact wording participants will see. Say what status AI output holds in your analysis.


Related briefings

Data protection and UK GDPR

Governance, audit and oversight

Part of a series of 60 briefings on deploying and governing AI in UK regulated organisations, archived with a DOI at 10.5281/zenodo.22975756.

Evaluating AI for a regulated organisation? Mickai runs on hardware you own, offline. Consequential actions wait for a named person to approve them, and what the AI did is sealed into a signed record an auditor can check without us. Applications for the invitation-only closed beta are open. Apply for the closed beta.

Written by Micky Irons, founder and chief executive of Mickai LTD.

Top comments (0)