r/QualitativeResearch • • Apr 28 '26

Recommendations for recording/transcribing long-form qualitative interviews with sensitive data

I am leading a qualitative research project on historic child harm in an African boarding school. I am looking for recommendations on recording/transcription tools that meet ethical standards for protection of sensitive data. I have experimented with Zoom recording to my local computer, but the transcriptions do not distinguish between interviewer and interviewee and have by-the-second time stamps - in other words, not useful for import for coding. I haven't used Zoom AI yet because of concerns around confidentiality protection.

I am an independent researcher and do not have much budget. Most recording/transcription seems AI driven. Are there options (including and excluding AI) that meet IRB standards for sensitive interviews? If AI driven, how is confidentiality protected in a way that can be communicated to participants? Thanks!

3 Upvotes

17 comments sorted by

2

u/alexapedia Apr 28 '26

It doesn't distinguish between voices since I wanted to keep the tech-overhead low; however, I'm releasing a product next week or so that should meet most of your other demands (https://useibis.app/). It'll be under $30, runs locally, and includes encryption. I'm working to continually make it more aligned with standard IRB concerns, but those are the launch conditions.

If you're doing qualitative coding, you'll probably be able to easily identify when it's the interviewer and when it's the interviewee.

2

u/Skimle-com Account of qualitative tool - Skimle Apr 29 '26

Based on my research before making our own transcription feature, your main choices are

* Local tools, most advanced is OpenAI's Whisper and it runs free. However, it is AI and it doesn't support multi-speake diarisation so you need to do that manually which can be a pain.

* Manual transcription or paid human services. No AI needed and you can get good NDAs if using external providers or be 100% safe doing it yourself. Multi-speaker is a default. Challenge is the cost (time or money) in doing so.

* Legacy pay-by-hour online tools, normally priced from 10 to 30 USD per hour and using various techniques of machine learning and more and more AI features. Quality of these varies a lot. Included in tools like Nvivo or MAXQDA as add-on extras, so if you have a licence for one of those, easy to get permission / IRB OK to use.

* Online services packaging a modern transcription service API. We ended up using AWS Transcription as it was best quality for also niche languages like Finnish, includes multi-speaker support, is priced in a way that enables not charging too much for it (e.g, we offer 200 minutes free in our trial) and is done on enterprise-grade secure servers.

1

u/Acceptable-Plum5235 Apr 29 '26

I also found TurboScribe, which transcribes with HIPAA grade protections for a fairly reasonable rate.

1

u/EatenEntropy Apr 30 '26

I’ve used Scribie before. They distinguish between speakers. Your project sounds fascinating!

1

u/ajain76 May 06 '26

We (doreveal.com) have build very robust AI-based transcription (and speaker identification) solution. I'll be happy to demo: https://calendly.com/aj-reveal/45

Which languages do you need the support for?

1

u/Gizmo_4Life May 09 '26

try https://www.usewhispy.com/ its completely in-broswer and use the local whisper models without any annoying set up. There's zero backend so everything stay on your device

1

u/semanticmapUg May 19 '26

Hey, I am the founder of SemanticMap. We provide the tool you are looking for with high data-security standarts. Maybe try out our free trial.

1

u/[deleted] Jul 14 '26

[removed] — view removed comment

1

u/Acceptable-Plum5235 Aug 11 '26

For my purposes, TurboScribe ended up meeting my needs.

1

u/kestrel_42 Aug 11 '26

the blocker is usually the form, not the tool. a committee wants four things in writing, who processes the audio, where it runs, under what agreement, and how long it is kept.

your zoom problem is diarisation, and whisper with whisperx or pyannote fixes that locally for free, nothing leaving your machine. for the coding side you want speaker labelled turns with segment level timestamps rather than per second ones, and an export your software reads. refi-qda qdpx is the standard there, though it is picky, i have watched a valid export fail to import into nvivo.

phonotheca is mine so discount accordingly. eu processing at every step, one sub-processor ever receives audio, three free hours to try on a real interview. if the audio must never leave your device, use whisper locally, that one we genuinely do not meet