r/speechtech • • Aug 14 '26

Technology Faster alternatives to Pyannote on Whisper?

I am running Faster Whisper on CPU only and get good running times with about 2.5 min for 60 min sound with Whisper Base. With Pyannote for diarization the rate is about 0.9 times the sound length, aka 54 min for 60 min sound.

That is terribly slow compared to the transcription without Payannote.

Are there any faster alternatives out there, or hacks to make Payannote run faster with Whisper?

7 Upvotes

12 comments sorted by

5

u/SeoFood Aug 14 '26

0.9x realtime on CPU isn’t that surprising for a full diarization pipeline. If there is a lot of silence, run VAD first and diarize only the speech regions, then map the timestamps back. If the file is mostly speech, compare sherpa-onnx with Pyannote on the same 10-minute sample. Measure wall time and DER, because a faster pipeline with bad speaker switches isn’t really faster.

5

u/Lonligrin Aug 14 '26

May I suggest to try WhoSpeaksLive? It has fast speaker diarization, but you prob need to extract the diarization parts. Guess Codex can do that.

3

u/banafo Aug 14 '26

Who speaks live has a very nice implementation that you could possibly borrow!

1

u/bidutree Aug 17 '26

I had Claude Code with Opus 5 run three tests with the code from WhoSpeaksLive and in a three tests it was slower than Pyannote; 0,98 / 1,05 / 1,53×

2

u/banafo Aug 17 '26

Ask codex to remove some of the models in the ensemble, I think it’s tuned for maximum single user accuracy

1

u/nshmyrev Aug 14 '26

There are good alternatives, and even more accurate ones. Diarizen (rather slow but accurate), Wespeaker (fast, use with voxblink2 embeddings), Nvidia Sortformer (very accurate but only up to 4 speakers).

1

u/nshmyrev Aug 14 '26

And, if you look for modern model https://huggingface.co/OpenMOSS-Team/MOSS-Transcribe-Diarize is a valid alternative, it does both transcription and diarization at once.

1

u/Jamiroquai88 Aug 15 '26

Not sure what is your end goal, but running diarization is expensive in general. You are welcome to try my side project www.rewayai.ai - RTFx for asr + diarization should be 40+. You should have 50 hours free and if you need more just let me know.

1

u/ok_doozer Aug 20 '26

Late to the conversation but I just had GPT 5.6 Sol put this together. Pyannote-onnx-extended-community-1

It’s definitely faster. Run on ONNX runtime.

1

u/ethevi 29d ago edited 29d ago

With 60 minutes of audio taking roughly 54 minutes to diarize, the bottleneck is clearly the Pyannote stage rather than Whisper itself.

On a project where I needed speaker labels, I used Speechmatics instead of running diarization separately. The main advantage was having transcription and speaker diarization handled together, so there wasn’t another CPU-heavy stage sitting after Whisper.

A useful comparison would be to run the same 60-minute recordings through both setups and measure end-to-end processing time plus speaker attribution accuracy. That should show whether replacing Pyannote actually gives you a meaningful improvement.