Skip to content

Let Whisper listen where the diarization found speakers, so a quiet second voice is transcribed instead of cut away by the VAD - #361

Draft
happyarts wants to merge 1 commit into
kaixxx:mainfrom
happyarts:fix/whisper-speech-map-from-diarization
Draft

happyarts wants to merge 1 commit into
kaixxx:mainfrom
happyarts:fix/whisper-speech-map-from-diarization

Conversation

@happyarts

Copy link
Copy Markdown
Contributor

With speaker detection on, Whisper now takes its speech map from the diarization instead of from Silero. A voice much quieter than the one beside it — an interviewer far from the microphone, a second person in the room — used to be cut away by the VAD before Whisper ever heard it. Now it is transcribed. Without speaker detection nothing changes.

What changes

noScribe/whisper_mp_worker.py

  • speech_map(turns, n_samples) turns the diarization's [start, end] turns into the chunks faster-whisper expects. Each turn gets 50 ms of padding, gaps under 0.5 s are closed, and the result is clipped to the audio. These are the same amounts as the Silero options a few lines below.
  • _speech_map_from(turns) replaces faster_whisper.transcribe.get_speech_timestamps, the one call through which faster-whisper gets its speech map, for the duration of detect_language() and transcribe() only. It always restores the original, also on an exception. Everything else in faster-whisper runs exactly as with Silero: collect_chunks and the mapping of timestamps back to the recording.
  • If there are no turns, or a future faster-whisper release no longer uses that name, Silero stays in charge. A test pins the name, so such a release fails the suite instead of silently switching the feature off.

noScribe/main.py

  • _run_whisper_subprocess_stream(..., diarization=None) passes the turns to the worker as speech_turns, in seconds.
  • diarization = None before the diarization block, so the call also works with speaker detection off.

tests/test_whisper_speech_map.py: 13 tests, all without a model:

  • the chunk arithmetic;
  • that faster-whisper still looks up the name that is swapped;
  • that the swap is undone, also after an exception;
  • that the worker really transcribes and detects the language inside the swap (with a stand-in model, in the style of test_whisper_audio_lifetime.py);
  • that the turns reach the worker through _process_single_job, and are None without speaker detection.

Why: measured

Silero drops quiet speech. Measured on the Krisp Voice Isolation benchmark set: real recordings, a primary speaker plus a second voice 15-20 dB lower in the same microphone, with windows where only the second voice speaks.

  • Silero let through 27-41 % of the second voice's speaking time. Of the words in what it dropped, Whisper wrote 2-3 %.
  • The diarization's turns covered 67 % of it. With them as the speech map, Whisper wrote 50-65 % of those words instead of 27-46 %.
  • It wrote no words at all in the set's pure-noise windows, and nothing changed on the phone calls, which have one speaker.

Against manual transcripts:

Silero (today) diarization as speech map
AMI, all 23 meetings (116 579 reference words) 76.22 % words found, 8.48 % extra 78.08 % found, 8.96 % extra
net errors per reference word 1.39 points fewer, 95 % [+0.93, +1.92] (paired over meetings); more words found in every meeting
four hand-checked German passages (podcast, raw Zoom call), WER 9.28 % 8.75 %, all four better

Whisper ran on the whole recording every time. Passages were cut out afterwards, because a passage cut out first gives the speech map an artificial edge mid-sentence, and Whisper hallucinated into it.

Tried and left out

  • Lowering Silero's threshold. 0.35 is within noise on AMI, 0.25 does worse than 0.35, and 0.15 lets room tone through.
  • No VAD at all. 190 invented words in the pure-noise windows.
  • Silero and the diarization combined. Nothing beyond the diarization alone.
  • Closing gaps up to 2 s instead of 0.5 s. One phone call became a single 31-s chunk, and Whisper decoded it into an invented sentence.
  • Moving adjust_for_pause to the turns as well. It still snaps segment edges to Silero's pauses. Simulated, the speaker found changes for 1-2 of 735 segments, so it stays as it is.

Limits

  • Speech the diarization misses is no longer transcribed, even where Silero would have heard it. None of the corpora above showed such a loss, but none had speech over long stretches of music.
  • Single files vary between runs with either speech map. faster-whisper re-decodes an unsure window by sampling, unseeded, and quiet speech is where it is unsure; one recording gave 219 to 292 words. Only the totals above say something.

It overlaps textually with #351 and #352, which add the same diarization = None line.

🤖 Generated with Claude Code

…ere Silero heard voice activity, so a voice far quieter than the one beside it is transcribed instead of cut away before Whisper hears it

With speaker detection, the diarization's turns replace Silero's speech map
for faster-whisper: only the one call that returns the map is swapped, so
the chunks are collected and the timestamps mapped back exactly as before.
Silero kept 27-41 % of a voice 15-20 dB below its neighbour, and Whisper
wrote almost none of what it dropped. On all 23 AMI meetings the
diarization's map makes 1.39 points fewer errors per reference word
(95 % [+0.93, +1.92]), with more words found in every meeting, and on four
hand-checked German passages WER falls from 9.28 % to 8.75 %. Without
speaker detection nothing changes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@kaixxx

kaixxx commented Sep 25, 2026

Copy link
Copy Markdown
Owner

VAD is very important with whisper to reduce hallucinations when background noise is interpreted as text. Maybe the voice detection in Nemotron can replace that, but this would need careful testing with real world data (including background noise that is not speech).
At the moment, I don't have time to look into this, as I am working on the client-server implementation. This also requires some changes to the internal architecture of the app to better separate frontend (UI) and backend from each other. As a result, our codebases might drift apart, so please slow down a little.... Your Nemotron integration looks very promising, I think this should be the next step once the new transformers version is final.

@happyarts

Copy link
Copy Markdown
Contributor Author

Thanks, Kai, understood. Good luck with the client-server split; I'll slow down.

You're right that the VAD is Whisper's guard against hallucinations. The case I have measured least is exactly the one you name: background that isn't speech. So far I only have 8 minutes of noise-only windows and phone calls with car and street noise, and Whisper invented no words in either. Music, TV and long stretches with nobody talking are not covered.

I'll turn this PR into a draft and not push it further before Nemotron. Once Nemotron replaces pyannote, its turns are what would feed Whisper, so that is the version worth testing. Its turns work as a speech map too: on the 15 AMI test meetings they are level with pyannote's (1.5 points fewer errors per reference word than Silero). Until then I'll collect test material with non-speech background in the background, and count the words invented where nobody speaks.

Where Nemotron stands, as a rough picture from the measurements in #360:

  • Speaker attribution: clearly better than pyannote. AMI test: 11.6 % vs 15.7 % of words under the wrong speaker, and 0.6 % vs 2.6 % in single-speaker speech. German CallHome: 6.8 % vs 11.8 % of time attributed wrongly. With the voice check on Nemotron's probabilities, roughly half the wrong words.
  • Speed and memory: about 7× faster on CPU, and a 4.8 h recording in 2.5–3.5 GB.
  • Level with pyannote: German two-track podcasts, and the speech map on AMI.
  • Open: background without speech, more than 8 speakers (they get merged), and more real-world interviews with difficult conditions.

So far it is better or level wherever it has been measured. What's missing are edge conditions nobody has tested yet, not a known weakness. I'll keep all of this ready for when transformers is final and your restructuring has settled.

No need to look at any of it now.

@happyarts
happyarts marked this pull request as draft September 25, 2026 08:27
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants