Let Whisper listen where the diarization found speakers, so a quiet second voice is transcribed instead of cut away by the VAD - #361
Conversation
…ere Silero heard voice activity, so a voice far quieter than the one beside it is transcribed instead of cut away before Whisper hears it With speaker detection, the diarization's turns replace Silero's speech map for faster-whisper: only the one call that returns the map is swapped, so the chunks are collected and the timestamps mapped back exactly as before. Silero kept 27-41 % of a voice 15-20 dB below its neighbour, and Whisper wrote almost none of what it dropped. On all 23 AMI meetings the diarization's map makes 1.39 points fewer errors per reference word (95 % [+0.93, +1.92]), with more words found in every meeting, and on four hand-checked German passages WER falls from 9.28 % to 8.75 %. Without speaker detection nothing changes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
VAD is very important with whisper to reduce hallucinations when background noise is interpreted as text. Maybe the voice detection in Nemotron can replace that, but this would need careful testing with real world data (including background noise that is not speech). |
|
Thanks, Kai, understood. Good luck with the client-server split; I'll slow down. You're right that the VAD is Whisper's guard against hallucinations. The case I have measured least is exactly the one you name: background that isn't speech. So far I only have 8 minutes of noise-only windows and phone calls with car and street noise, and Whisper invented no words in either. Music, TV and long stretches with nobody talking are not covered. I'll turn this PR into a draft and not push it further before Nemotron. Once Nemotron replaces pyannote, its turns are what would feed Whisper, so that is the version worth testing. Its turns work as a speech map too: on the 15 AMI test meetings they are level with pyannote's (1.5 points fewer errors per reference word than Silero). Until then I'll collect test material with non-speech background in the background, and count the words invented where nobody speaks. Where Nemotron stands, as a rough picture from the measurements in #360:
So far it is better or level wherever it has been measured. What's missing are edge conditions nobody has tested yet, not a known weakness. I'll keep all of this ready for when transformers is final and your restructuring has settled. No need to look at any of it now. |
With speaker detection on, Whisper now takes its speech map from the diarization instead of from Silero. A voice much quieter than the one beside it — an interviewer far from the microphone, a second person in the room — used to be cut away by the VAD before Whisper ever heard it. Now it is transcribed. Without speaker detection nothing changes.
What changes
noScribe/whisper_mp_worker.pyspeech_map(turns, n_samples)turns the diarization's[start, end]turns into the chunks faster-whisper expects. Each turn gets 50 ms of padding, gaps under 0.5 s are closed, and the result is clipped to the audio. These are the same amounts as the Silero options a few lines below._speech_map_from(turns)replacesfaster_whisper.transcribe.get_speech_timestamps, the one call through which faster-whisper gets its speech map, for the duration ofdetect_language()andtranscribe()only. It always restores the original, also on an exception. Everything else in faster-whisper runs exactly as with Silero:collect_chunksand the mapping of timestamps back to the recording.noScribe/main.py_run_whisper_subprocess_stream(..., diarization=None)passes the turns to the worker asspeech_turns, in seconds.diarization = Nonebefore the diarization block, so the call also works with speaker detection off.tests/test_whisper_speech_map.py: 13 tests, all without a model:test_whisper_audio_lifetime.py);_process_single_job, and areNonewithout speaker detection.Why: measured
Silero drops quiet speech. Measured on the Krisp Voice Isolation benchmark set: real recordings, a primary speaker plus a second voice 15-20 dB lower in the same microphone, with windows where only the second voice speaks.
Against manual transcripts:
Whisper ran on the whole recording every time. Passages were cut out afterwards, because a passage cut out first gives the speech map an artificial edge mid-sentence, and Whisper hallucinated into it.
Tried and left out
adjust_for_pauseto the turns as well. It still snaps segment edges to Silero's pauses. Simulated, the speaker found changes for 1-2 of 735 segments, so it stays as it is.Limits
It overlaps textually with #351 and #352, which add the same
diarization = Noneline.🤖 Generated with Claude Code