Skip to content

Add Voxtral as an optional second transcription engine (Apple Silicon) - #352

Open
happyarts wants to merge 5 commits into
kaixxx:mainfrom
happyarts:feature/voxtral-engine
Open

happyarts wants to merge 5 commits into
kaixxx:mainfrom
happyarts:feature/voxtral-engine

Conversation

@happyarts

@happyarts happyarts commented Sep 20, 2026 •

Copy link
Copy Markdown
Contributor

This adds Mistral's Voxtral as a second transcription engine next to
faster-whisper, running on the Apple Silicon GPU through MLX. It is opt-in and
gated three times over — platform, installed packages, selected model — so
nothing changes for anyone who does not ask for it.

No hurry with this one. I know a release comes first and a feature this size
does not belong in front of it. I am opening it now because it has been in
regular use on my fork since July and has been stable for weeks, and because it
reads best next to #351: Voxtral is where the voice check earns the most (see
Relation to #351 below), so the two can be looked at together whenever it
suits you.

Why Voxtral is worth a look

Voxtral is not "a better Whisper" — on clean read-aloud audio it is mid-table.
What it is good at is exactly what noScribe is for: real conversation.
Measured on a hand-corrected two-minute passage of German interview audio
picked for being hard (crosstalk, brand names, a dialect speaker), with
noScribe's own Whisper settings:

Model WER CER Sub Del Ins Speed
voxtral-mini-8bit (3B) 4.27 % 3.39 % 10 8 0 6.8x
whisper fast 8.06 % 3.34 % 22 4 8 2.4x
whisper precise 14.22 % 8.75 % 22 4 34 2.7x
  • Half the word errors of fast, at almost three times the speed (measured
    on an M1 Max; the 3B build fits a 16 GB Mac). The character error is level
    with fast — what Voxtral adds is
    consistent spelling and sentences that read as sentences (denser, more
    sensible punctuation), which is what costs the time when proof-reading.
  • It omits where Whisper invents. Zero insertions against 8 and 34: on this
    passage precise ended in 27 consecutive hallucinated words, including a
    product that does not exist. An omission is visible when you proof-read
    against the audio; a fluent hallucination is not. For research transcripts
    that difference matters more than the headline rate.
  • This does not say fast beats precise. Both Whisper models make the
    same 22 substitutions and 4 deletions here; the whole gap is that one
    hallucination at the end of the clip, without which precise would sit at
    about 7.8 %. On FLEURS the two produce identical output. The precise row
    shows a failure mode on one passage, not a rate.
  • It does not win everywhere, and the write-up says so. On FLEURS German
    (public, read-aloud) both Whisper models score 4.05 % WER and the 3B Voxtral
    4.81 %. docs/other-asr-engines.md measures five further engines (Parakeet,
    Qwen3-ASR, Cohere, VibeVoice, transcribe.cpp) on the same three yardsticks;
    the recurring finding is that a model can top the read-aloud benchmark and be
    unusable for interviews, and none of them has displaced this build.

Two hand-corrected references come from private recordings and cannot be
shipped; the FLEURS numbers reproduce with docs/scripts/fleurs.py. All names
in the write-ups are fictional stand-ins.

Mac-only — and why I think it is still in good hands here

The engine needs MLX, so it runs on arm64 Macs only. That is a real limitation,
and it is why everything about it is built so the other platforms never notice:

  • No new dependency for anyone. The MLX stack lives in its own file,
    environments/requirements_voxtral_macOS_arm64.txt, layered on the macOS
    requirements. No existing requirements file gains a package.
  • No new import for anyone. Every Voxtral import is function-local and
    behind engine == "voxtral". Models are registered only on Darwin/arm64 and
    only if is_available() (a find_spec, no import) finds the packages. A Mac
    without the stack sees no models and no warnings.
  • Frozen builds stay as they are. The macOS spec excludes the MLX stack,
    so a packaged app never carries it, whatever the build machine has installed;
    there the feature is simply absent. Nothing to build, sign or ship differently.
  • CI stays green on Linux. All Voxtral tests either run without the stack
    or skip through pytest.importorskip. Checked locally with mlx,
    mlx_voxtral, mlx_lm, transformers and torchaudio blocked at import:
    386 passed, 11 skipped, no errors. With the stack: 403 passed, 3 skipped.
  • The seam is general, not Voxtral-shaped. WhisperModel gains an engine
    field, and the worker speaks the same queue messages as
    whisper_mp_worker.py, so main.py consumes segments identically. Whatever
    engine comes next — on any platform — plugs into the same place.
    (docs/other-asr-engines.md already measures one route to this same model
    off the Mac, transcribe.cpp, and finds it on par.)
  • The maintenance risk is named and has an exit. The coupling to
    mlx-voxtral is five library calls plus a short list of attribute reaches,
    all listed in VOXTRAL.md; versions are pinned and tests/test_voxtral_pin.py
    checks them. The library's author merged every fix this work reported and
    relicensed it to MIT when asked. Should it ever go quiet,
    docs/migration-mlx-audio.md is a verified work order for the move to
    mlx-audio.

For everyone on an Apple Silicon Mac this is a markedly better transcript on
the hardware they already own, fully offline — which is the point of noScribe.
Whether that justifies a platform-specific engine in the main tree is of course
your call — I am happy to keep carrying it on the fork if not, and equally
happy to maintain it here.

What is in it

Five commits, meant to be read in order:

Commit Contents Size
Engine voxtral_engine.py, voxtral_mp_worker.py, ctc_align.py, transcript_corrections.py, the requirements file ~4 100 lines
Integration main.py, transcription.py, four new UI strings in all nine languages, one .gitignore line, the MLX stack in the macOS spec's excludes +280 / −52 in main.py
Tests 23 test files, 311 tests, one reference file (forced_align_ref.npz, 0.9 MB) ~4 800 lines
Tools quantize_voxtral.py, calibrate_loop_detection.py ~500 lines
Docs VOXTRAL.md, seven write-ups in docs/, the scripts behind them ~8 000 lines

The part that touches existing behaviour is the second commit, and it is
small.
In main.py:

  • _run_whisper_subprocess_stream is split, not rewritten: it still builds the
    same arguments, and its message pump moves verbatim into
    _run_engine_subprocess_stream, which both engines share. It still returns
    the _Info object, and Voxtral's result carries duration too, so whatever
    you plan to read from it works for both.
  • The model picker shows Voxtral builds with the RAM they need
    (voxtral-mini-8bit · 13 GB RAM, with a ⚠ if the machine has less), because
    a model that does not fit does not fail — it swaps, and looks like a hang.
    Whisper entries are unchanged; model_key() maps the label back, so the
    config still stores the plain name.
  • The CUDA-to-CPU retry is now explicitly Whisper's; a Voxtral error propagates
    instead of offering a pointless "retry Whisper on CPU".
  • Worker log messages start on a fresh line instead of being glued to the
    transcript text streaming in, and the log header names the model used. Both
    apply to Whisper as well; both are one-liners.
  • Per-segment progress stays Whisper's source of progress. Voxtral sends its
    own progress messages, so its runs skip the per-segment estimate.

transcription.py: WhisperModel gains engine and repo, and the model
scanner no longer warns about a directory that holds safetensors instead of
model.bin (it still warns about a directory that holds neither).

The engine itself is large because Voxtral out of the box is not usable
for hour-long interviews, and most of the file is what makes it so:

  • pass length sized from installed RAM, long files cut at pauses;
  • repetition loops detected and repaired — a temperature ladder, then a split
    (at a speaker turn where a diarization exists), with the good prefix of a
    failed pass salvaged rather than redone;
  • guards against a pass that comes back translated, and recovery of an opening
    the model dropped;
  • word timestamps by CTC forced alignment: wav2vec2 emissions through
    transformers, the Viterbi in numpy (ctc_align.py). torchaudio's kernel is
    deliberately not used — it can crash on long inputs (pytorch/audio issue
    4208) and resolves exact ties to the worse path, which this work reported
    upstream (issue 4221 there, with a fix).
    tests/test_forced_align_stability.py pins the numpy DP to a reference
    recorded from torchaudio;
  • the aligner's input normalised per window, as its models were trained (the
    feature extractor that would do it is bypassed), and a pass whose words leave
    diarized speech empty aligned a second time, level-matched: a voice 15-20 dB
    below its neighbour had its words stamped 10-20 s early, into the other
    speaker's time. Realigning the same text: words more than 0.5 s off AMI's
    manual stamps 7.1 % → 3.1 %; words under the wrong speaker after the voice
    check of Check each passage's speaker against the voice itself, which halves the words written under the wrong speaker #351 −0.35 points over 105 AMI, CallHome and VoxConverse files, and
    on a second-voice corpus a further −1.04 points in call-centre audio, checked
    on recordings that took no part in the choice;
  • the log-Mel clamp floor taken from a percentile instead of the maximum, so
    one door slam no longer degrades a whole pass (measured +1.52 WER points at
    28 dB; free on clean audio). Done by swapping the processor's feature
    extractor, so the library stays unpatched;
  • a user-maintained find/replace list for brand and product names, since
    Voxtral has no hotword hook, plus the spelling of the entered speaker names:
    a word that sounds exactly like one of them and is spelled another way
    ("Mohna", "Steffy") takes the name's spelling, but only when the macOS
    dictionary of the transcript's language does not know it. Over 407k words of
    transcripts checked against 70-150 common first names per language, that
    changed one word, rightly.

Every tuning constant carries the measurement behind it in the comment above
it, and most tests are regression guards for a defect found on a real
recording; their docstrings say which.

About docs/

I kept the write-ups and scripts in, because the code comments cite them and a
number without its method is an assertion. They are self-contained and nothing
imports them, so if you would rather not carry ~8 000 lines of measurement
record in the main tree, say so and I will cut the PR down to VOXTRAL.md and
move the rest to the fork, with links. Note that docs/scripts/ needs
git add -f: .gitignore has a scripts/ entry.

Relation to #351 and #350

This branch is cut from main and contains neither, so the diff is Voxtral
only. It overlaps with both in main.py textually; whichever merges first, I
will rebase the rest — the combined state is what my fork has been running.

In practice Voxtral should go in after #351. The engine merges one-word
cues into their neighbour so a sentence is not torn apart by a stray label,
which can in turn put a short interjection under the wrong speaker; the voice
check is what repairs that. It works without #351, but it is better with it.

How to try it

On an Apple Silicon Mac, from source, on top of the usual
requirements_macOS_arm64.txt install:

pip install -r environments/requirements_voxtral_macOS_arm64.txt
python -m noScribe

voxtral-mini-8bit appears in the model dropdown and downloads (~6 GB) on
first use. Headless works too:
python -m noScribe audio.wav out.html --no-gui --model voxtral-mini-8bit.
On a 32 GB M1 Max a one-minute clip with speaker detection and word timestamps
takes about 25 s end to end, model load included (models already downloaded).

🤖 Generated with Claude Code

@happyarts
happyarts force-pushed the feature/voxtral-engine branch from cb9d428 to 0bca7bd Compare September 23, 2026 13:18
happyarts and others added 5 commits September 24, 2026 13:31
…hat speaks the same queue protocol as the Whisper one

voxtral_engine.py decodes with mlx-voxtral on the GPU and owns everything around it: pass sizing from installed RAM, pause-aligned cuts, repetition-loop detection with a temperature ladder and prefix salvage, language guards, recovery of a dropped opening, the percentile log-Mel clamp floor, and word timestamps by CTC forced alignment (emissions from transformers, the Viterbi in numpy in ctc_align.py). A model that is not downloaded and cannot be fetched says so, decided by asking the hub. transcript_corrections.py is the user-maintained find/replace list that stands in for the hotword hook Voxtral does not have, and it sets a speaker name's spelling where the model wrote one that sounds the same and no dictionary knows. Nothing here is imported unless a Voxtral model is selected.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…the stack installed, and run both engines through one subprocess pump

WhisperModel gains engine and repo; main.py registers the builds only where voxtral_engine.is_available(), labels them with the RAM they need, and dispatches on the engine. The Whisper path is the same code, moved: _run_whisper_subprocess_stream now builds its arguments and hands them to the shared pump, which still returns the info object. The CUDA-to-CPU retry stays Whisper's, since Voxtral has no CUDA path to fall back from. A build saved under its earlier name means the one that replaced it, and the macOS spec keeps the MLX stack out of the packaged app, where a partial copy would let the picker offer Voxtral that then fails to load.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ed, all of which run or skip cleanly without the MLX stack

Most of these pin a failure found on real recordings: loop breaking, forced-alignment caps and density, a lost opening, prefix salvage, translated passes, the log-Mel floor. test_forced_align_stability.py holds the numpy Viterbi to a reference recorded from torchaudio (tests/data/forced_align_ref.npz), and test_worker_import_lightweight.py now also keeps the Voxtral worker stdlib-only at import. test_voxtral_transcribe_guards.py drives transcribe() with the model and the aligners stubbed out, for what lies between the pieces: the language that reaches the model, an evicted aligner, an empty file, a model that cannot be fetched.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ve the loop-detection thresholds from logs

quantize_voxtral.py produces the shipped layout (weights quantised, audio encoder and projector left in bf16); calibrate_loop_detection.py recomputes the thresholds the engine's comments cite.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…nd its decisions together with the scripts that made them

VOXTRAL.md is the user documentation. docs/ holds what the code comments cannot carry: which build ships and why, the comparison with six other ASR engines, the log-Mel clamp floor, the audio path, the numpy Viterbi, and a conditional brief for moving to mlx-audio should mlx-voxtral ever go quiet.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant