Conversation
happyarts
force-pushed
the
feature/voxtral-engine
branch
from
September 23, 2026 13:18
cb9d428 to
0bca7bd
Compare
…hat speaks the same queue protocol as the Whisper one voxtral_engine.py decodes with mlx-voxtral on the GPU and owns everything around it: pass sizing from installed RAM, pause-aligned cuts, repetition-loop detection with a temperature ladder and prefix salvage, language guards, recovery of a dropped opening, the percentile log-Mel clamp floor, and word timestamps by CTC forced alignment (emissions from transformers, the Viterbi in numpy in ctc_align.py). A model that is not downloaded and cannot be fetched says so, decided by asking the hub. transcript_corrections.py is the user-maintained find/replace list that stands in for the hotword hook Voxtral does not have, and it sets a speaker name's spelling where the model wrote one that sounds the same and no dictionary knows. Nothing here is imported unless a Voxtral model is selected. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…the stack installed, and run both engines through one subprocess pump WhisperModel gains engine and repo; main.py registers the builds only where voxtral_engine.is_available(), labels them with the RAM they need, and dispatches on the engine. The Whisper path is the same code, moved: _run_whisper_subprocess_stream now builds its arguments and hands them to the shared pump, which still returns the info object. The CUDA-to-CPU retry stays Whisper's, since Voxtral has no CUDA path to fall back from. A build saved under its earlier name means the one that replaced it, and the macOS spec keeps the MLX stack out of the packaged app, where a partial copy would let the picker offer Voxtral that then fails to load. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ed, all of which run or skip cleanly without the MLX stack Most of these pin a failure found on real recordings: loop breaking, forced-alignment caps and density, a lost opening, prefix salvage, translated passes, the log-Mel floor. test_forced_align_stability.py holds the numpy Viterbi to a reference recorded from torchaudio (tests/data/forced_align_ref.npz), and test_worker_import_lightweight.py now also keeps the Voxtral worker stdlib-only at import. test_voxtral_transcribe_guards.py drives transcribe() with the model and the aligners stubbed out, for what lies between the pieces: the language that reaches the model, an evicted aligner, an empty file, a model that cannot be fetched. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ve the loop-detection thresholds from logs quantize_voxtral.py produces the shipped layout (weights quantised, audio encoder and projector left in bf16); calibrate_loop_detection.py recomputes the thresholds the engine's comments cite. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…nd its decisions together with the scripts that made them VOXTRAL.md is the user documentation. docs/ holds what the code comments cannot carry: which build ships and why, the comparison with six other ASR engines, the log-Mel clamp floor, the audio path, the numpy Viterbi, and a conditional brief for moving to mlx-audio should mlx-voxtral ever go quiet. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
happyarts
force-pushed
the
feature/voxtral-engine
branch
from
September 24, 2026 11:39
0bca7bd to
2de31c8
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This adds Mistral's Voxtral as a second transcription engine next to
faster-whisper, running on the Apple Silicon GPU through MLX. It is opt-in and
gated three times over — platform, installed packages, selected model — so
nothing changes for anyone who does not ask for it.
No hurry with this one. I know a release comes first and a feature this size
does not belong in front of it. I am opening it now because it has been in
regular use on my fork since July and has been stable for weeks, and because it
reads best next to #351: Voxtral is where the voice check earns the most (see
Relation to #351 below), so the two can be looked at together whenever it
suits you.
Why Voxtral is worth a look
Voxtral is not "a better Whisper" — on clean read-aloud audio it is mid-table.
What it is good at is exactly what noScribe is for: real conversation.
Measured on a hand-corrected two-minute passage of German interview audio
picked for being hard (crosstalk, brand names, a dialect speaker), with
noScribe's own Whisper settings:
fastprecisefast, at almost three times the speed (measuredon an M1 Max; the 3B build fits a 16 GB Mac). The character error is level
with
fast— what Voxtral adds isconsistent spelling and sentences that read as sentences (denser, more
sensible punctuation), which is what costs the time when proof-reading.
passage
preciseended in 27 consecutive hallucinated words, including aproduct that does not exist. An omission is visible when you proof-read
against the audio; a fluent hallucination is not. For research transcripts
that difference matters more than the headline rate.
fastbeatsprecise. Both Whisper models make thesame 22 substitutions and 4 deletions here; the whole gap is that one
hallucination at the end of the clip, without which
precisewould sit atabout 7.8 %. On FLEURS the two produce identical output. The
preciserowshows a failure mode on one passage, not a rate.
(public, read-aloud) both Whisper models score 4.05 % WER and the 3B Voxtral
4.81 %.
docs/other-asr-engines.mdmeasures five further engines (Parakeet,Qwen3-ASR, Cohere, VibeVoice, transcribe.cpp) on the same three yardsticks;
the recurring finding is that a model can top the read-aloud benchmark and be
unusable for interviews, and none of them has displaced this build.
Two hand-corrected references come from private recordings and cannot be
shipped; the FLEURS numbers reproduce with
docs/scripts/fleurs.py. All namesin the write-ups are fictional stand-ins.
Mac-only — and why I think it is still in good hands here
The engine needs MLX, so it runs on arm64 Macs only. That is a real limitation,
and it is why everything about it is built so the other platforms never notice:
environments/requirements_voxtral_macOS_arm64.txt, layered on the macOSrequirements. No existing requirements file gains a package.
behind
engine == "voxtral". Models are registered only on Darwin/arm64 andonly if
is_available()(afind_spec, no import) finds the packages. A Macwithout the stack sees no models and no warnings.
so a packaged app never carries it, whatever the build machine has installed;
there the feature is simply absent. Nothing to build, sign or ship differently.
or skip through
pytest.importorskip. Checked locally withmlx,mlx_voxtral,mlx_lm,transformersandtorchaudioblocked at import:386 passed, 11 skipped, no errors. With the stack: 403 passed, 3 skipped.
WhisperModelgains anenginefield, and the worker speaks the same queue messages as
whisper_mp_worker.py, somain.pyconsumes segments identically. Whateverengine comes next — on any platform — plugs into the same place.
(
docs/other-asr-engines.mdalready measures one route to this same modeloff the Mac, transcribe.cpp, and finds it on par.)
mlx-voxtralis five library calls plus a short list of attribute reaches,all listed in
VOXTRAL.md; versions are pinned andtests/test_voxtral_pin.pychecks them. The library's author merged every fix this work reported and
relicensed it to MIT when asked. Should it ever go quiet,
docs/migration-mlx-audio.mdis a verified work order for the move tomlx-audio.For everyone on an Apple Silicon Mac this is a markedly better transcript on
the hardware they already own, fully offline — which is the point of noScribe.
Whether that justifies a platform-specific engine in the main tree is of course
your call — I am happy to keep carrying it on the fork if not, and equally
happy to maintain it here.
What is in it
Five commits, meant to be read in order:
voxtral_engine.py,voxtral_mp_worker.py,ctc_align.py,transcript_corrections.py, the requirements filemain.py,transcription.py, four new UI strings in all nine languages, one.gitignoreline, the MLX stack in the macOS spec'sexcludesmain.pyforced_align_ref.npz, 0.9 MB)quantize_voxtral.py,calibrate_loop_detection.pyVOXTRAL.md, seven write-ups indocs/, the scripts behind themThe part that touches existing behaviour is the second commit, and it is
small. In
main.py:_run_whisper_subprocess_streamis split, not rewritten: it still builds thesame arguments, and its message pump moves verbatim into
_run_engine_subprocess_stream, which both engines share. It still returnsthe
_Infoobject, and Voxtral's result carriesdurationtoo, so whateveryou plan to read from it works for both.
(
voxtral-mini-8bit · 13 GB RAM, with a ⚠ if the machine has less), becausea model that does not fit does not fail — it swaps, and looks like a hang.
Whisper entries are unchanged;
model_key()maps the label back, so theconfig still stores the plain name.
instead of offering a pointless "retry Whisper on CPU".
transcript text streaming in, and the log header names the model used. Both
apply to Whisper as well; both are one-liners.
own
progressmessages, so its runs skip the per-segment estimate.transcription.py:WhisperModelgainsengineandrepo, and the modelscanner no longer warns about a directory that holds safetensors instead of
model.bin(it still warns about a directory that holds neither).The engine itself is large because Voxtral out of the box is not usable
for hour-long interviews, and most of the file is what makes it so:
(at a speaker turn where a diarization exists), with the good prefix of a
failed pass salvaged rather than redone;
the model dropped;
transformers, the Viterbi in numpy (
ctc_align.py). torchaudio's kernel isdeliberately not used — it can crash on long inputs (
pytorch/audioissue4208) and resolves exact ties to the worse path, which this work reported
upstream (issue 4221 there, with a fix).
tests/test_forced_align_stability.pypins the numpy DP to a referencerecorded from torchaudio;
feature extractor that would do it is bypassed), and a pass whose words leave
diarized speech empty aligned a second time, level-matched: a voice 15-20 dB
below its neighbour had its words stamped 10-20 s early, into the other
speaker's time. Realigning the same text: words more than 0.5 s off AMI's
manual stamps 7.1 % → 3.1 %; words under the wrong speaker after the voice
check of Check each passage's speaker against the voice itself, which halves the words written under the wrong speaker #351 −0.35 points over 105 AMI, CallHome and VoxConverse files, and
on a second-voice corpus a further −1.04 points in call-centre audio, checked
on recordings that took no part in the choice;
one door slam no longer degrades a whole pass (measured +1.52 WER points at
28 dB; free on clean audio). Done by swapping the processor's feature
extractor, so the library stays unpatched;
Voxtral has no hotword hook, plus the spelling of the entered speaker names:
a word that sounds exactly like one of them and is spelled another way
("Mohna", "Steffy") takes the name's spelling, but only when the macOS
dictionary of the transcript's language does not know it. Over 407k words of
transcripts checked against 70-150 common first names per language, that
changed one word, rightly.
Every tuning constant carries the measurement behind it in the comment above
it, and most tests are regression guards for a defect found on a real
recording; their docstrings say which.
About
docs/I kept the write-ups and scripts in, because the code comments cite them and a
number without its method is an assertion. They are self-contained and nothing
imports them, so if you would rather not carry ~8 000 lines of measurement
record in the main tree, say so and I will cut the PR down to
VOXTRAL.mdandmove the rest to the fork, with links. Note that
docs/scripts/needsgit add -f:.gitignorehas ascripts/entry.Relation to #351 and #350
This branch is cut from
mainand contains neither, so the diff is Voxtralonly. It overlaps with both in
main.pytextually; whichever merges first, Iwill rebase the rest — the combined state is what my fork has been running.
In practice Voxtral should go in after #351. The engine merges one-word
cues into their neighbour so a sentence is not torn apart by a stray label,
which can in turn put a short interjection under the wrong speaker; the voice
check is what repairs that. It works without #351, but it is better with it.
How to try it
On an Apple Silicon Mac, from source, on top of the usual
requirements_macOS_arm64.txtinstall:voxtral-mini-8bitappears in the model dropdown and downloads (~6 GB) onfirst use. Headless works too:
python -m noScribe audio.wav out.html --no-gui --model voxtral-mini-8bit.On a 32 GB M1 Max a one-minute clip with speaker detection and word timestamps
takes about 25 s end to end, model load included (models already downloaded).
🤖 Generated with Claude Code