Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
29 commits
Select commit Hold shift + click to select a range
088c6fe
Redesign RecognitionSubsystem to async Engine/Session API
Oct 1, 2026
0cfb9db
Redesign SynthesisSubsystem to async Engine/Session API
Oct 1, 2026
cf82e9b
Migrate DemaConsulting.Speech.Cli to async Engine/Session API
Oct 1, 2026
a7580d6
Migrate Demo app to async Engine/Session API and fix synthesis re-cre…
Oct 1, 2026
5d91a19
fix: resolve cancellation-vs-timeout race in SilenceTimeoutRecognizer…
Oct 1, 2026
ecdf5a6
test: rename recog* variables to recognition* for spelling compliance
Oct 1, 2026
b707784
docs(reqstream): fix orphaned requirements and stale test-name refere…
Oct 1, 2026
df904bf
docs: fix MD013 line-length violations and apply fix.ps1 formatting
Oct 1, 2026
d81d3ab
Fix stale doc-comment references to deleted/renamed API surface
Oct 1, 2026
226c722
docs: update ApiMark/NuGet descriptions to reflect async Engine/Sessi…
Oct 1, 2026
b2184df
fix: close concurrency bugs in synthesis/recognition session lifecycle
Oct 1, 2026
70190ff
docs: fix stale async-API references found by formal review across 32…
Oct 1, 2026
27fc279
Close 6 formal-review test-coverage gaps
Oct 1, 2026
9b2a7f9
docs: fix ApiMark generated-doc defects in factory remarks
Oct 1, 2026
d1fbadd
fix: correct stale/missing file references in design doc definition.yaml
Oct 1, 2026
9e42c42
Fix verification-doc build and reqstream orphans from recognition rename
Oct 2, 2026
60d9ca5
docs: correct UnavailableRecognitionSession.State verification docs t…
Oct 2, 2026
a9eb9b2
fix(recognition): close session/engine lifecycle races from PR #40 re…
Oct 2, 2026
35c9c56
fix(synthesis): close session/engine lifecycle races from PR #40 review
Oct 2, 2026
cae762b
fix(demo): make shutdown cleanup complete before shutdown proceeds
Oct 2, 2026
14331de
fix(recognition,synthesis): honor StopAsync cancellation token and ve…
Oct 2, 2026
9dd20da
Fix PR #40 review findings 21-33: event-ordering, backend-reuse, and …
Oct 2, 2026
27a1f25
Fix stale test-name references in reqstream after StopAsync rename
Oct 2, 2026
d65b5f7
fix(recognition): yield before fire-and-forget teardown to prevent de…
Oct 2, 2026
1ff8f69
Fix 4 new PR review findings: async device start, doc corrections
Oct 2, 2026
f6371cd
Fix two more PR review findings: synthesis stop/dispose race and reco…
Oct 2, 2026
6e5e0e1
Fix PR review finding: correct stale ChunkedPipeline requirement wording
Oct 2, 2026
778ed9b
Fix 4 PR bot findings: result ordering, lease-release races, synthesi…
Oct 2, 2026
9cd13c3
Fix finding 30: CancelAndAwaitOperationAsync skipped pending native c…
Oct 2, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
7 changes: 7 additions & 0 deletions .cspell.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -53,6 +53,7 @@ words:
- unsupplied
- denoise
- Denoise
- backpressure
- Downmix
- downmix
- Downmixes
Expand Down Expand Up @@ -153,6 +154,12 @@ words:
- undiscarded
- unreset
- undecoded
- overwritable
- provisionals
- precheck
- unrequested
- unawaited
- unconfigured

# Exclude common build artifacts, dependencies, and vendored third-party code
ignorePaths:
Expand Down
190 changes: 140 additions & 50 deletions .reviewmark.yaml

Large diffs are not rendered by default.

52 changes: 31 additions & 21 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -122,18 +122,29 @@ var captureDevice = new AudioDeviceFactory().CreateCaptureDevice(
AudioDeviceSelection.SystemDefault,
model.AudioFormat);

// 5. Compose the recognizer and stream recognized text as it arrives. Create never throws for an
// ordinary machine state (model not installed, no microphone) - check IsAvailable instead.
using var recognizer = SpeechRecognizerFactory.Create(model, catalog, captureDevice);
if (recognizer.IsAvailable)
// 5. Load the engine once, create a session bound to the capture device, and stream
// recognized text as it arrives. LoadAsync never throws for an ordinary machine state
// (model not installed, no microphone) - check IsAvailable instead.
await using var engine = await SpeechRecognizerFactory.LoadAsync(model, catalog);
if (engine.IsAvailable)
{
recognizer.ResultReceived += (_, args) =>
Console.WriteLine($"{(args.Result.IsFinal ? "final" : "partial")}: {args.Result.Text}");
await using var session = await engine.CreateSessionAsync(captureDevice);

recognizer.Start();
await session.StartAsync();
Console.WriteLine("Listening - press any key to stop...");

var resultsTask = Task.Run(async () =>
{
await foreach (var evt in session.GetResultsAsync())
{
var status = evt.Result.IsFinal ? "final" : "partial";
Console.WriteLine($"{status}: {evt.Result.Text}");
}
});

Console.ReadKey(intercept: true);
recognizer.Stop();
await session.StopAsync();
await resultsTask;
}
```

Expand Down Expand Up @@ -164,16 +175,16 @@ var playbackDevice = new AudioDeviceFactory().CreatePlaybackDevice(
AudioDeviceSelection.SystemDefault,
model.PreferredAudioFormat);

// 5. Compose the synthesizer and speak. Create never throws for an ordinary machine state (model
// not installed, no speakers) - check IsAvailable instead.
using var synthesizer = SpeechSynthesizerFactory.Create(model, catalog, playbackDevice);
if (synthesizer.IsAvailable)
// 5. Load the engine and speak a one-shot phrase. LoadAsync never throws for an ordinary
// machine state (model not installed, no speakers) - check IsAvailable instead.
await using var engine = await SpeechSynthesizerFactory.LoadAsync(model, catalog);
if (engine.IsAvailable)
{
await synthesizer.SpeakAsync("To be, or not to be. [short pause] That is the question.");
await engine.SpeakAsync(playbackDevice, "To be, or not to be. [short pause] That is the question.");
}
```

Both `Create(...)` factories never throw for an ordinary machine state: a model that isn't
Both `LoadAsync(...)` factories never throw for an ordinary machine state: a model that isn't
installed, a machine with no microphone/speakers, and a missing speech-engine native runtime all
return `IsAvailable == false` instead of an exception. `AudioDeviceFactory` also exposes
`RefreshDevices()` to re-scan for hot-plugged hardware, surfacing `AudioDeviceInUseException` if a
Expand All @@ -182,17 +193,16 @@ device from the factory is currently active.
`SpeakAsync` recognizes Natural Language Audio Tags (such as `[whispers]`, `[short pause]`, or
`[excited]`), renders each one per the model's own declared capability, chunks narration into
sentence-sized pieces, and pipelines synthesis with playback - an earlier chunk plays while a
later chunk is still synthesizing. `Stop()` cancels an in-flight `SpeakAsync` call deterministically
and is a safe no-op when nothing is speaking.
later chunk is still synthesizing. Passing a cancelled `CancellationToken` to `SpeakAsync` cancels
an in-flight call deterministically.

For a model that declares tunable parameters - such as Kokoro's `voice` choice or VITS/Piper's
numeric `speaker` id - pass a `parameterValues` bag keyed by each parameter's `Id`:
numeric `speaker` id - pass a `parameterValues` bag keyed by each parameter's `Id` to `LoadAsync(...)`:

```csharp
using var synthesizer = SpeechSynthesizerFactory.Create(
await using var engine = await SpeechSynthesizerFactory.LoadAsync(
model,
catalog,
playbackDevice,
parameterValues: new Dictionary<string, object> { ["voice"] = "af_bella" });
```

Expand All @@ -203,8 +213,8 @@ different models without breaking composition. A supplied value for a parameter
declare, but that fails that parameter's own validation - the wrong CLR type, a number outside
its declared range, a fractional value for a whole-number-only parameter, or a string that
matches none of a `ChoiceParameter`'s declared options - throws `ArgumentException` synchronously
from `Create()`, naming the parameter, the model, and the reason the value is invalid. This same
rule applies to `SpeechRecognizerFactory.Create`'s `parameterValues` argument.
from `LoadAsync(...)`, naming the parameter, the model, and the reason the value is invalid. This
same rule applies to `SpeechRecognizerFactory.LoadAsync`'s `parameterValues` argument.

See the [user guide][link-user-guide] for the full API walkthrough, voice/speaker catalogs, and
Natural Language Audio Tag vocabulary.
Expand Down
17 changes: 14 additions & 3 deletions docs/design/definition.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -41,10 +41,21 @@ input-files:
- docs/design/speech/model-management-subsystem/speech-model-descriptor.md
- docs/design/speech/model-management-subsystem/speech-model-catalog.md
- docs/design/speech/recognition-subsystem.md
- docs/design/speech/recognition-subsystem/i-speech-recognizer.md
- docs/design/speech/recognition-subsystem/i-speech-recognizer-engine.md
- docs/design/speech/recognition-subsystem/i-recognition-session.md
- docs/design/speech/recognition-subsystem/speech-recognizer-factory.md
- docs/design/speech/recognition-subsystem/sherpa-onnx-speech-recognizer.md
- docs/design/speech/recognition-subsystem/unavailable-speech-recognizer.md
- docs/design/speech/recognition-subsystem/sherpa-onnx-recognition-session.md
- docs/design/speech/recognition-subsystem/unavailable-speech-recognizer-engine.md
- docs/design/speech/recognition-subsystem/unavailable-recognition-session.md
- docs/design/speech/synthesis-subsystem.md
- docs/design/speech/synthesis-subsystem/i-speech-synthesizer-engine.md
- docs/design/speech/synthesis-subsystem/i-synthesis-session.md
- docs/design/speech/synthesis-subsystem/speech-synthesizer-factory.md
- docs/design/speech/synthesis-subsystem/sherpa-onnx-speech-synthesizer-engine.md
- docs/design/speech/synthesis-subsystem/sherpa-onnx-synthesis-session.md
- docs/design/speech/synthesis-subsystem/unavailable-speech-synthesizer-engine.md
- docs/design/speech/synthesis-subsystem/unavailable-synthesis-session.md
- docs/design/speech/synthesis-subsystem/audio-tag-parser.md
- docs/design/speech-demo.md
- docs/design/speech-demo/shell-subsystem.md
- docs/design/speech-demo/device-selection-subsystem.md
Expand Down
Loading
Loading