You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Redesign RecognitionSubsystem/SynthesisSubsystem to async Engine/Session API - #40
Redesigns the public recognition and synthesis APIs around a consistent, fully
async Engine/Session model: ISpeechRecognizerEngine/ISpeechSynthesizerEngine
load models and create sessions asynchronously; sessions stream results as IAsyncEnumerable/events, and are explicitly cancelled (Cancel()) and
disposed (DisposeAsync()) rather than relying on synchronous disposal of
blocking native calls. Native/blocking work is moved off the caller's thread
(worker threads / long-running Task.Run hints) so the public API surface is
async end-to-end, while the underlying native libraries remain blocking
internally. One-active-session-per-engine cardinality is enforced via RecognitionEngineBusyException, and the existing "honest unavailable, no
throw" convention is preserved for LoadAsync/CreateSessionAsync failure
paths.
This is a deliberate 0.x.y breaking API change, approved by the project
owner, to avoid the library's public shape ossifying around a blocking-first
design that the rest of the .NET ecosystem has moved away from.
Includes:
Core library redesign (RecognitionSubsystem, SynthesisSubsystem) to the
async Engine/Session API, including internal IRecognitionBackend/ ISynthesisBackend renames to free up the Engine name for the new
public types.
Migration of DemaConsulting.Speech.Cli and the Demo app to the new API,
including a fix for a synthesis session recreate-per-click bug uncovered
during migration.
Concurrency-bug fixes in session lifecycle (cancellation-vs-timeout race in SilenceTimeoutRecognizerSession, double-dispose/busy-session edge cases).
Full documentation pass: ApiMark/NuGet package descriptions, design docs, and
stale doc-comment references updated to describe the new async API so
downstream coding agents are guided toward the new surface.
Reqstream/requirements updates (fixed orphaned requirements, stale test-name
references, line-length/formatting compliance).
Bug fix (non-breaking change which fixes an issue)
New feature (non-breaking change which adds functionality)
Breaking change (fix or feature that would cause existing functionality to not work as expected)
Documentation update
Code quality improvement
Related Issues
Closes #
Pre-Submission Checklist
Build and Test
Code builds successfully and all tests pass: pwsh ./build.ps1 (2655/2655 passed)
Code produces zero warnings
Code Quality
New code has appropriate XML documentation comments
Static analyzer warnings have been addressed
Quality Checks
All linters pass: pwsh ./lint.ps1
Testing
Added unit tests for new functionality
Updated existing tests if behavior changed
All tests follow the AAA (Arrange, Act, Assert) pattern
Test coverage is maintained or improved
Documentation
Updated README.md (if applicable)
Updated docs/ documentation (if applicable)
Added code examples for new features (if applicable)
Updated requirements.yaml (if applicable)
Additional Notes
Three independent review passes (code-review agent + formal-review agent across
all affected review-sets) were run after implementation and all findings were
addressed, including this final test-coverage-gap closure commit. This is a
0.x.y breaking change; downstream consumers will need to migrate to the new
Engine/Session API (documented in ApiMark/design docs).
Replace sync ISpeechRecognizer with ISpeechRecognizerEngine (Layer 3, loaded
model) and IRecognitionSession (Layer 5, bound to one device). Rename
internal IRecognitionEngine/IRecognitionEngineFactory to IRecognitionBackend/
IRecognitionBackendFactory. SpeechRecognizerFactory.Create -> LoadAsync.
Adds state machine, engine exclusivity lease, GetResultsAsync backpressure
buffer, and dedicated-worker cancel/abandon policy per the approved plan.
Replace ISpeechSynthesizer with ISpeechSynthesizerEngine (Layer 3) and
ISynthesisSession (Layer 5). Rename internal ISynthesisEngine/
ISynthesisEngineFactory to ISynthesisBackend/ISynthesisBackendFactory.
SpeechSynthesizerFactory.Create -> LoadAsync. Session-scoped SpeakAsync/
SynthesizeAsync with an overlap guard, plus engine-level one-shot
convenience overloads. Fixes stale ISynthesisEngine cref in ISynthesisModel.
RecognizeCommand/SpeakCommand/AskCommand and the model-catalog adapter now
load engines via LoadAsync and drive sessions via CreateSessionAsync/
StartAsync/StopAsync/GetResultsAsync/SpeakAsync. SilenceTimeoutRecognizerSession
redesigned as a stateless IAsyncEnumerable<SpeechRecognitionEvent> decorator
racing idle-timeout against MoveNextAsync, replacing the old lock/wait-handle
machinery built around the sync ISpeechRecognizer.ResultReceived event.
…ate bug
RecognitionPanelViewModel/SynthesisPanelViewModel now use async StartAsync/
StopAsync commands, cache the loaded engine, and drive RecognitionStreamingState/
SynthesisPlaybackState from session StateChanged events instead of ad hoc
assignment. SynthesisPanelViewModel caches its ISpeechSynthesizerEngine/
ISynthesisSession across Play calls (session-scoped parameters), fixing the
known bug where a new synthesizer was constructed on every Play. Cheap device
preconditions are now checked before the async engine load so an honest
'no device' outcome never depends on what the engine-loading seam returns.
…Session
GetResultsAsync raced Task.Delay against the inner session's
MoveNextAsync via Task.WhenAny, but only checked which task won
without checking whether the delay actually completed successfully.
When the enumeration's own cancellation token was cancelled at the
same moment the delay task completed (in a Canceled state, not
RanToCompletion), the code incorrectly treated it as a genuine idle
timeout and called StopAsync(), which gracefully completed the
wrapped session's channel and silently swallowed the expected
OperationCanceledException.
Fix: require delayTask.IsCompletedSuccessfully before treating the
race as a genuine timeout.
Renamed recogSession/recogEngine to recognitionSession/recognitionEngine
throughout AskCommandTests.cs (36 occurrences) to resolve a cspell
'Unknown word: recog' lint failure without polluting the spell-check
dictionary with an informal abbreviation. No functional change.
…nces
- Link Speech-Recognition-EngineExclusivity into
Speech-Recognition-StreamingRecognizer's children.
- Link Speech-Synthesis-EngineExclusivity and
Speech-Synthesis-SessionOverlapAndLifecycle into
Speech-Synthesis-StreamingSynthesizer's children.
- Link SpeechDemo-Synthesis-EngineSessionReuse into
SpeechDemo-Synthesis-TextToSpeech's children.
- Replace stale pre-redesign test-name references
(SpeechRecognizerFactory_Create_.../SpeechSynthesizerFactory_Create_...,
SynthesizeStreamAsync_.../PlayStreamAsync_...) with the current
LoadAsync-era and system-integration test names in
docs/reqstream/ots/sherpa-onnx.yaml,
docs/reqstream/speech/model-management-subsystem/speech-model-contract.yaml,
and docs/reqstream/speech.yaml.
Verified via 'dotnet reqstream --enforce' against local TestResults/*.trx:
no orphaned requirements and no Speech-specific unsatisfied requirements
remain; only pre-existing CI-only OTS/Quality/Template requirements
(requiring BuildMark/SarifMark/SonarMark/ReviewMark/Pandoc/WeasyPrint/
FileAssert/cross-platform-matrix evidence not produced by a local
build.ps1 run) remain unsatisfied.
- Shorten overlong table rows in docs/design/speech.md,
docs/design/speech-cli.md and its subsystem docs, docs/design/speech-demo.md
and recognition-panel-subsystem.md so no line exceeds the 120-character
MD013 limit, without altering the technical content of the tables.
- Apply fix.ps1 markdown normalization (trailing blank line removal,
emphasis marker normalization) across the affected design/verification
files.
Also add cspell.yaml entries for genuine technical/compound words introduced
by this redesign's documentation: overwritable, provisionals, precheck,
unrequested, unawaited, unconfigured.
- Rename stale Create -> LoadAsync references in doc comments, design docs,
verification docs, and reqstream YAML across the demo, library, and CLI
projects (ModelSettingsViewModel.cs plus 13 additional files found via
repo-wide grep).
- Fix stale internal seam references: IRecognitionEngine/IRecognitionEngineFactory
-> IRecognitionBackend/IRecognitionBackendFactory in docs/design/ots/sherpa-onnx.md,
and ISynthesisEngine -> ISynthesisBackend / ISpeechSynthesizer -> ISpeechSynthesizerEngine
in ISynthesisModel.cs, SherpaOnnxVitsLibriTtsEnglishSynthesisModel.cs,
FakeSynthesisModel.cs, and docs/reqstream/speech.yaml.
- Rewrite README.md's Speech-to-text and Text-to-speech usage examples to use
the current async Engine/Session API (LoadAsync, CreateSessionAsync,
GetResultsAsync, SpeakAsync), mirroring docs/user_guide/introduction.md.
No behavioral code changes; doc-comment/prose corrections only.
…on API
The package Description and ApiMarkLibraryDescription still pointed
readers (including coding agents consuming the packaged API docs) at
the old synchronous SpeechRecognizerFactory/SpeechSynthesizerFactory
composition pattern. Update both to start with LoadAsync/
CreateSessionAsync and the session streaming/SpeakAsync surface,
matching the README and user guide.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
- SherpaOnnxSynthesisSession.DisposeAsync now awaits the in-flight
operation's completion (bounded by DedicatedWorker's cooperative-
cancel-then-abandon policy) before releasing the engine's exclusivity
lease, closing the window where a new session could call into the
same backend while the prior one's native call might still be running.
- Fixed a TOCTOU race where StopAsync/DisposeAsync could call
CancelAsync() on an operation's CancellationTokenSource concurrently
with RunOperationAsync disposing that same source, surfacing an
unhandled ObjectDisposedException.
- SherpaOnnxRecognitionSession.StopAsync now distinguishes the caller
that owns teardown from concurrent callers that only await the
shared pump task, so non-idempotent teardown (backend.Reset(),
device.Stop(), etc.) never runs twice.
- Fixed a pre-existing flaky test
(SherpaOnnxSpeechRecognizerEngine_CreateSessionAsync_PriorSessionDisposing_ThrowsRecognitionEngineBusyException):
its mock device.Stop() used a 5-second bounded wait as a stall
mechanism, which could elapse on its own under heavy parallel
test-run CPU contention before the race assertion even ran,
releasing the lease early and intermittently failing the test for
the wrong reason. Now blocks until explicitly released, with the
test's own cancellation token as a safety net instead of a fixed
timeout.
- Added regression tests for all of the above.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
… review-sets
Pure documentation/requirements correction pass addressing every finding from
the formal-review agent's 'dotnet reviewmark --elaborate' sweep of the 27
in-scope review-sets (25 Fail + 2 Pass-with-minor-findings) affected by the
async Engine/Session redesign. No functional/behavioral code changes; all
edits are design/verification/reqstream prose or XmlDoc comments.
Categories of fixes:
- Stale pre-redesign API references (Create/GenerateSegment/ISpeechRecognizer/
SherpaOnnxSpeechSynthesizer) replaced with current async LoadAsync/
GenerateSegmentAsync/SherpaOnnxSynthesisSession names across design docs,
verification docs, reqstream yaml justifications, and a handful of XmlDoc
<remarks>/<exception> blocks (SpeechRecognizerFactory, SpeechModelCatalog,
UnavailableSpeechRecognizerEngine, SpeechModelParameterDiagnostics,
IRecognitionSession).
- Special case: corrected unavailable-recognition-session.md and the embedded
summary in recognition-subsystem.md to match the already-correct
UnavailableRecognitionSession implementation (State=Created, StopAsync/
DisposeAsync safe no-ops, GetResultsAsync throws eagerly) - the code was not
changed, only the doc that had drifted from it.
- Corrected the "never throws" claim for UnavailableSpeechRecognizerEngine to
note it still throws for a null device/already-cancelled token (caller
errors), matching the implementation.
- Fixed requirement-text/behavior mismatches: RecognitionCommandSubsystem's
--start-timeout default (fixed 8s, not a silence-timeout fallback), the
recognition CatalogSeamExtension requirement narrowed to reflect
GetAudioFormat having no production call site, and the SpeechModelContract
SynthesisEngineDeclaration compound requirement split into two IDs with its
cross-subsystem test link corrected.
- Updated SpeechCli subcommand/subsystem counts (ten->eleven subcommands,
four->five subsystems) across introduction.md, speech-cli.yaml, and
verification/speech-cli.md to include the ask/ConversationCommandSubsystem
addition, including a new SpeechCli-ModelCommands-RecognitionSeamExtension
requirement mirroring the existing synthesis-side requirement.
- Added missing documentation: SpeechDemo shell-subsystem's async
DisposeAsync shutdown ordering, and SpeechDemo RecognitionPanelSubsystem's
undocumented CaptureDebugRecorder/WavFileWriter diagnostic instrumentation.
- Added 7 missing test names to conversation-command-subsystem.yaml's
requirement test lists to match the verification doc.
Validated via fix.ps1 (no-op), lint.ps1 (yamllint/cspell/markdownlint-cli2/
reqstream/versionmark/reviewmark/sysml2tools/dotnet format all clean), a
reqstream --enforce run confirming no new orphaned/unsatisfied requirements
beyond the pre-existing CI-only baseline (Quality-*/Template-OTS-*/
Template-Platform-*), and build.ps1 (2630/2630 tests passing).
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Adds targeted unit tests (plus companion reqstream requirement updates)
for six deferred findings from the async Engine/Session redesign's
formal-review pass:
1. RecognitionEngineBusyException constructors - new test file covering
the default, message, and message+inner-exception constructors.
2. SherpaOnnxVitsLibriTtsEnglishSynthesisModel.InstallAsync - covers
null/empty stagedFilesDirectory argument validation.
3. UnavailableSpeechRecognizerEngine.CreateSessionAsync - covers an
already-cancelled token throwing OperationCanceledException.
4. SherpaOnnxSpeechRecognizerEngine.CreateSessionAsync - covers the
unavailable-device fallback (Warning diagnostic + Unavailable
session) with a new requirement
Speech-Recognition-RecognitionBackend-UnavailableDeviceFallback.
5. RecognizeCommand's mic-mode start-timeout default - extracts the
inline timeout-resolution expressions into testable internal static
ResolveStartTimeout/ResolveSilenceTimeout helpers (no behavior
change) and adds tests proving the 8-second start-timeout default
is fixed and independent of --silence-timeout.
6. ask subcommand help listing - extends
SpeechCli_HelpFlag_Provided_ListsAllSubcommands to name-check ask
and adds Program_Run_WithAskCommand_DoesNotThrowNotImplemented to
parallel the existing one-per-subsystem dispatch-smoke-test pattern.
All 2655 tests pass; fix.ps1/lint.ps1 clean.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Found by an independent sub-agent reviewing only the packaged ApiMark
doc tree as a fresh consumer would:
1. SpeechSynthesizerFactory's remarks referenced 'the former synchronous
factory' - a leftover comparison to the pre-redesign API that is
meaningless/confusing to a reader who never saw the old shape. Removed.
2. Both SpeechRecognizerFactory and SpeechSynthesizerFactory's remarks
described their three LoadAsync overloads via adjacent <see cref>
links; ApiMark renders every overload link with identical 'LoadAsync()'
text, so the generated prose read as 'LoadAsync(), LoadAsync(), or
LoadAsync()' with no way to tell the overloads apart. Reworded to name
each overload's distinguishing parameter (installedModelDirectory path,
SpeechModelStore, SpeechModelCatalog) directly in prose, so the
generated doc is unambiguous regardless of link-text collapsing.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
CI's pandoc design-doc build failed with:
docs/design/speech/recognition-subsystem/i-speech-recognizer.md: does not exist
The RecognitionSubsystem/SynthesisSubsystem async redesign renamed and split
several design doc files (e.g. i-speech-recognizer.md -> i-speech-recognizer-
engine.md + i-recognition-session.md), but docs/design/definition.yaml's
input-files list was never updated to match:
- Fixed 4 stale recognition-subsystem filenames to their current names.
- Added 2 new recognition-subsystem files that had no entry at all
(i-recognition-session.md, unavailable-recognition-session.md).
- Added the entire synthesis-subsystem.md + 7 synthesis-subsystem/*.md files,
which were completely absent from the design doc build despite existing on
disk since the redesign.
Verified by running the exact CI pandoc command locally against all 78
referenced input files (all now resolve) and confirming design.html
generates successfully.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
This testRef points to SherpaOnnxSynthesisSessionTests, which exercises a fake ISynthesisBackend; there is no direct test of SherpaOnnxSynthesisEngine in the test tree. Per the repository's SysML artifact convention, an indirectly exercised unit should omit testRef rather than point at an unrelated test, or add a direct backend test before referencing one.
Update verification to reference GenerateSegmentAsync
The redesigned API no longer has SherpaOnnxSpeechSynthesizer.GenerateSegment; this verification note still names that removed type/method after updating the factory call. Update it to SherpaOnnxSynthesisSession.GenerateSegmentAsync so the documented end-to-end proof points to a real implementation.
Update verification to reference GenerateSegmentAsync
The redesigned API no longer has SherpaOnnxSpeechSynthesizer.GenerateSegment; this verification note still names that removed type/method after updating the factory call. Update it to SherpaOnnxSynthesisSession.GenerateSegmentAsync so the documented end-to-end proof points to a real implementation.
- docs/verification/definition.yaml had the same stale/missing-file-reference
bug as the design-doc definition.yaml: 4 stale recognition-subsystem
filenames from the pre-redesign IRecognitionSession/IRecognitionSession
split, and the entire synthesis-subsystem section was missing. Corrected
the filenames and added all 8 synthesis-subsystem verification files.
Verified by running the exact CI pandoc command locally (success).
- docs/reqstream/speech/recognition-subsystem.yaml: the rename left two
requirements unreachable from any root-tagged requirement (reqstream
--enforce orphan detection): Speech-Recognition-RecognitionBackend-
UnavailableDeviceFallback and Speech-Recognition-RecognitionEngineBusy
Exception-Constructors existed but were never added to a parent's
children list. Added them to UnavailableFallback and EngineExclusivity
respectively. Verified via dotnet reqstream --enforce (0 orphans).
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…o Created
Findings 16-18 (PR #40 review). UnavailableRecognitionSession.State always
returns RecognitionSessionState.Created (verified against the shipped
implementation), but the sysml2 model and two verification docs recorded
Faulted, contradicting the implementation and the companion design docs.
Aligned all three artifacts with the actual runtime behavior.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…view
Findings 1-6, 14 (PR #40 review):
- SherpaOnnxRecognitionSession.StartAsync no longer releases _syncRoot before
device start/pump creation completes; a concurrent StopAsync now serializes
against startup instead of racing it back to Running (finding 1).
- File-mode pump ordering: the pump is now started before the capture device,
with startup rollback if device.Start() fails, so the channel is being
drained while a file-backed device synchronously emits its frames instead
of overflowing before a reader exists (finding 2).
- FaultSession now runs the same teardown path as StopAsync (completing
_pendingFrames, stopping the device, awaiting the pump) while preserving
the fault for result consumers, instead of returning early and leaking the
capture stream/pump (finding 3).
- The engine lease is now released only after the pump's raw completion is
observed, not merely after the abandon-aware wrapper task gives up; an
abandoned pump worker therefore keeps the lease (and blocks backend reuse/
disposal) until it genuinely exits, while StopAsync/FaultSession themselves
remain fast and abandon-tolerant as the existing abandon-timeout regression
test requires (finding 4).
- DisposeAsync now shares a single in-flight teardown task across concurrent
callers instead of letting every call after the first return immediately,
so concurrent disposal genuinely awaits the same teardown (finding 5).
- SherpaOnnxSpeechRecognizerEngine.CreateSessionAsync now performs the
disposed check and lease acquisition under one synchronization boundary
with DisposeAsync, rolling back the lease if session construction fails,
closing the create-versus-dispose race (finding 6).
- SpeechRecognizerFactory's detached load-worker completion now disposes any
backend it eventually creates after the caller has already observed
cancellation, preventing a leaked native backend (finding 14).
Added regression tests exercising each race: concurrent StartAsync/StopAsync,
synchronous file-emission pump draining via spin-wait synchronization,
fault-then-dispose teardown, lease retention until an abandoned worker
exits, concurrent double-DisposeAsync, and an unsynchronized
CreateSessionAsync/DisposeAsync race stress loop.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…rify operation-tracking race (PR #40 findings 19-20)
Finding 19: SherpaOnnxRecognitionSession.StopAsync accepted a
CancellationToken parameter but never observed it, so a caller could
not bound a stuck wait despite the documented contract. StopAsync now
captures the shared teardown task under the lock and wraps the
caller''s own wait with Task.WaitAsync(cancellationToken) when the
token can be canceled, letting that caller''s returned task complete
early with OperationCanceledException without aborting the shared
teardown relied on by every other concurrent caller (including
DisposeAsync). Updated IRecognitionSession.StopAsync''s XML doc to
describe this precisely: the token bounds only this caller''s own
wait, never the underlying drain/stop.
Finding 20: reported that RunOperationAsync/TrackOperation in
SherpaOnnxSynthesisSession could let a rejected overlapping call
overwrite the tracked _operationTask, so a racing DisposeAsync would
await only the rejected call''s faulted task instead of the original
in-flight operation. Verified against current code (post finding-8
redesign) that StartOperation already validates overlap/dispose/fault
state entirely before ever touching _operationTask, all inside one
lock, so a rejected overlapping call cannot reach the assignment - the
race no longer reproduces. Added a regression test proving
DisposeAsync issued immediately after a rejected overlapping call
still awaits the original operation to genuine native completion
before completing.
New regression tests:
- SherpaOnnxRecognitionSession_StopAsync_CallerTokenCanceled_ReturnsEarlyWithoutAbortingSharedTeardown
- SherpaOnnxSynthesisSession_DisposeAsync_AfterRejectedOverlappingCall_StillAwaitsOriginalOperation
Validated with fix.ps1, build.ps1 (2694/2694 tests passing), and
lint.ps1 (clean).
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…cancellation races
Findings 21-25 (SherpaOnnxRecognitionSession.cs):
- Split TransitionTo/RaiseStateChanged so StateChanged handlers never run while
_syncRoot is held (closes a deadlock risk if a handler calls back into
StopAsync/DisposeAsync).
- Fault the session (not just log) when AcceptSamples/TryDecode throws on the
pump thread, completing the result buffer so GetResultsAsync cannot hang.
- Complete the result buffer when StopAsync tears down a never-started
(Created-state) session.
- Defer only the backend reset/device stop - never the session's own state
convergence - until an abandoned pump worker's raw completion genuinely
happens, so DisposeCoreAsync cannot release the lease while that worker may
still be touching the shared backend.
- Restructure teardown creation (new TeardownStart type) so starting the
shared teardown operation is deferred until after the caller has raised its
own Stopping/Faulted transition, instead of EnsureTeardownStartedLocked
starting it inline - closing a StateChanged ordering race (Stopped observed
before Stopping) exposed under Release/net8.0/net9.0 timing.
Findings 26-28 (SherpaOnnxSynthesisSession.cs):
- Classify an OperationCanceledException from a still-running abandoned
native Generate call as a session fault, not cooperative cancellation, so
the session cannot be reused while that call is still in flight.
- StopAsync now captures and awaits the raw native-call completion (not just
the abandon-aware wrapper task) so it only converges once the native call
has genuinely finished.
- StopAsync's cancellationToken now only bounds the caller's own wait via
WaitAsync, never the shared teardown, mirroring the recognition session.
Finding 29 (App.axaml.cs): set the shutdown guard flag synchronously inside
the ShutdownRequested handler, before the async cleanup begins, so a second
shutdown request cannot start an overlapping cleanup.
Findings 30-31 (SpeechRecognizerFactory.cs / SpeechSynthesizerFactory.cs):
check cancellation after a dedicated-worker load completes successfully but
within the abandon grace period, disposing the backend and throwing instead
of reporting a loaded engine.
Findings 32-33: correct stale SherpaOnnxSpeechSynthesizer.GenerateSegment
references to SherpaOnnxSynthesisSession.GenerateSegmentAsync in two
synthesis-model verification docs.
Adds regression tests for every fix above using deterministic
semaphore/ManualResetEvent-gated fakes (no Thread.Sleep), including a
200-iteration StateChanged ordering test and new blocking backend-factory
test doubles for findings 30/31.
pwsh ./fix.ps1, ./build.ps1 (2727/2727 tests across net8.0/net9.0/net10.0),
and ./lint.ps1 all pass clean.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The PR #40 review-fix pass renamed
SherpaOnnxSynthesisSession_StopAsync_WhileSpeaking_CancelsInFlightOperationOnlyAfterInFlightGenerateReturns
to
SherpaOnnxSynthesisSession_StopAsync_WhileSpeaking_DoesNotCompleteUntilInFlightGenerateReturns
(reflecting that StopAsync's token now only bounds the caller's own wait rather than
genuinely canceling the drain), but 3 reqstream requirement files still referenced the
old test name, leaving 5 requirements unsatisfied under --enforce:
Speech-Synthesis-StreamingSynthesizer, Speech-Synthesis-MockableContract,
Speech-Synthesis-FaultContainment, Speech-Synthesis-ISynthesisSession-StreamingContract,
and Speech-Synthesis-SherpaOnnxSynthesisSession-CancellationAndLifecycle. Updated all 3
references to the current test name. Verified via dotnet reqstream --enforce against the
local TestResults/*.trx: all 5 requirements now satisfied (491/519, remaining 28 are
CI-only Template-OTS/Quality self-validation items not run locally).
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…adlock
RunTeardownAsync was invoked fire-and-forget (_ = RunTeardownAsync(...))
from StopAsync, DisposeCoreAsync, and FaultSession. When its internal
awaits (pumpCts.CancelAsync()/the pump-task await) happened to complete
synchronously - which they often do - the entire method, including the
blocking device.Stop() native call, ran inline on the calling thread
instead of in the background.
This defeated DisposeAsync/StopAsync's documented
teardown-runs-in-the-background contract and deadlocked
SherpaOnnxSpeechRecognizerEngine_CreateSessionAsync_PriorSessionDisposing_ThrowsRecognitionEngineBusyException,
whose whole premise is that disposal can be observed in-flight, not yet
complete.
Fix: add an unconditional await Task.Yield() as the first statement in
RunTeardownAsync so control always returns to the fire-and-forget caller
before any synchronous/native teardown work runs, regardless of whether
the internal awaits happen to complete synchronously.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
- SherpaOnnxRecognitionSession.StartAsync: run the native, potentially slow
IAudioCaptureDevice.Start() call through the session's DedicatedWorker
instead of inline under _syncRoot, so this method no longer blocks the
caller/UI thread for the life of that call. The documented invariant that a
concurrent StopAsync/DisposeAsync can never converge the session to Stopped
and return while the device is still starting is preserved by publishing the
device-start task to a new _deviceStartTask field under the lock before it
is released, and having RunTeardownAsync always await that same task before
it ever stops the device.
- SherpaOnnxSynthesisSession.GenerateAndOptionallyPlayAsync: move the
synchronous, native-backed IAudioPlaybackDevice.Start() call behind a
Task.Run hand-off so this already-async method's first blocking device call
happens after a genuine asynchronous boundary, instead of inline before its
first await.
- WavFileAudioCaptureDevice: corrected the class remarks, which referenced
SpeechRecognizerFactory.LoadAsync as the IAudioCaptureDevice composition
seam; it is actually ISpeechRecognizerEngine.CreateSessionAsync.
- SpeechSynthesizerUnavailableException: corrected the class remarks, which
no longer matched shipped behavior - a real session's playback failures
propagate as AudioDeviceUnavailableException and calls after a session has
faulted throw SynthesisSessionFaultedException, neither of which is this
type.
Added unit tests proving StartAsync/SpeakAsync return control to the caller
without blocking while the respective device Start() call is still in
flight, following the existing ManualResetEventSlim-based fake-device idiom
used elsewhere in both test suites.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…gnition start cancellation gap
- SherpaOnnxSynthesisSession: StopAsync/DisposeAsync now re-read
_pendingNativeCompletion after the operation task settles instead of
snapshotting it upfront. GenerateSegmentAsync publishes that field strictly
before it awaits the native call, so the earlier upfront snapshot could
still be null/stale while the operation was genuinely in flight, letting
StopAsync/DisposeAsync report the operation stopped (and release the
engine lease) while an abandoned native Generate call was still running.
This is the root cause of the macOS CI failure in
SherpaOnnxSynthesisSession_StopAsync_AbandonedGenerate_DoesNotCompleteUntilNativeCallGenuinelyReturns
(confirmed: 10/10 local reruns now pass).
- SherpaOnnxRecognitionSession: StartAsync's device-start worker call now
passes the caller's own cancellationToken (previously CancellationToken.None),
so cancelling StartAsync while the device is still starting is honored via
the same cooperative-cancel-then-abandon policy used elsewhere, instead of
being ignored indefinitely. Split the published task into the abandon-aware
task StartAsync itself awaits and a raw completion (_deviceStartRawCompletion,
mirroring _pumpRawCompletion) that RunTeardownAsync awaits for genuine
completion, preserving the invariant that device.Stop() never races a
still-in-flight device.Start(). Added a new regression test.
Validation: build.ps1 (all TFMs) passed; net8.0 suite run twice with no hang;
fix.ps1/lint.ps1 clean; reqstream 491/519 (expected, no regression).
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Speech-Synthesis-SherpaOnnxSynthesisSession-ChunkedPipeline previously said
SynthesizeAsync should 'yield each resulting segment as it becomes ready,'
which no longer matches the shipped API: SynthesizeAsync returns the full
ordered segment list only once synthesis completes (chunked
synthesis/playback pipelining is an internal optimization SpeakAsync relies
on directly, not a streaming result contract SynthesizeAsync exposes to its
caller). Reworded the requirement to describe the actual, already-reviewed
behavior rather than changing the public API.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…s device-start cancellation
- RecognitionResultBuffer.AddResult: supersede a not-yet-consumed
provisional result when a final result arrives for the same
utterance, so consumers never see a stale provisional trailing its
own final. Updated class remarks and AddResult XML docs. Added
dedicated RecognitionResultBufferTests.cs covering both the
supersede case and the case where an already-consumed provisional
for a different utterance is unaffected. Updated the integration
test SpeechTests.Speech_SystemIntegration_StreamingRecognition_
CapturedAudioProducesRecognitionResults to assert the corrected
single-final-result contract.
- SherpaOnnxSpeechRecognizerEngine.ReleaseLease and
SherpaOnnxSpeechSynthesizerEngine.ReleaseLease: move the
_lease.Release() call inside the _syncRoot lock to close a race
where a concurrent lease acquisition could observe inconsistent
engine state between the release signal and the lock-protected
bookkeeping update.
- SherpaOnnxSynthesisSession.GenerateAndOptionallyPlayAsync: route
the playback device-start call through DedicatedWorker (passing the
real cancellation token) instead of an unabandon-able Task.Run, and
publish its raw completion through the same _pendingNativeCompletion
field GenerateSegmentAsync already uses. This lets the existing
Faulted-vs-Stopped decision in RunOperationAsync work correctly for
device-start abandonment (fault fast, matching the Generate
abandonment precedent) while still deferring the actual
_device.Stop() call to a new continuation
(StopDeviceAfterGenuineStartCompletionAsync) when the start was
genuinely abandoned, so StopAsync/DisposeCoreAsync still correctly
wait for the device to genuinely stop. Added regression test
SherpaOnnxSynthesisSession_StopAsync_AbandonedDeviceStart_
DoesNotCallDeviceStopUntilStartGenuinelyReturns.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…ompletion when operation task already cleared
SherpaOnnxSynthesisSession.CancelAndAwaitOperationAsync returned
immediately whenever the caller's snapshot of _operationTask was null,
without ever checking _pendingNativeCompletion. RunOperationAsync's
abandon-fault handler clears _operationTask (and _operationCancellation)
eagerly as part of transitioning to Faulted, while deliberately leaving
_pendingNativeCompletion set to the still-running abandoned native
call. A StopAsync/DisposeAsync call made after that fault was already
observed would therefore see operationTask as null and skip the
pending-native-completion wait entirely, letting DisposeAsync release
the engine's exclusivity lease (and let the backend be reused) while
the abandoned call was still genuinely executing.
Fixed by always re-reading and awaiting _pendingNativeCompletion,
regardless of whether operationTask was present. Updated the method's
XML docs to explain the null-operationTask case, and added a
regression test
(SherpaOnnxSynthesisSession_DisposeAsync_CalledAfterAbandonmentFault_DoesNotReleaseLeaseUntilNativeCallGenuinelyReturns)
that disposes only after the abandonment fault has already been
observed, proving the lease is still held until the native call
genuinely returns.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Pull Request
Description
Redesigns the public recognition and synthesis APIs around a consistent, fully
async Engine/Session model:
ISpeechRecognizerEngine/ISpeechSynthesizerEngineload models and create sessions asynchronously; sessions stream results as
IAsyncEnumerable/events, and are explicitly cancelled (Cancel()) anddisposed (
DisposeAsync()) rather than relying on synchronous disposal ofblocking native calls. Native/blocking work is moved off the caller's thread
(worker threads / long-running
Task.Runhints) so the public API surface isasync end-to-end, while the underlying native libraries remain blocking
internally. One-active-session-per-engine cardinality is enforced via
RecognitionEngineBusyException, and the existing "honest unavailable, nothrow" convention is preserved for
LoadAsync/CreateSessionAsyncfailurepaths.
This is a deliberate 0.x.y breaking API change, approved by the project
owner, to avoid the library's public shape ossifying around a blocking-first
design that the rest of the .NET ecosystem has moved away from.
Includes:
RecognitionSubsystem,SynthesisSubsystem) to theasync Engine/Session API, including internal
IRecognitionBackend/ISynthesisBackendrenames to free up theEnginename for the newpublic types.
DemaConsulting.Speech.Cliand the Demo app to the new API,including a fix for a synthesis session recreate-per-click bug uncovered
during migration.
SilenceTimeoutRecognizerSession, double-dispose/busy-session edge cases).stale doc-comment references updated to describe the new async API so
downstream coding agents are guided toward the new surface.
references, line-length/formatting compliance).
argument validation, cancellation handling, device-unavailable fallback,
testable timeout-resolution helpers,
asksubcommand help-listing).Type of Change
Related Issues
Closes #
Pre-Submission Checklist
Build and Test
pwsh ./build.ps1(2655/2655 passed)Code Quality
Quality Checks
pwsh ./lint.ps1Testing
Documentation
Additional Notes
Three independent review passes (code-review agent + formal-review agent across
all affected review-sets) were run after implementation and all findings were
addressed, including this final test-coverage-gap closure commit. This is a
0.x.y breaking change; downstream consumers will need to migrate to the new
Engine/Session API (documented in ApiMark/design docs).