-
Notifications
You must be signed in to change notification settings - Fork 30
Add Qwen3-TTS voice design, consented cloning, and optional fine-tuning #5381
Copy link
Copy link
Closed
Labels
area:mediaarea:voiceVoice stack: STT/TTS pipeline, proactive speech, voice tools, call bridgeVoice stack: STT/TTS pipeline, proactive speech, voice tools, call bridgedependenciesProposed from a dependency-freedom auditProposed from a dependency-freedom auditeffort:xhighDispatch reasoning effort: extra highDispatch reasoning effort: extra highmodel:heavyDispatch capability: heavyDispatch capability: heavyplanTracked by /do:replanTracked by /do:replanplan-featureFeature plan filed by the plan-feature brainstormFeature plan filed by the plan-feature brainstorm
Description
Activity
Metadata
Metadata
Assignees
Labels
area:mediaarea:voiceVoice stack: STT/TTS pipeline, proactive speech, voice tools, call bridgeVoice stack: STT/TTS pipeline, proactive speech, voice tools, call bridgedependenciesProposed from a dependency-freedom auditProposed from a dependency-freedom auditeffort:xhighDispatch reasoning effort: extra highDispatch reasoning effort: extra highmodel:heavyDispatch capability: heavyDispatch capability: heavyplanTracked by /do:replanTracked by /do:replanplan-featureFeature plan filed by the plan-feature brainstormFeature plan filed by the plan-feature brainstorm
Part of #5377. Blocked by #5380.
Goal
Add one thoroughly integrated richer local voice backend for original voice design, consented rapid cloning, instruction-controlled synthesis, streaming, and later optional fine-tuning.
Proposed approach
studioandinteractiveroutes; benchmark character similarity and first-audio latency before enabling interactive use.Acceptance criteria
Out of scope