Repository navigation
Add microphone recording to Play Audio - #520
Open
nikhil-vincent wants to merge 3 commits into
Open
nikhil-vincent wants to merge 3 commits into
nikhil-vincent wants to merge 3 commits into
Conversation
F4310
suggested changes
Aug 13, 2026
F4310
left a comment
Contributor
There was a problem hiding this comment.
Summary of Key Risks & Action Items
🔴 Critical Issues
- Memory Leak (
AudioContext):recContext.close()is called loosely incleanupRecGraph()without state validation or guarantee, accumulating audio context instances over multiple recordings. - Uncapped Recording Size: Lack of duration/size limits allows multi-minute recordings (~50 MB+ uncompressed raw PCM), risking client crash or server network DoS.
- Generic Error Handling:
getUserMediafailures yield generic alerts; users can't tell if the issue is a rejected permission or missing hardware.
🟠 Medium Issues
- Sample Rate Mismatch:
recSampleRateis read before the browser finalizes audio context creation, causing potential pitch and playback speed skew during downsampling. - Aggressive Normalization: High peak target (0.82) and max gain (4.0x) without noise floor or VAD checks severely amplifies static on quiet or empty clips.
📋 Action Plan
- [High] Enforce a max recording limit (e.g., 30 seconds).
- [High] Ensure deterministic cleanup of
AudioContextincleanupRecGraph(). - [High] Handle
getUserMediaexception types explicitly (NotAllowedError,NotFoundError). - [Medium] Read
recContext.sampleRatedynamically post-resume(). - [Medium] Implement basic noise floor checks prior to peak normalization.
- [Low] Extend audio fade times from 12 ms to 30–50 ms to prevent 8 kHz click artifacts.
Author
|
Thanks for the review — addressed on
|
The Play Audio section only accepted a WAV file. This adds a Record / Stop / Send flow that captures raw PCM from the browser mic, downsamples it to 8 kHz mono 16-bit WAV, and posts it to the existing play_sound API. File uploads also accept common formats (mp3, ogg, m4a, webm) and go through the same conversion path.
… errors - Stop automatically at 30s and cap captured PCM - Close AudioContext only when not already closed; await previous close before a new record - Map getUserMedia errors (NotAllowedError, NotFoundError, etc.) - Read sampleRate after resume() and from input buffers - Skip peak normalize below a noise floor; lower max gain - Lengthen edge fades to 40ms to reduce 8 kHz clicks
Guard startRecording with recStarting before any await, hide Record immediately, and discard a late getUserMedia stream if the user cancels during the permission prompt.
nikhil-vincent
force-pushed
the
audio-record-and-stream
branch
from
September 4, 2026 14:49
fe91c28 to
57c1c24
Compare
Author
|
@F4310 Friendly ping for a re-review when you have a chance. The earlier review feedback is addressed on this branch (30s recording cap, safer |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
The Vector control page's Play Audio section only accepted a
.wavfile. This adds a Record / Stop / Send flow so you can capture from the browser microphone and play it on Vector through the existing/api-sdk/play_soundendpoint.Recording uses raw PCM (
ScriptProcessor) instead ofMediaRecorder/WebM, then converts to 8 kHz mono 16-bit WAV with a simple anti-alias downsample, light speech normalize, and short fades. File uploads also accept common formats (mp3, ogg, m4a, webm) and go through the same conversion.Assume behavior control first, same as the rest of the control page.
Test plan
.wavfile and send (existing path still works).mp3(or other listed format) and send