Skip to content

Add microphone recording to Play Audio - #520

Open
nikhil-vincent wants to merge 3 commits into
kercre123:mainfrom
nikhil-vincent:audio-record-and-stream
Open

nikhil-vincent wants to merge 3 commits into
kercre123:mainfrom
nikhil-vincent:audio-record-and-stream

Conversation

@nikhil-vincent

@nikhil-vincent nikhil-vincent commented Aug 13, 2026 •

Copy link
Copy Markdown

Summary

The Vector control page's Play Audio section only accepted a .wav file. This adds a Record / Stop / Send flow so you can capture from the browser microphone and play it on Vector through the existing /api-sdk/play_sound endpoint.

Recording uses raw PCM (ScriptProcessor) instead of MediaRecorder/WebM, then converts to 8 kHz mono 16-bit WAV with a simple anti-alias downsample, light speech normalize, and short fades. File uploads also accept common formats (mp3, ogg, m4a, webm) and go through the same conversion.

Assume behavior control first, same as the rest of the control page.

Test plan

  • Open Vector control, assume behavior control
  • Click Record, speak, click Stop recording, then Send — Vector should play the clip
  • Confirm the preview player does not auto-play after recording
  • Upload a .wav file and send (existing path still works)
  • Upload an .mp3 (or other listed format) and send
  • Confirm Record without a serial in the URL shows an error
  • Confirm Record works after granting mic permission, and is blocked if permission is denied

@F4310 F4310 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Summary of Key Risks & Action Items

🔴 Critical Issues

  • Memory Leak (AudioContext): recContext.close() is called loosely in cleanupRecGraph() without state validation or guarantee, accumulating audio context instances over multiple recordings.
  • Uncapped Recording Size: Lack of duration/size limits allows multi-minute recordings (~50 MB+ uncompressed raw PCM), risking client crash or server network DoS.
  • Generic Error Handling: getUserMedia failures yield generic alerts; users can't tell if the issue is a rejected permission or missing hardware.

🟠 Medium Issues

  • Sample Rate Mismatch: recSampleRate is read before the browser finalizes audio context creation, causing potential pitch and playback speed skew during downsampling.
  • Aggressive Normalization: High peak target (0.82) and max gain (4.0x) without noise floor or VAD checks severely amplifies static on quiet or empty clips.

📋 Action Plan

  1. [High] Enforce a max recording limit (e.g., 30 seconds).
  2. [High] Ensure deterministic cleanup of AudioContext in cleanupRecGraph().
  3. [High] Handle getUserMedia exception types explicitly (NotAllowedError, NotFoundError).
  4. [Medium] Read recContext.sampleRate dynamically post-resume().
  5. [Medium] Implement basic noise floor checks prior to peak normalization.
  6. [Low] Extend audio fade times from 12 ms to 30–50 ms to prevent 8 kHz click artifacts.

@nikhil-vincent

Copy link
Copy Markdown
Author

Thanks for the review — addressed on c760967:

  1. 30s cap — timer plus sample-count stop; UI notes the limit.
  2. AudioContext cleanup — close() only if state !== "closed", previous close is awaited before a new record.
  3. getUserMedia errors — explicit NotAllowedError, NotFoundError, NotReadableError, OverconstrainedError, SecurityError.
  4. Sample rate — read after resume(), then from inputBuffer.sampleRate on each process callback.
  5. Noise floor — skip normalize when RMS/peak are at the floor; max gain 2.5x.
  6. Fades — 12ms → 40ms.

The Play Audio section only accepted a WAV file. This adds a Record /
Stop / Send flow that captures raw PCM from the browser mic, downsamples
it to 8 kHz mono 16-bit WAV, and posts it to the existing play_sound API.

File uploads also accept common formats (mp3, ogg, m4a, webm) and go
through the same conversion path.
… errors

- Stop automatically at 30s and cap captured PCM
- Close AudioContext only when not already closed; await previous close before a new record
- Map getUserMedia errors (NotAllowedError, NotFoundError, etc.)
- Read sampleRate after resume() and from input buffers
- Skip peak normalize below a noise floor; lower max gain
- Lengthen edge fades to 40ms to reduce 8 kHz clicks
Guard startRecording with recStarting before any await, hide Record
immediately, and discard a late getUserMedia stream if the user cancels
during the permission prompt.
@nikhil-vincent
nikhil-vincent force-pushed the audio-record-and-stream branch from fe91c28 to 57c1c24 Compare September 4, 2026 14:49
@nikhil-vincent

Copy link
Copy Markdown
Author

@F4310 Friendly ping for a re-review when you have a chance.

The earlier review feedback is addressed on this branch (30s recording cap, safer AudioContext cleanup, explicit getUserMedia errors, sample-rate handling after resume, noise-floor normalize, longer fades, plus a double-start guard). I also just rebased onto current main so the PR is up to date — happy to adjust further if anything still looks off.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants