Turn YouTube videos, channels, and playlists into a searchable local library of MP3 audio, Whisper transcripts, and structured AI summaries.
VidSift is a private, local-first Node.js application for personal research and media review. It runs on your workstation, binds to 127.0.0.1 by default, and keeps generated content inside the project directory.
- Process individual YouTube videos, entire channels, and playlists.
- Extract MP3 audio with
yt-dlp. - Transcribe locally with the Whisper CLI.
- Generate configurable GitHub Flavored Markdown summaries.
- Use an OpenAI-compatible HTTP endpoint or signed-in Codex and Claude CLI sessions.
- Run transcription-only or collection summary-only workflows when an LLM or media download is unnecessary.
- Browse persistent history with full-text search, favorites, folders, pinning, and drag-and-drop ordering.
- Edit transcripts, regenerate outputs, play saved audio, and open the original YouTube video.
- Resume active jobs after a browser reload and retry failed collection items without reprocessing completed videos.
- Keep multiple reusable summary templates for different research workflows.
- Quick Start
- How It Works
- Installation
- Summary Providers
- Configuration
- Library and Data
- Architecture
- API
- Development
- Troubleshooting
- Security
- License
- Node.js 22 or newer
- Python
yt-dlpwith its default dependencies- FFmpeg available on
PATH - Whisper CLI
- At least one summary provider, unless you only use transcription-only mode
git clone https://github.com/jamesbubenik/VidSift.git
cd VidSift
npm install
Copy-Item .env.example .env
npm startOpen http://127.0.0.1:3000, select Process URL, and submit a YouTube video, channel, or playlist URL.
Note
.env is optional. The application has local defaults, and summary settings can be changed from the configuration screen.
- VidSift validates the YouTube URL and reads its metadata with
yt-dlp. - For a channel or playlist, it discovers the collection and creates a dedicated Library folder.
- It downloads each video's audio as MP3.
- The local Whisper CLI produces a transcript.
- Unless Transcription only is selected, the chosen provider turns the transcript into a structured Markdown summary.
- VidSift saves the generated files, metadata, and workflow state to the local Library.
Collection jobs process one video at a time. Completed videos are skipped during refreshes; failed, pending, or interrupted videos remain eligible for retry. Members-only videos are detected and skipped without stopping public videos in the same collection.
- Minimize an active workflow while processing continues.
- Reload or reopen the browser without losing the active job view.
- Gracefully stop a channel or playlist after its current video finishes.
- Refresh a collection to process new and previously failed videos.
- Use Summary only during collection refresh to regenerate summaries from saved transcripts.
- Review readable per-video errors without aborting the rest of a collection.
Install or update yt-dlp:
python -m pip install -U "yt-dlp[default]"VidSift uses python -m yt_dlp on Windows to avoid command-shim and permission issues. Confirm all required commands are available:
node --version
python -m yt_dlp --version
whisper --help
ffmpeg -versionInstall dependencies and start the server:
npm install
npm startPowerShell start and stop helpers are also included:
.\start-server.ps1
.\stop-server.ps1Install Node.js, Python, FFmpeg, and Whisper with the appropriate package manager, then install yt-dlp:
python3 -m pip install -U "yt-dlp[default]"
npm install
npm startOn non-Windows systems, VidSift tries the yt-dlp executable first and falls back to the Python module.
VidSift supports three provider modes:
| Provider | How it works | Setup |
|---|---|---|
| OpenAI-compatible / Local | Sends chat-completions requests to a configurable HTTP endpoint. | Set the base URL, optional API key, and model in VidSift. |
| OpenAI Account | Runs an ephemeral, read-only codex exec using the Codex CLI's saved OAuth session. |
Install Codex, run codex login, and verify with codex login status. |
| Anthropic Account | Runs Claude with tools disabled and session persistence off using the Claude CLI's saved OAuth session. | Install Claude, run claude auth login, and verify with claude auth status. |
VidSift does not read or store OAuth tokens. The official CLIs own sign-in, credential storage, token refresh, and request authentication. API-key environment variables are removed from account-provider child processes to prevent an OAuth selection from silently falling back to API-key billing.
The configuration screen includes connection checks, model discovery for compatible endpoints, OpenAI account model and reasoning controls, and reusable output templates. Claude model and effort settings remain controlled by the Claude CLI.
Copy .env.example to .env to override the defaults:
| Variable | Default | Description |
|---|---|---|
HOST |
127.0.0.1 |
Server bind address. |
PORT |
3000 |
Server port. |
LLM_BASE_URL |
http://127.0.0.1:1234/v1 |
OpenAI-compatible API base URL. |
LLM_API_KEY |
empty | Optional API key for the compatible endpoint. |
LLM_MODEL |
model-name |
Chat model exposed by the endpoint. |
LLM_TIMEOUT_MS |
0 |
Header/body inactivity timeout; 0 disables it. |
WHISPER_MODEL |
turbo |
Whisper model passed to the CLI. |
LOG_LEVEL |
DEBUG |
One of DEBUG, INFO, WARN, or ERROR. |
Settings saved through the application override the environment defaults for the LLM provider, endpoint, model, prompt, rules, and output sections. The local API key is stored in DATA/summary-config.json when supplied, so protect that file appropriately.
The compatible service must expose:
POST /chat/completions
Model discovery uses its OpenAI-compatible models endpoint. Logs are written to log.txt and rotate to log.1.txt at 5 MB.
The Library groups saved work into Individual Videos, channel folders, and playlist folders. Search covers titles, collection names, source URLs, transcripts, and summaries. Favorites are stored in application history; visual ordering, pins, and collapsed folders are stored in browser-local storage.
Generated content uses this layout:
CHANNELS/
|-- Individual Videos/
| |-- MP3/
| |-- TRANSCRIPTION/
| `-- SUMMARY/
|-- channel-name/
| |-- MP3/
| |-- TRANSCRIPTION/
| `-- SUMMARY/
`-- playlist-playlist-name-id/
|-- MP3/
|-- TRANSCRIPTION/
`-- SUMMARY/
Files share a workflow timestamp and use sanitized, lowercase titles:
YYYYMMDDHHMMSS-sanitized-video-title.ext
Persistent application metadata is stored in:
DATA/history.json— history, favorites, workflow state, collection metadata, and file referencesDATA/summary-config.json— provider and summary configurationDATA/output-templates.json— reusable prompts, rules, and output sections
Generated content and runtime data are ignored by Git. Deleting an item or collection from the Library also removes its generated files and history records.
Browser UI
|
| POST /api/process
v
Express backend
|
+-- youtubeService -> discovery and MP3 extraction with yt-dlp
+-- whisperService -> local Whisper CLI transcription
+-- summaryService -> provider routing and summary formatting
+-- oauthLlmService -> Codex/Claude CLI OAuth summarization
+-- historyService -> persistent history and file metadata
+-- summaryConfigService -> provider and output configuration
+-- fileService -> output directories and cleanup
The browser starts background jobs and polls the Express API for progress. The repository is intentionally small and uses the built-in Node test runner:
.
|-- public/ # Browser UI
|-- src/
| |-- routes/ # Express API routes
| |-- services/ # YouTube, Whisper, summary, history, and file services
| |-- utils/ # Validation, filenames, timestamps, and logging
| |-- config.js
| `-- server.js
|-- tests/ # Unit and service tests
|-- screenshots/ # Project screenshots
|-- CHANNELS/ # Generated media and documents
|-- DATA/ # Persistent application state
|-- .env.example
`-- package.json
Start a video, channel, or playlist job:
POST /api/process
Content-Type: application/json
{
"url": "https://www.youtube.com/playlist?list=PLresearch"
}Set "transcriptionOnly": true to skip summarization. Collection refresh requests can instead use summary-only mode to rebuild summaries from existing transcripts.
Available endpoints
GET /api/health
POST /api/process
GET /api/jobs/active
GET /api/jobs/:id
POST /api/jobs/:id/stop
GET /api/history
GET /api/history/:id
DELETE /api/history/:id
PATCH /api/history/:id/favorite
PATCH /api/history/:id/transcription
POST /api/history/:id/refresh-summary
POST /api/history/:id/refresh-transcription
POST /api/collections/:channelKey/refresh
DELETE /api/collections/:channelKey
GET /api/files/:kind/:filePath
GET /api/config
PUT /api/config
POST /api/config/reset
POST /api/config/models
GET /api/config/auth-status/:provider
GET /api/config/account-models/:provider
GET /api/output-templates
POST /api/output-templates
PUT /api/output-templates/:id
DELETE /api/output-templates/:id
The older /api/channels/:channelKey refresh and delete routes remain available as compatibility aliases.
npm run dev # Start with Node watch mode
npm test # Run the full test suiteTests mock network and external-process boundaries; they do not download from YouTube, invoke Whisper, or call a live LLM. Coverage includes URL classification, command construction, retry behavior, history persistence, collection identity, summary formatting, configuration, output templates, OAuth CLI behavior, ordering, and logging.
GitHub Actions runs the same suite on Windows with Node.js 22 for every push to main and every pull request.
yt-dlp fails or returns HTTP 403
Update yt-dlp and its default/EJS dependencies:
python -m pip install -U "yt-dlp[default]"Confirm Node.js 22 or newer is available, then check the application error and log.txt. VidSift automatically reruns transient downloads up to three total attempts so yt-dlp can obtain a fresh media URL. If YouTube is rate-limiting requests, wait and refresh the collection later.
yt-dlp or Whisper cannot be launched
- Run the failing command from the same terminal used to start VidSift.
- On Windows, verify
python -m yt_dlp --versionandwhisper --help. - Confirm FFmpeg is on
PATH. - Try a smaller Whisper model such as
baseortinyif resources are limited.
Summarization fails
- For a compatible endpoint, confirm the service is running, the base URL is correct, and the selected chat model is available.
- For OpenAI Account, run
codex login status; for Anthropic Account, runclaude auth status. - Use Check Connection after changing CLI authentication.
- Leave
LLM_TIMEOUT_MS=0when local inference can take more than five minutes.
Library ordering does not change
Clear search and favorites-only filters before dragging. Individual Videos, pinned folders, and starred videos are intentionally locked. Ordering is saved per browser profile.
Important
VidSift is a personal workstation utility, not a hardened multi-user service. Do not expose it publicly without adding authentication, authorization, request limits, and deployment hardening.
- Submitted URLs are validated as YouTube URLs.
- External commands use argument arrays instead of user-built shell strings.
- Generated-file routes validate paths and restrict access to project output locations.
- The server binds to the loopback interface by default.
- OAuth-backed modes use only the signed-in workstation user's account.
VidSift is available under the MIT License.
