Personal content capture, extraction & controlled work intake.
Drop a URL for structured, actionable markdown. An authorised π€ reaction can record a pending work intake.
MegaMind is a content capture and work intake system. You drop a link (primarily YouTube, also GitHub or an article) and MegaMind:
- Extracts the content using the best available method per source
- Distils it into structured markdown with insights, actions, and implementation prompts
- Posts the output to a Discord Forum channel with auto-categorised topic tags
- Indexes it in a centralised catalogue with category, tags, and status tracking
- Stores it in Obsidian for offline access across all devices
- Captures work intake β an authorised π€ reaction records a pending request for separate approval
The goal: capture useful content on the go, then review implementation requests before any agent acts.
βββββββββββββββββββββββββββ INPUTS βββββββββββββββββββββββββββββ
β β
β Discord #extract YouTube Playlist "extract" β
β Drop any URL from any (auto-watched, posts to β
β device #extract as audit trail) β
β β
β GitHub Issues (extract label) β
β Mobile capture via GH app β
β β
βββββββββββββββββββββββββββββ¬ββββββββββββββββββββββββββββββββββββ
β
βΌ
ββββββββββββββββββββ EXTRACTION ENGINE βββββββββββββββββββββββββ
β β
β Source Router β detects URL type β dispatches: β
β β
β YouTube β verified captions + preserved local snapshot β
β X/Twitter β existing URL handler; further work in backlog β
β Articles β readability-lxml + BeautifulSoup β
β GitHub β GitHub API + README scrape β
β β
βββββββββββββββββββββββββββββ¬ββββββββββββββββββββββββββββββββββββ
β
βΌ
βββββββββββββββββββββ AI PROCESSING ββββββββββββββββββββββββββββ
β β
β GPT-6 Astra via Codex, with Claude Opus 5.5 API fallback: β
β β
β β Summary β
β β Four source-specific insights β
β β Three concrete next steps β
β β Zero or one copy-ready implementation prompt β
β β Tags + Category (auto-mapped to Forum topic tags) β
β β Source links + references β
β β
βββββββββββββββββββββββββββββ¬ββββββββββββββββββββββββββββββββββββ
β
βΌ
βββββββββββββββββββββ OUTPUTS ββββββββββββββββββββββββββββββββββ
β β
β Discord Forum (#output) β
β β Each extraction = Forum post with topic tags β
β β Auto-categorised with up to 5 tags per post β
β β Filterable by tag β browse by topic β
β β Authorised π€ reaction records a pending work intake β
β β
β Obsidian Vault β
β β Full markdown synced across all devices β
β β
β extractions/INDEX.md β
β β Central catalogue: title, source, category, status β
β β
β Web Dashboard (localhost:8050) β
β β Searchable library, source review, Forum activity β
β β
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
git clone https://github.com/onekiller89/MegaMind.git
cd MegaMind
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txtcp .env.example .envEdit .env with your API keys:
| Key | Required | Purpose |
|---|---|---|
ANTHROPIC_API_KEY |
Yes | Claude API for AI processing |
XAI_API_KEY |
Optional | Grok API for X/Twitter extraction |
DISCORD_BOT_TOKEN |
For bot | Discord bot token for MegaMind |
DISCORD_SERVER_ID |
For bot | RussHub server ID |
DISCORD_EXTRACT_CHANNEL_ID |
For bot | #extract channel ID |
DISCORD_OUTPUT_CHANNEL_ID |
For bot | Forum channel ID (1478880776291487785) |
YOUTUBE_API_KEY |
Optional | YouTube Data API v3 for playlist watcher |
OBSIDIAN_VAULT_PATH |
Optional | Path to Obsidian vault for auto-sync |
DISCORD_WORK_INTAKE_ENABLED |
Optional | Enables pending intake capture; defaults to false |
DISCORD_WORK_INTAKE_REACTOR_IDS |
Optional | Comma-separated IDs authorised to capture intakes |
WORK_INTAKE_ALLOWED_EXECUTORS |
Optional | Executor names allowed for later approval |
WORK_INTAKE_ALLOWED_TARGETS |
Optional | Project aliases allowed for later approval |
# CLI β one-off extraction
python coord.py https://www.youtube.com/watch?v=example
# Discord bot (includes dashboard, YouTube watcher)
python discord_bot.py
# Or via systemd (recommended)
systemctl --user start megamind.servicepython coord.py <URL> Extract from URL
python coord.py --paste <type> Manual paste (youtube|twitter|github|article)
python coord.py --list Show all extractions
python coord.py --list --filter "TODO" Filter by status
python coord.py --status 3 "In Progress" Update entry #3 status
| Command | Description |
|---|---|
/extract <url> |
Extract insights from a URL |
/check |
Force check the YouTube playlist now |
/status |
Show MegaMind bot status |
/search <query> |
Find Forum posts by title or topic tag, with links |
/stats |
Forum post counts, topic activity and recent posts |
/budget |
Show API usage and cost tracking |
/dashboard |
Get the dashboard link |
Drop any URL in #extract and MegaMind processes it automatically. Results appear in the #output Forum as tagged posts.
Auto-starts with the Discord bot (or run standalone with python dashboard.py).
- Searchable, filterable note library with readable insights and actions
- One-click source text review, method labels, flags and reviewer notes
- Workflow status editing that preserves tags in the catalogue
- Forum topics and recent posts, with links back to Discord
- Model and partial API cost provenance; Codex subscription runs have no API dollar charge
Disable auto-start with MEGAMIND_DASHBOARD=0.
Source review records and raw source snapshots stay under ignored local data/ files. For older YouTube notes, View source text can recover captions from the same video ID; it labels these as recovered, not as the original ingestion snapshot. The source method says how text was obtained, not whether every claim in the source is true.
Dashboard workflow status edits update the local extractions/INDEX.md. GitHub receives them when the next extraction is published from a current production main. Source review decisions remain private to this WSL installation.
Model selection lives in .env: MODEL_PROVIDER=codex and CODEX_MODEL=gpt-6-astra use the local Codex CLI sign-in and subscription allowance. If Codex is unavailable, MegaMind uses the Anthropic API key with CLAUDE_MODEL=claude-opus-5-5. ChatGPT/Codex subscriptions do not pay ordinary OpenAI API charges. The concise analysis prompt is in prompts/analysis-v2.md.
The #output channel is a Discord Forum (channel type 15). Each extraction becomes a Forum post, auto-tagged based on AI-detected category.
| Tag | Emoji | Covers |
|---|---|---|
| AI Agents | π€ | Agentic AI, agent frameworks, orchestration |
| AI Tools | π§ | AI products, SDKs, Claude Code, APIs |
| AI Strategy | π§ | AI business strategy, adoption, industry trends |
| Prompting | π¬ | Prompt engineering, system prompts, techniques |
| Automation | β‘ | Workflow automation, scripting, scheduling |
| Productivity | π | PKM, Obsidian, tools, time management |
| Development | π» | Python, JS, web dev, software patterns |
| DevOps | π | Docker, Kubernetes, CI/CD, IaC, GitOps |
| Content Creation | π¬ | Video, writing, design, media production |
| Data Science | π | ML, data engineering, analytics, models |
| Security | π | Cybersecurity, hardening, compliance |
| Fitness | πͺ | Training, nutrition, health, mindfulness |
| Finance | π° | Budgeting, investing, tax, financial planning |
Posts can have up to 5 tags (Discord's limit). The AI processor assigns a primary category, which maps to one or more Forum tags via CATEGORY_TO_FORUM_TAGS. For example:
- "ai image generation" β AI Tools + Content Creation
- "prompt engineering" β Prompting + AI Strategy
- "data science" β Data Science
Tags are resolved in discord_bot.py via resolve_forum_tags() and matched against channel.available_tags.
- Click any tag chip at the top of the Forum to filter by topic
- Posts auto-archive after inactivity but remain visible and filterable
- Archived posts unarchive when someone replies
Every extraction produces a markdown file and a Forum post:
# [Title]
> Source: YouTube | Extracted: 2026-09-23 14:30 UTC | Method: youtube_transcript_api | Analysis: gpt-6-astra
> URL: https://...
### Summary β Two plain sentences about value
### Key Insights β Four source-specific takeaways
### Actions β Three concrete next steps
### Implementation Prompts β At most one useful prompt, or None
### Links & Resources β Original URL and explicitly mentioned resources
### Tags β For categorisation
### Category β Primary category (maps to Forum tags)The analysis target is 400 words or fewer. Claims stay attributed to the speaker when MegaMind has not independently verified them.
| Source | Method | Why |
|---|---|---|
| YouTube | YouTube captions or yt-dlp subtitles | Transcript tied to the requested video ID |
| X/Twitter | Existing Grok URL handler | Improvements deferred; YouTube is the priority |
| Articles/Blogs | readability-lxml + BeautifulSoup | Clean extraction, handles most sites |
| GitHub repos | GitHub API + README scrape | Structured repo info + documentation |
YouTube refuses automated extraction when no verifiable captions are available. Interactive use can accept a manually copied transcript.
YouTube metadata supplies the video title when available. A model's description of a URL alone is never accepted as the video's content.
Older URL-only Grok outputs with unverified content are preserved under extractions/quarantine/. Their index rows are marked Unverified; use a transcript-backed re-extraction before relying on them.
All extractions are tracked in extractions/INDEX.md:
| # | Title | Source | Category | Tags | Status | Date | File |
|---|---|---|---|---|---|---|---|
| 1 | Claude Code Tips | YouTube | Claude Code | #claude-code |
Backlog | 2025-01-15 | view |
| 2 | AI Agents Thread | Twitter/X | AI Agents | #agents |
TODO | 2025-01-16 | view |
Status flow: Backlog β TODO β In Progress β Done
Drop URLs from your phone β they get processed automatically.
Drop a URL in #extract from the Discord mobile app. MegaMind picks it up automatically. Results appear in the Forum with proper tagging.
- Open the repo in the GitHub mobile app
- Create a new issue β paste the URL as the title
- Add the
extractlabel - GitHub Actions processes it, commits the extraction, and closes the issue
Phone Cloud Desktop
βββββ βββββ βββββββ
Share URL
ββ Discord #extract βββββββ MegaMind bot βββββ Extraction
ββ GitHub App β Issue βββββ GitHub Actions committed
runs coord.py to repo
closes issue updates INDEX
posts to Forum
git pull
Obsidian sync
Pick up & implement
MegaMind automatically polls a YouTube playlist for new videos:
- Add videos to your "extract" playlist from any device
- MegaMind detects them on the next poll (default: every hour)
- Posts an audit trail message to #extract
- Extracts and processes the video
- Moves the video to a "completed" playlist (requires OAuth2)
OAuth2 setup (for playlist management):
python youtube_auth.pyMegaMind counts Codex subscription analyses separately from tracked Anthropic API calls:
- Anthropic API cost estimate from input/output tokens and model rate
- Codex run count without a fabricated API charge
- Last 100 entries in history
- Available via
/budget,/stats, or the dashboard footer
Data is persisted to api_budget.json and survives restarts.
| Channel | Type | ID | Purpose |
|---|---|---|---|
| #extract | Text | 1476145053721301149 |
INPUT β Drop URLs here from any device. YouTube watcher posts here as audit trail. |
| #output | Forum | 1478880776291487785 |
OUTPUT β Forum with 13 topic tags. Each extraction = tagged post. Filter by topic. |
| #general | Text | 1474002242175762549 |
General chat β OpenClaw responds here. |
| #ai-control | Text | 1474022861936132178 |
OpenClaw admin β model switching, status, skill management. |
| #testing | Text | 1474023043885174836 |
Experimentation β test both bots freely. |
Server: RussHub (1474002241319866439)
| Bot | ID | Channels |
|---|---|---|
| OpenClaw | 1474002760612708544 |
#general, #ai-control, #testing, #reports |
| MegaMind | 1476156237904085032 |
#extract, #output (Forum) |
When intake is enabled and an authorised user reacts with π€ on a MegaMind implementation prompt in the Forum:
- MegaMind detects the reaction
- Extracts the prompt text from the code block
- Records a
pending_confirmationintake with source IDs and an idempotency key - Appends an audit event locally
This does not execute code, create an issue or branch, or message anyone. A separate approval records the executor, target and bounded scope. Executor handoff still requires its own reviewed implementation.
Intakes are written to data/work-intakes.json and data/work-intake-audit.jsonl, which are ignored by Git. Keep their contents private.
MegaMind runs as a systemd user service on WSL2 Ubuntu:
# Service management
systemctl --user start megamind.service
systemctl --user stop megamind.service
systemctl --user restart megamind.service
systemctl --user status megamind.service
# View logs
journalctl --user -u megamind.service -f
# Service file location
~/.config/systemd/user/megamind.serviceThe service auto-starts on WSL boot (user linger enabled). The dashboard auto-starts as a subprocess.
MegaMind/
βββ coord.py # CLI entry point
βββ discord_bot.py # MegaMind Discord bot (Forum posting, auto-tagging)
βββ dashboard.py # Web dashboard (library + source review + Forum)
βββ dashboard_ui.html # Browser interface
βββ source_evidence.py # Private raw snapshots and review decisions
βββ forum_index.py # Cached Discord Forum post catalogue
βββ budget.py # API usage and cost tracking
βββ youtube_auth.py # YouTube OAuth2 setup helper
βββ config.py # Configuration (.env, paths, API keys)
βββ requirements.txt # Python dependencies
βββ .env # API keys and config (not committed)
βββ .env.example # Template for API keys
βββ .github/workflows/
β βββ extract.yml # GitHub Actions extraction workflow
βββ extractors/
β βββ detector.py # URL β source type detection
β βββ base.py # Base extractor interface
β βββ youtube.py # YouTube via captions or manual transcript
β βββ twitter.py # Twitter/X via Grok API
β βββ github.py # GitHub via API + scraping
β βββ article.py # Articles via readability + scraping
βββ processors/
β βββ ai_processor.py # Codex analysis + Claude API fallback
βββ outputs/
β βββ formatter.py # Discord embed + Forum post formatting
β βββ index.py # Central INDEX.md management
β βββ storage.py # File storage (repo + Obsidian)
βββ watchers/
β βββ youtube_playlist.py # YouTube playlist auto-watcher
βββ extractions/
β βββ INDEX.md # Centralised extraction tracker
βββ prompts/ # Session prompts for future work
βββ upgrade-audit-prompt.md
βββ forum-stats-prompt.md
βββ cross-bot-integration-prompt.md
| Component | Tool |
|---|---|
| Runtime | Python 3.11+ on WSL2 (Ubuntu) |
| Service | systemd user service (megamind.service) |
| Discord Bot | discord.py β watches #extract, posts to Forum with auto-tagging |
| YouTube Watcher | google-api-python-client β playlist polling + OAuth2 management |
| Grok/xAI | OpenAI-compatible client for X/Twitter threads |
| LLM | Codex GPT-6 Astra (subscription), Claude Opus 5.5 (API fallback) |
| Articles | readability-lxml + BeautifulSoup |
| Obsidian | File-based via WSL mount, synced via Obsidian Sync |
| Dashboard | Local Python HTTPServer + searchable browser interface |
| Budget | JSON-based tracking with per-model pricing |
MegaMind runs alongside OpenClaw (AI assistant bot) on the same Discord server:
- OpenClaw handles chat, model switching, skills (email, GitHub, Obsidian, weather, etc.)
- MegaMind handles content extraction and knowledge capture
- Both run as systemd user services on the same WSL2 instance
- Future: cross-bot integration (OpenClaw triggering extractions, shared search)
| Device | Input | View Output |
|---|---|---|
| Phone (Discord) | Drop URL in #extract | Browse Forum by tag |
| Work Desktop | Discord web + CLI | Forum + Obsidian |
| Home PC | CLI + Discord + vault | Full access |
| Any Device | Obsidian Sync | Read-only |
| Any Browser | GitHub repo | Read-only |
- Core extraction pipeline (YouTube, Twitter, GitHub, Article)
- AI processing with Claude (context-aware prompts)
- CLI tool (
coord.py) - Discord bot β watch #extract, post to #output
- YouTube playlist watcher with auto-extraction
- Obsidian vault + GitHub storage
- Central INDEX.md with status tracking
- Mobile capture (GitHub Actions)
- Controlled π€ reaction β pending local intake (disabled by default)
- API budget tracking
- Web dashboard (searchable library, status, source review, Forum activity)
- systemd user service (
megamind.service) - Discord Forum channel with 13 topic tags
- Multi-tag auto-categorisation (up to 5 tags per post)
- YouTube oEmbed title resolution
- Requester ID tracking (thread visibility fix)
-
/statscommand β Forum analytics (posts by tag, recent activity) - Improved
/searchβ Forum thread title and tag search - X/Twitter integration improvements β deferred while YouTube is prioritised
- Re-extraction β re-process existing URLs with updated AI processing
- OpenClaw skill integration β trigger extractions from any channel
- Cross-bot status awareness β each bot knows if the other is online
- Shared knowledge search β OpenClaw queries MegaMind's extraction index
- Risk classification per prompt (Low / Medium / High)
- OpenClaw dispatch β route low/medium prompts directly for execution
- PDF/DOCX document extraction
- Jina Reader API as alternative article extractor
- Python 3.11+
- WSL2 (Ubuntu) on Windows 11
- Anthropic API key (for Claude processing)
- xAI API key (optional, for X/Twitter extraction)
- Discord bot token (for MegaMind bot)
MIT β Russ Thompson