Aura is a zero-surface, multi-agent creative platform: a team of AI Virtuosos that turn your intent into generated media, interactive courses, and automation for the creative tools you already use — with deep multimodal understanding as the layer beneath, not the destination. Step into the conductor's shoes and let the Virtuosos do the heavy lifting.
📖 Full documentation: aura-symphony.netlify.app/docs
The complete docs hub covers getting started, architecture, the Virtuosos, Lenses, core patterns, backend services, adaptive learning, and operations — searchable, in-depth guides that go well beyond this README.
Aura is powered by a multi-agent architecture. Each agent, or "Virtuoso," is an expert in a specific domain:
- The Conductor: The orchestrator. It parses your intent and routes tasks.
- The Visionary: The visual analyst. It uses a multimodal's capabilities to extract deep insights from video frames and images.
- The Scholar: The researcher. It grounds analysis in real-world facts using Google Search.
- The Artisan: The creator. It uses Veo and Imagen to generate new media assets.
- The Analyst: The logician. It synthesizes data, creates structured courses, and tracks learning progress.
- The Chronicler: The documentarian. It summarizes sessions, generates TTS, and exports data.
- The Critic: The quality gate. It adversarially evaluates Virtuoso outputs on relevance, factual consistency, and quality — triggering retry loops when standards aren't met.
- Semantic Video Search: Ask the Conductor questions like "Where did they discuss the black hole?" and it will instantly drop markers on the timeline at the exact timestamps. Powered by adaptive semantic chunking and ChromaDB vector search.
- Streaming Responses: All Virtuoso outputs stream token-by-token via
generateContentStream, reducing perceived latency by 60-80%. - ReAct Planning: Complex multi-step queries are automatically routed through a Reason+Act loop — the Conductor plans, executes, observes, and adapts.
- Voice-Activated Conductor: Click the microphone icon or use the wake word to speak your commands naturally, achieving a true "zero-surface" experience.
- WebWorker Pool: Frame extraction and heavy processing are distributed across a pool of N workers with work-stealing for load balancing.
- NLE Integration: Export your timeline, annotations, and generated assets directly to Premiere Pro, Final Cut Pro, or DaVinci Resolve via FCPXML, EDL, or CSV.
- Agent Studio & Plugin Marketplace: Build custom AI agents or install third-party Virtuosos from the marketplace with sandboxed execution and SHA-256 integrity verification.
- CRDT Collaboration: Real-time multi-user editing via Yjs CRDTs with WebSocket transport, conflict-free shared state, and peer cursor visualization.
- Offline-First PWA: Service Worker caching + IndexedDB mutation queue enables full offline operation with background sync on reconnect.
- Multimodal RAG: CLIP ViT-B/32 frame embeddings enable true visual search alongside text-based retrieval, with fusion scoring across modalities.
Valhalla is Aura's bridge to the outside world: an LLM agent writes an automation script and a hardened in-browser WebAssembly sandbox runs it safely — then hands you the validated, tool-ready script to apply in your software. Letting a model generate code is easy; running it without giving an attacker or a hallucination the keys to the machine is the hard part — and that's the part Valhalla is built around.
- Defense in depth (5 layers): static analysis (regex + AST) → a 22-module runtime import firewall → WASM capability isolation (no network, no filesystem, no subprocess) → a 30s CPU timeout → a memory cap. Layers 2–5 are structural — they hold even if detection is fooled.
- Quantitative safety scoring:
100 − 30·critical − 10·warning; any critical finding blocks the script before it runs. A dedicated AST pass catches the obfuscated escapes (getattrdunder,globals(),__builtins__,eval→variable). - Tool-agnostic substrate: the target is a parameter, not hard-wired code — point it at any scriptable tool (Blender, Houdini, Ableton, Figma, DaVinci…). Each step is the same primitive — generate a vetted script for tool X — so they compose into multi-tool pipelines (e.g. course → Blender for 3D → Ableton Live for audio).
- PMDE + human override: API/scripting first, with the generated script, safety report, and live output surfaced for your approval at every step.
This is a proof of concept that's deliberately transparent about what's built versus what's next — the value is the architecture, and it's meant to grow with its users. Read the full deep-dive: PROJECT_VALHALLA.md · or in the docs at aura-symphony.netlify.app/docs/valhalla.
Creation extends to teaching. The Create Course lens authors a structured, interactive learning module from your material — turning a finished piece into a course others can learn from. As learners progress, Aura personalizes the path with Bayesian Knowledge Tracing (BKT) using temporal decay and prerequisite-aware content selection; the Digital Learner Profile (DLP) gives calibrated probability-of-mastery estimates, and Federated Learning with differential privacy lets the system improve across users while preserving individual privacy.
- Node.js 18+
- A model of your choice API key
-
Clone and install:
git clone <repo-url> && cd aura-symphony npm install
-
Create your environment file:
cp .env.example .env
-
Edit
.envand add your Gemini API key:GEMINI_API_KEY=AIzaSy...your_actual_key -
Start the dev server:
npm run dev
The app runs at http://localhost:3000.
Aura supports two ways to supply an AI provider key:
| Method | Description |
|---|---|
| In-App Settings (recommended) | Click the ⚙️ Settings icon in the toolbar to open the AI Provider Settings panel. From there you can add one or more providers, each with its own Base URL, API Key, and Model. The active provider is used for all AI calls. A Test Connection button validates the key and model before you commit. Settings are persisted in localStorage. |
.env file (fallback) |
Set GEMINI_API_KEY in a .env file at the project root. Vite injects it as process.env.API_KEY at build time via define in vite.config.ts. This key is used only when no custom provider is configured in Settings. |
Resolution order:
getAI()checks for an active provider with a non-empty API key first. If none exists, it falls back to the.env-based default client.getEffectiveModel(registryModel)similarly returns the user's custom model when set, or the per-virtuoso default otherwise.
Compatibility: The Settings panel supports any OpenAI-compatible API (Google AI, Anthropic via proxy, Ollama, local LLMs, etc.).
Aura is a polyglot microservices architecture with a rich React 19 + Vite 8 frontend. Seven AI agents communicate via the SymphonyBus (custom EventTarget-based event bus) with commission chaining for multi-step orchestration.
| Service | Stack | Port | Role |
|---|---|---|---|
| Frontend | React 19 + Vite 8 + TypeScript | 3000 | SPA, agent orchestration, frame extraction |
| API Proxy | Express | 3005 | Gemini key isolation, rate limiting, usage metering |
| Vector Search | FastAPI + ChromaDB | 3001 | Semantic search with adaptive chunking |
| Graph Knowledge | Express + SQLite | 4004 | Concept graph traversal, learning paths |
| Media Pipeline | Express + FFmpeg + WebSocket | 3002/3003 | Cloud-side frame extraction, transcription |
| CLIP Embeddings | FastAPI + CLIP ViT-B/32 | 3006 | Multimodal visual search |
All backend services are optional — the frontend degrades gracefully to browser-local alternatives. Orchestrated via docker-compose.yml.
Defense-in-depth AI safety: Zod schema validation on all LLM function calls → Critic agent adversarial quality gate → Valhalla 3-layer sandbox → Plugin sandboxed execution.
438 tests across 20 files. Production build in ~29s.
The workspace UI has been modernized with a comprehensive design system overhaul:
- Design Token System — 16 modular CSS files (all under 250 lines) with full token coverage: color (light/dark), spacing, radius, shadow, typography, transitions, z-index
- Light/Dark Theme — Brand dark theme default (matches marketing site), with localStorage persistence and a toggle button in the app header. Users can switch to light mode if preferred.
- Font Migration — Inter for body/UI text (replacing Space Mono), Space Mono reserved for timestamps, code blocks, and data tables only
- Icon Consolidation — All Material Symbols icons migrated to Lucide React (single icon system, reduced bundle size)
- Modal Consolidation — All modals use the shared accessible
Modal.tsxcomponent (focus trap, ARIA, escape, body scroll lock). Zero nativealert()/prompt()/confirm()calls remain - Toast Notifications — Success/error/info toasts with auto-dismiss, ARIA live regions, and keyboard support
- Skeleton Loading — Shimmer-animated skeletons replace spinners for content loading (insight cards, chat, library)
- Keyboard Shortcuts — Space (play/pause), J/K/L (seek), arrows (frame step),
/(focus conductor), Cmd+K (command palette) - Command Palette — Cmd/Ctrl+K fuzzy-searchable action launcher with keyboard navigation
- Responsive Design — 768px and 1024px breakpoints with drawer-mode Lens Laboratory and mobile reflow
- Lens Labels — Persistent labels under lens icons at wide breakpoints, tooltips for disabled lenses
See CHANGELOG.md for detailed implementation progress.
All Rights Reserved. See LICENSE for details. Unauthorized copying, modification, distribution, or use is strictly prohibited.