Skip to content

Repository files navigation

📚 Doubt.AI — Real-Time Online Study Doubt Solver

A premium, glassmorphic desktop assistant designed to help online students clear doubts instantly. With a simple global hotkey (Ctrl + Shift + H), the app captures a screenshot of your screen, dumps the last 1 minute of system audio (what the lecturer was saying), and streams a comprehensive, context-aware explanation using Gemini 2.5 Flash. Uniquely built as a zero-dependency, native C++ application, it requires no Python, Node.js, or complex runtime dependencies to operate.


✨ Key Features

  • 🔇 WASAPI Loopback Capture: Low-level, high-efficiency C++ background service captures speaker output (not microphone) in a rolling 60-second memory buffer. No virtual audio cables or drivers required.
  • 🖼️ Fast Screen Capture: GDI+ screen capturing captures the active screen and saves it as a lightweight PNG immediately on hotkey trigger.
  • ⚡ Hotkey Trigger Indicator: Double-beep audio confirmation lets you know your doubt has been captured.
  • 🪐 Glassmorphic Dashboard: A beautiful, modern dark-mode web application styled with vibrant gradients, glowing accent orbs, glass cards, and micro-animations.
  • 📝 Live Markdown Streaming: Real-time server-sent events (SSE) stream the AI tutor's explanation word-by-word, rendered natively in fully formatted Markdown.
  • 💬 Interactive Follow-Up: Chat back-and-forth with the AI to ask questions, clarify equations, or request code modifications.

🛠️ Architecture

graph TD
    User([User Press Hotkey]) -->|Ctrl + Shift + H| Cpp[C++ App: doubt_helper.exe]
    Cpp -->|1. Takes screenshot.png| Disk[(Local Disk)]
    Cpp -->|2. Dumps rolling 60s buffer as audio.wav| Disk
    Cpp -->|3. Broadcasts event: new_doubt| SSE[SSE Connection]
    SSE -->|Triggers UI reload| Web[Web Interface: templates/index.html]
    Web -->|POSTs history & prompt| Cpp
    Cpp -->|4. Multimodal HTTPS Payload| Gemini[Gemini 2.5 Flash API]
    Gemini -->|Streams text response| Cpp
    Cpp -->|SSE streams explanation| Web
Loading

🚀 Getting Started

Prerequisites

  1. Windows OS (required for GDI+ screenshotting, WASAPI capturing, and Win32 hotkeys).
  2. MinGW-W64 (g++) installed (only if you wish to recompile the C++ source). (Note: Python is no longer required!)

Installation & Setup

  1. Clone the Repository:

    git clone https://github.com/your-username/doubt-ai.git
    cd doubt-ai
  2. Configure Environment Variables: Create a .env file in the root directory (or copy .env.example):

    GEMINI_API_KEY=your_gemini_api_key_here
    PORT=5000
    HOTKEY_MODIFIERS=MOD_CONTROL|MOD_SHIFT
    HOTKEY_KEY=H

    Replace your_gemini_api_key_here with a valid key from Google AI Studio.

  3. Compile the Application (Optional — precompiled doubt_helper.exe is included): If you make changes to the C++ code, compile it using:

    compile.bat
  4. Launch the Application: Simply run the start script:

    run.bat

    This script will launch the unified C++ helper and web server directly in your terminal.


📖 How to Use

  1. Start run.bat and leave the terminal minimized.
  2. Go to your online lecture (e.g., YouTube, Zoom, Coursera, or a PDF textbook).
  3. The moment you have a doubt or the lecturer explains a complex topic:
    • Press Ctrl + Shift + H.
    • You will hear a double beep indicating successful capture.
  4. Your browser will automatically open (or refresh if already open) to http://localhost:5000.
  5. Review the captured screenshot and listen to the audio playback. The AI Doubt Tutor will automatically stream a detailed explanation matching what was spoken in the lecture with the slides on your screen!
  6. Type any follow-up questions in the input box at the bottom of the chat to interact with the tutor.

🎛️ Configuration

You can customize the global hotkey by editing the .env file:

  • Modifiers: Combines standard Windows hotkey modifiers: MOD_CONTROL, MOD_SHIFT, MOD_ALT, MOD_WIN (separated by |).
  • Key: Any alphabetical character (A-Z) or function keys (F1-F12).
    • Example: HOTKEY_KEY=D and HOTKEY_MODIFIERS=MOD_CONTROL|MOD_ALT registers Ctrl + Alt + D.

⚙️ Tech Stack

  • System APIs: C++, Windows API, GDI+ (Screenshot), WASAPI Loopback (System Audio Capture)
  • Web Server: Crow C++ Framework (v1.3.2), Standalone Asio (Header-only Networking)
  • Frontend: HTML5, Vanilla CSS3 (Custom Glassmorphism Design System), Javascript (ES6), Marked.js (Markdown rendering)
  • AI Integration: WinHTTP HTTPS Request Client, Google Gemini REST API (Gemini 2.5 Flash model)

📄 License

This project is licensed under the MIT License - see the LICENSE file for details.

About

Doubt.AI is a real-time AI-powered study assistant that captures screenshots and the last minute of lecture audio with a single hotkey to instantly solve academic doubts. Built entirely in native C++, it uses WASAPI, GDI+, and Gemini 2.5 Flash to provide context-aware explanations without requiring Python, Node.js, or external runtimes.

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages