Skip to content

Latest commit

 

History

6 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 

Repository files navigation

🧠 Local Web LLM

Run powerful quantized LLMs 100% in your browser — no backend, no API keys, no cloud. All inference runs locally on your GPU via WebGPU and WebAssembly.

React TypeScript Vite WebLLM License


✨ Features

Feature Details
🔒 100% Private Zero network calls after model download. Your prompts never leave your machine.
⚡ Q4F16 Quantized Models 4-bit float-16 quantization delivers ~20% faster token generation vs. f32 baselines.
🚫 Unrestricted Model NeuralHermes (7B) — DPO-trained, minimal guardrails, explicit content capable.
💬 Streaming Chat Real-time token-by-token streaming with live tokens/sec metrics.
🗂️ Persistent History Multi-session chat history stored in browser IndexedDB — no account required.
🗑️ Model Cache Manager View cached models and delete them directly from the UI to free disk space.
⚙️ Hardware Controls Context window slider, WebGPU flag helper, and discrete GPU instructions.

📸 Screenshots

Models View Chats View Performance Settings
Models View Chats View Settings

🤖 Available Models

Model ID VRAM Tag
Llama 3.1 8B Llama-3.1-8B-Instruct-q4f16_1-MLC ~5 GB 🔵 Q4F16 Fast
Mistral 7B Mistral-7B-Instruct-v0.3-q4f16_1-MLC ~4.5 GB 🔵 Q4F16 Fast
Hermes-3 3B Hermes-3-Llama-3.2-3B-q4f16_1-MLC ~2.3 GB ⚫ Compact · Fast
NeuralHermes 7B NeuralHermes-2.5-Mistral-7B-q4f16_1-MLC ~4.6 GB 🔴 ⚠ No Guardrails
Qwen2 1.5B Qwen2-1.5B-Instruct-q4f16_1-MLC ~1.5 GB ⚫ Compact
Phi-3 Mini Phi-3-mini-4k-instruct-q4f16_1-MLC ~3 GB ⚫ Compact
SmolLM2 135M SmolLM2-135M-Instruct-q0f32-MLC ~500 MB ⚫ Ultra-Light

⚠️ NeuralHermes has significantly reduced content restrictions. Use responsibly.


🚀 Getting Started

Prerequisites

  • A browser with WebGPU support (Chrome 113+ recommended)
  • Enable the flag if needed: chrome://flags/#enable-unsafe-webgpu
  • A GPU with at least 4 GB VRAM for 7B models (2 GB for compact models)

Install & Run

git clone https://github.com/Codesensitive/Local-Web-LLM.git
cd Local-Web-LLM/app
npm install
npm run dev

Then open http://localhost:5173 in Chrome.


🏗️ Architecture

Local-Web-LLM/
├── app/
│   └── src/
│       ├── components/
│       │   ├── ModelSidebar.tsx   # Model registry, cache management, chat history
│       │   ├── ChatCanvas.tsx     # Streaming chat UI with markdown rendering
│       │   └── SettingsModal.tsx  # Context window & optimization controls
│       ├── hooks/
│       │   └── useLLM.ts          # WebLLM engine lifecycle & streaming hook
│       └── utils/
│           ├── storage.ts         # IndexedDB persistence layer
│           └── gpuCheck.ts        # WebGPU capability detection

Tech stack: React 18 · TypeScript (strict) · Vite · Tailwind CSS · @mlc-ai/web-llm

Inference engine: @mlc-ai/web-llm — WebGPU-accelerated inference via precompiled MLC WASM shaders.


⚙️ How It Works

  1. Select a model from the sidebar — models are streamed from HuggingFace and cached in browser IndexedDB on first load.
  2. WebLLM initializes a WebGPU-backed engine in a Web Worker thread, keeping the UI thread responsive.
  3. Send a message — tokens are streamed back iteratively via the OpenAI-compatible chat completions API, updating the UI on each chunk.
  4. Chats are auto-saved to IndexedDB and restored on next visit.

🔧 Performance Tips

  • Use Chrome with hardware acceleration enabled
  • For iGPU users, reduce context window size in ⚙️ Optimizations to lower VRAM usage
  • Start with SmolLM2 135M or Qwen2 1.5B on lower-end hardware
  • Force discrete GPU via Chrome flags for best throughput

📄 License

MIT © Codesensitive

About

Run quantized LLMs (Llama, Mistral, Hermes, NeuralHermes) 100% in-browser via WebGPU. No backend, no API keys. Built with React, Vite, TypeScript and @mlc-ai/web-llm. Features q4f16 models, streaming chat, IndexedDB history, token/s metrics, and an unrestricted model.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages