Run powerful quantized LLMs 100% in your browser — no backend, no API keys, no cloud. All inference runs locally on your GPU via WebGPU and WebAssembly.
| Feature | Details |
|---|---|
| 🔒 100% Private | Zero network calls after model download. Your prompts never leave your machine. |
| ⚡ Q4F16 Quantized Models | 4-bit float-16 quantization delivers ~20% faster token generation vs. f32 baselines. |
| 🚫 Unrestricted Model | NeuralHermes (7B) — DPO-trained, minimal guardrails, explicit content capable. |
| 💬 Streaming Chat | Real-time token-by-token streaming with live tokens/sec metrics. |
| 🗂️ Persistent History | Multi-session chat history stored in browser IndexedDB — no account required. |
| 🗑️ Model Cache Manager | View cached models and delete them directly from the UI to free disk space. |
| ⚙️ Hardware Controls | Context window slider, WebGPU flag helper, and discrete GPU instructions. |
| Models View | Chats View | Performance Settings |
|---|---|---|
![]() |
![]() |
![]() |
| Model | ID | VRAM | Tag |
|---|---|---|---|
| Llama 3.1 8B | Llama-3.1-8B-Instruct-q4f16_1-MLC |
~5 GB | 🔵 Q4F16 Fast |
| Mistral 7B | Mistral-7B-Instruct-v0.3-q4f16_1-MLC |
~4.5 GB | 🔵 Q4F16 Fast |
| Hermes-3 3B | Hermes-3-Llama-3.2-3B-q4f16_1-MLC |
~2.3 GB | ⚫ Compact · Fast |
| NeuralHermes 7B | NeuralHermes-2.5-Mistral-7B-q4f16_1-MLC |
~4.6 GB | 🔴 ⚠ No Guardrails |
| Qwen2 1.5B | Qwen2-1.5B-Instruct-q4f16_1-MLC |
~1.5 GB | ⚫ Compact |
| Phi-3 Mini | Phi-3-mini-4k-instruct-q4f16_1-MLC |
~3 GB | ⚫ Compact |
| SmolLM2 135M | SmolLM2-135M-Instruct-q0f32-MLC |
~500 MB | ⚫ Ultra-Light |
⚠️ NeuralHermes has significantly reduced content restrictions. Use responsibly.
- A browser with WebGPU support (Chrome 113+ recommended)
- Enable the flag if needed:
chrome://flags/#enable-unsafe-webgpu - A GPU with at least 4 GB VRAM for 7B models (2 GB for compact models)
git clone https://github.com/Codesensitive/Local-Web-LLM.git
cd Local-Web-LLM/app
npm install
npm run devThen open http://localhost:5173 in Chrome.
Local-Web-LLM/
├── app/
│ └── src/
│ ├── components/
│ │ ├── ModelSidebar.tsx # Model registry, cache management, chat history
│ │ ├── ChatCanvas.tsx # Streaming chat UI with markdown rendering
│ │ └── SettingsModal.tsx # Context window & optimization controls
│ ├── hooks/
│ │ └── useLLM.ts # WebLLM engine lifecycle & streaming hook
│ └── utils/
│ ├── storage.ts # IndexedDB persistence layer
│ └── gpuCheck.ts # WebGPU capability detection
Tech stack: React 18 · TypeScript (strict) · Vite · Tailwind CSS · @mlc-ai/web-llm
Inference engine: @mlc-ai/web-llm — WebGPU-accelerated inference via precompiled MLC WASM shaders.
- Select a model from the sidebar — models are streamed from HuggingFace and cached in browser IndexedDB on first load.
- WebLLM initializes a WebGPU-backed engine in a Web Worker thread, keeping the UI thread responsive.
- Send a message — tokens are streamed back iteratively via the OpenAI-compatible chat completions API, updating the UI on each chunk.
- Chats are auto-saved to IndexedDB and restored on next visit.
- Use Chrome with hardware acceleration enabled
- For iGPU users, reduce context window size in ⚙️ Optimizations to lower VRAM usage
- Start with SmolLM2 135M or Qwen2 1.5B on lower-end hardware
- Force discrete GPU via Chrome flags for best throughput
MIT © Codesensitive


