Skip to content

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

PETIT — a tiny, CPU-first, quantized LLM inference engine

PETIT is a dependency-free C++17 inference engine for small decoder-only LLMs, engineered to use as little RAM, CPU and GPU as possible — it runs entirely on a laptop CPU with no GPU. Weights are block-quantized (q8_0 int8, or q4_0 4-bit) so they shrink up to ~7× and stay cache-resident.

It comes with a pure-NumPy trainer and an exporter that quantizes a checkpoint into the engine's own binary format, so the full loop train → quantize → serve runs end-to-end with one toolchain. The engine's output is verified against an independent NumPy implementation by comparing full logit vectors at every position (tools/verify.py).


What makes it fast & small

Lever Design pays off in
int8 block quantization (q8_0) Every 32 weights share one fp32 scale; stored as int8. Dequantization happens on the fly; we never keep a full fp32 matmul copy. ~4× less weight RAM
q4_0 (4-bit) quantization Same block scheme but with a fp16 scale + 4-bit nibbles (18 bytes/block = 4.5 bits/weight). ~1.9× smaller than q8_0, ~7× smaller than fp32, at a small quality cost. ~7× less weight RAM
Quantized GEMM gemv_q8 has an AVX2 path (FMA dot product) and a portable scalar fallback chosen by -march=native; gemv_q4 uses the scalar fallback. fewer CPU cycles per token
KV-cache Past attentions are cached; each new token is a single forward on the latest token, not a re-read of the context. inference cost ∝ 1 token, not ∝ context²
RoPE + RMSNorm + SwiGLU Llama-style block, no learned position embeddings. smaller weights; no length dependence
Zero deps, single binary petit.exe is a few hundred KB; nothing to pip install at runtime. runs literally anywhere

The tiny model in this repo runs close to cache-resident on a typical laptop and generates thousands of tokens/s from a <0.05 MB q4 weight file (48 KB, ~4.4× smaller than its fp32 form).


Repository layout

src/
  config.h        hyperparameters + binary header (mirrors the exporter)
  quant.h/.cpp   q8_0 + q4_0 quant / dequant / dot-products (fp16 helpers)
  matmul.h/.cpp  quantized gemv (AVX2 for q8, scalar always for q4), gemm
  model.h/.cpp   forward pass: RMSNorm, RoPE, KV-cache, SwiGLU, sampling hooks
  tokenizer.h/.cpp byte-level BPE (with optional merges file; char mode = no merges)
  sampler.h/.cpp greedy / temperature / top-k / top-p
  main.cpp       CLI: load weights, tokenize, run generation
scripts/
  train_tiny.py  pure-NumPy GPT trainer (matches the engine exactly, with a
                 numeric gradient check --check)
  corpus.txt     small demo corpus
tools/
  export_model.py  numpy checkpoint -> .petit quantized binary
  ref_infer.py     NumPy reference inference (dequantizes .petit exactly)
  verify.py        compares engine logits vs ref_infer (CI-runnable)
models/
  .npz (weights), .petit (quantized), vocab.txt (tokenizer)
.github/workflows/
  ci.yml          build + run verify.py on push to main / any PR

Build

MinGW-w64 (g++ 16) or any C++17 compiler:

mingw32-make          # Windows → petit.exe
make -C . CXX=g++     # POSIX: adapt, or
g++ -O2 -std=c++17 -march=native -o petit src/*.cpp

End-to-end quickstart

pip install numpy

# 1) train a tiny character-level LM (NumPy, CPU, ~5 min)
python scripts/train_tiny.py --corpus scripts/corpus.txt \
    --embd 40 --layers 2 --heads 2 --ctx 32 --steps 3000 --out models/tiny.npz

# 2) quantize to the engine's binary (default q8_0; --q4 for 4-bit, --no-quant for fp32)
python tools/export_model.py models/tiny.npz models/tiny.petit \
    --vocab 30 --embd 40 --layers 2 --heads 2 --ctx 256 --q4

# 3) generate
./petit -m models/tiny.petit -v models/vocab.txt \
        -p "jack and jill went up the hill" -n 200 -t 0.8 -s 0.9

CLI options

-m <weights>   quantized model (required)
-v <vocab>     tokenizer vocab, one token per line (required)
-g <merges>    optional BPE merges file ("left right" pairs, low rank first)
-p <prompt>    prompt text
-n <tokens>    number of tokens to generate  (default 64)
-t <temp>      temperature (0 => greedy)      (default 0.8)
-k <top-k>     top-k filter                    (default 0 = off)
-s <top-p>     top-p (nucleus)                 (default 0 = off)
--seed <n>     RNG seed

How the pieces match

  • Weights in the checkpoint and the .petit file are stored transposed out × in, and all norms are plain vectors, so the engine never needs to reshape anything — a projection is one gemv_q8 (see src/model.cpp).
  • Quantization in tools/export_model.py writes q8_0 blocks identical to struct QBlock in src/quant.h: a fp32 scale followed by 32 int8 values, zero-padded to a multiple of 32 so the AVX2 kernel needs no tail handling. With --q4 it writes struct Q4Block (a fp16 scale + 16 bytes of 4-bit nibbles per 32 weights); the engine picks gemv_q4/gemv_q8 from the header's quant field (0=fp32, 1=q8_0, 2=q4_0).
  • RoPE is parameterized base=10000; it is applied per head. There are no learned absolute/relative embeddings, so context extension is parameter-free.

Verification

The repo ships two reference implementations that must agree with the C++ binary:

python tools/verify.py models/tiny.petit models/vocab.txt "the sun" -n 40

verify.py runs the engine with --dump-logits and compares the full logit vector to ref_infer at every position (worst error and argmax parity). Comparing raw logits — not just greedy text — makes the check immune to a single-ulp argmax tie flip, so VERIFY PASS is a strong, deterministic proof that quantize, loader, forward-pass and (for --q4) the nibble decode agree. Combined with train_tiny.py --check (a finite-difference test of every gradient of the autodiff), this gives high confidence that trainer, quantizer, loader and engine all speak one identical format.

Try at ~100M–1B

The engine is size-agnostic: bump --embd/--layers in the exporter. This codebase is a reference, so a realistic ~1B model is trained with a proper accelerator and exported through export_model.py as-is — the engine will load it without further changes (weights + int8 memory = the whole point).

Roadmap (small-to-medium, later)

  • int4 (q4_0) blocks for another ~2× memory drop
  • AVX2 dot-product for q4_0 (needs a packed shuffle to match the interleave)
  • larger-hidden matmul tiling/blocking for >192-dim models
  • streaming / sliding-window attention to hard-bound the cache
  • true speculative decoding / batching for peak CPU utilization

Built as a single-from-scratch study in how far a small inference engine can go.

About

PETIT is a dependency-free C++17 inference engine for small decoder-only LLMs, engineered to use as little **RAM, CPU and GPU** as possible it runs entirely on a laptop CPU with no GPU.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages