PETIT is a dependency-free C++17 inference engine for small decoder-only LLMs, engineered to use as little RAM, CPU and GPU as possible — it runs entirely on a laptop CPU with no GPU. Weights are block-quantized (q8_0 int8, or q4_0 4-bit) so they shrink up to ~7× and stay cache-resident.
It comes with a pure-NumPy trainer and an exporter that quantizes a
checkpoint into the engine's own binary format, so the full loop
train → quantize → serve runs end-to-end with one toolchain. The engine's
output is verified against an independent NumPy implementation by comparing
full logit vectors at every position (tools/verify.py).
| Lever | Design | pays off in |
|---|---|---|
| int8 block quantization (q8_0) | Every 32 weights share one fp32 scale; stored as int8. Dequantization happens on the fly; we never keep a full fp32 matmul copy. |
~4× less weight RAM |
| q4_0 (4-bit) quantization | Same block scheme but with a fp16 scale + 4-bit nibbles (18 bytes/block = 4.5 bits/weight). ~1.9× smaller than q8_0, ~7× smaller than fp32, at a small quality cost. | ~7× less weight RAM |
| Quantized GEMM | gemv_q8 has an AVX2 path (FMA dot product) and a portable scalar fallback chosen by -march=native; gemv_q4 uses the scalar fallback. |
fewer CPU cycles per token |
| KV-cache | Past attentions are cached; each new token is a single forward on the latest token, not a re-read of the context. | inference cost ∝ 1 token, not ∝ context² |
| RoPE + RMSNorm + SwiGLU | Llama-style block, no learned position embeddings. | smaller weights; no length dependence |
| Zero deps, single binary | petit.exe is a few hundred KB; nothing to pip install at runtime. |
runs literally anywhere |
The tiny model in this repo runs close to cache-resident on a typical laptop and generates thousands of tokens/s from a <0.05 MB q4 weight file (48 KB, ~4.4× smaller than its fp32 form).
src/
config.h hyperparameters + binary header (mirrors the exporter)
quant.h/.cpp q8_0 + q4_0 quant / dequant / dot-products (fp16 helpers)
matmul.h/.cpp quantized gemv (AVX2 for q8, scalar always for q4), gemm
model.h/.cpp forward pass: RMSNorm, RoPE, KV-cache, SwiGLU, sampling hooks
tokenizer.h/.cpp byte-level BPE (with optional merges file; char mode = no merges)
sampler.h/.cpp greedy / temperature / top-k / top-p
main.cpp CLI: load weights, tokenize, run generation
scripts/
train_tiny.py pure-NumPy GPT trainer (matches the engine exactly, with a
numeric gradient check --check)
corpus.txt small demo corpus
tools/
export_model.py numpy checkpoint -> .petit quantized binary
ref_infer.py NumPy reference inference (dequantizes .petit exactly)
verify.py compares engine logits vs ref_infer (CI-runnable)
models/
.npz (weights), .petit (quantized), vocab.txt (tokenizer)
.github/workflows/
ci.yml build + run verify.py on push to main / any PR
MinGW-w64 (g++ 16) or any C++17 compiler:
mingw32-make # Windows → petit.exe
make -C . CXX=g++ # POSIX: adapt, or
g++ -O2 -std=c++17 -march=native -o petit src/*.cpppip install numpy
# 1) train a tiny character-level LM (NumPy, CPU, ~5 min)
python scripts/train_tiny.py --corpus scripts/corpus.txt \
--embd 40 --layers 2 --heads 2 --ctx 32 --steps 3000 --out models/tiny.npz
# 2) quantize to the engine's binary (default q8_0; --q4 for 4-bit, --no-quant for fp32)
python tools/export_model.py models/tiny.npz models/tiny.petit \
--vocab 30 --embd 40 --layers 2 --heads 2 --ctx 256 --q4
# 3) generate
./petit -m models/tiny.petit -v models/vocab.txt \
-p "jack and jill went up the hill" -n 200 -t 0.8 -s 0.9-m <weights> quantized model (required)
-v <vocab> tokenizer vocab, one token per line (required)
-g <merges> optional BPE merges file ("left right" pairs, low rank first)
-p <prompt> prompt text
-n <tokens> number of tokens to generate (default 64)
-t <temp> temperature (0 => greedy) (default 0.8)
-k <top-k> top-k filter (default 0 = off)
-s <top-p> top-p (nucleus) (default 0 = off)
--seed <n> RNG seed
- Weights in the checkpoint and the
.petitfile are stored transposedout × in, and all norms are plain vectors, so the engine never needs to reshape anything — a projection is onegemv_q8(seesrc/model.cpp). - Quantization in
tools/export_model.pywrites q8_0 blocks identical tostruct QBlockinsrc/quant.h: a fp32 scale followed by 32 int8 values, zero-padded to a multiple of 32 so the AVX2 kernel needs no tail handling. With--q4it writesstruct Q4Block(a fp16 scale + 16 bytes of 4-bit nibbles per 32 weights); the engine picksgemv_q4/gemv_q8from the header'squantfield (0=fp32,1=q8_0,2=q4_0). - RoPE is parameterized
base=10000; it is applied per head. There are no learned absolute/relative embeddings, so context extension is parameter-free.
The repo ships two reference implementations that must agree with the C++ binary:
python tools/verify.py models/tiny.petit models/vocab.txt "the sun" -n 40verify.py runs the engine with --dump-logits and compares the full logit
vector to ref_infer at every position (worst error and argmax parity).
Comparing raw logits — not just greedy text — makes the check immune to a
single-ulp argmax tie flip, so VERIFY PASS is a strong, deterministic proof
that quantize, loader, forward-pass and (for --q4) the nibble decode agree.
Combined with train_tiny.py --check (a finite-difference test of every
gradient of the autodiff), this gives high confidence that trainer, quantizer,
loader and engine all speak one identical format.
The engine is size-agnostic: bump --embd/--layers in the exporter. This
codebase is a reference, so a realistic ~1B model is trained with a proper
accelerator and exported through export_model.py as-is — the engine will load
it without further changes (weights + int8 memory = the whole point).
- int4 (q4_0) blocks for another ~2× memory drop
- AVX2 dot-product for q4_0 (needs a packed shuffle to match the interleave)
- larger-
hiddenmatmul tiling/blocking for >192-dim models - streaming / sliding-window attention to hard-bound the cache
- true speculative decoding / batching for peak CPU utilization
Built as a single-from-scratch study in how far a small inference engine can go.