Sai Coumar
This repository contains my configurations for locally hosted Sub-Agent AI. This repository is focused on providing local coding agents for frontier models to delegate work to in order to conserve token usage with expensive models. The docker compose brings up llama.cpp agents on separate GPUs.
- GPU1 (
llama-1, port 18081) runsQwen3-Coder-30B-A3Bfor fast code generation, 8 parallel slots at 4096 tokens each. This is the batch/sub-agent backend the MCP dispatcher drives, and it can achieve ~200 generated tok/s with parallel jobs on an A5000. - GPU2 (
llama-2, port 18082) runsQwen3.8-27Bas a single-session driving model for the Claude Code TUI:--parallel 1with a 131072-token window. This replaced the old Gemma-4-31B reviewer.
Note that both dispatcher roles (coder and review) now point at qwen3-coder on GPU1, so the review pass is no longer a second opinion from a different model. GPU3 is idle if you want to bring back a dedicated reviewer.
Qwen3.8's native window is 262144, but on a 24 GB A5000 you are limited by VRAM, not the model. Measured with q8_0 KV cache and -ub 2048:
-c |
VRAM used | Result |
|---|---|---|
| 131072 | 21.7 GB | works, ~2.8 GB headroom (current setting) |
| 163840 | 23.1 GB | works, ~1.4 GB headroom |
| 204800 | — | OOM allocating the 1920 MiB compute buffer |
128k fits at all only because of the model's hybrid attention (full_attention_interval = 4 — only every 4th of its 65 layers is full-attention). Dense KV at that depth would be ~17 GB.
If you raise -c, raise CLAUDE_CODE_MAX_CONTEXT_TOKENS in the settings file below to match. Claude Code assumes a 200k window for models it doesn't recognize, so without that variable auto-compact aims past what the server accepts and you hit a hard wall mid-session instead of a graceful compact.
llama-swap/qwen38-claudecode.jinja is the model's own chat template with one change. The stock template raises System message must be at the beginning. for any system/developer message after the leading block, and Claude Code sends one mid-conversation, so every request failed. The patched copy renders it as an inline system turn instead:
{%- if message.role == "system" or message.role == "developer" %}
{{- '<|im_start|>system\n' + content + '<|im_end|>' + '\n' }}Do not work around this with --chat-template chatml. That parses, but the model stops seeing tool definitions entirely and <think> tags leak into visible output, which makes it useless for Claude Code.
Authentication matters here. Setting ANTHROPIC_AUTH_TOKEN flips Claude Code into third-party "API Usage Billing" mode, which disables claude.ai account features — connectors are explicitly disabled, and Remote Control depends on the same account auth (it connects to a hardcoded wss://bridge.claudeusercontent.com, independent of ANTHROPIC_BASE_URL).
Since llama-server requires no authentication at all, the token is unnecessary. Leaving it out keeps your normal OAuth login active, so --remote-control and connectors keep working while inference goes to the local model.
mkdir -p ~/.claude && cat > ~/.claude/local-qwen.json <<'EOF'
{
"env": {
"ANTHROPIC_BASE_URL": "http://localhost:18082",
"ANTHROPIC_MODEL": "qwen3.8",
"ANTHROPIC_SMALL_FAST_MODEL": "qwen3.8",
"CLAUDE_CODE_MAX_CONTEXT_TOKENS": "131072",
"ANTHROPIC_CUSTOM_MODEL_OPTION": "qwen3.8",
"ANTHROPIC_CUSTOM_MODEL_OPTION_NAME": "Qwen3.8-27B (local, GPU2)",
"ANTHROPIC_CUSTOM_MODEL_OPTION_DESCRIPTION": "Local llama.cpp on GPU2, 128k ctx"
}
}
EOF
alias claude-local='claude --settings ~/.claude/local-qwen.json'
Deliberately omitted:
ANTHROPIC_AUTH_TOKEN— see above; adding it back disables Remote Control and connectors.CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC— in OAuth mode the calls to Anthropic are the account/control-plane traffic Remote Control needs.
ANTHROPIC_SMALL_FAST_MODEL is required, not optional: without it Claude Code reaches out to the real Anthropic API for background tasks like title generation.
Plain claude is unaffected and still uses your Anthropic account. claude-local is fully local.
/modelcannot list local and cloud models side by side. Nothing in Claude Code attaches a base URL to a model —ANTHROPIC_CUSTOM_MODEL_OPTIONonly adds a label to the picker, andmodelOverridesremaps names, not endpoints. Mixing the two in one session requires a gateway (e.g. LiteLLM) fronting both, withANTHROPIC_BASE_URLpointed at it.- First request after an idle period pays a ~40 s model load.
ttlon theqwen3.8entry is 1800 rather than 300 so this doesn't fire every time you stop to read something. - A 30B-class Q4 local model is a real step down from a frontier model on long agentic loops — expect weaker multi-step tool sequencing. Solid for single-file edits and Q&A.