Skip to content

Repository files navigation

Locally Hosted Sub-Agent AI

Sai Coumar

This repository contains my configurations for locally hosted Sub-Agent AI. This repository is focused on providing local coding agents for frontier models to delegate work to in order to conserve token usage with expensive models. The docker compose brings up llama.cpp agents on separate GPUs.

  • GPU1 (llama-1, port 18081) runs Qwen3-Coder-30B-A3B for fast code generation, 8 parallel slots at 4096 tokens each. This is the batch/sub-agent backend the MCP dispatcher drives, and it can achieve ~200 generated tok/s with parallel jobs on an A5000.
  • GPU2 (llama-2, port 18082) runs Qwen3.8-27B as a single-session driving model for the Claude Code TUI: --parallel 1 with a 131072-token window. This replaced the old Gemma-4-31B reviewer.

Note that both dispatcher roles (coder and review) now point at qwen3-coder on GPU1, so the review pass is no longer a second opinion from a different model. GPU3 is idle if you want to bring back a dedicated reviewer.

Context sizing

Qwen3.8's native window is 262144, but on a 24 GB A5000 you are limited by VRAM, not the model. Measured with q8_0 KV cache and -ub 2048:

-c VRAM used Result
131072 21.7 GB works, ~2.8 GB headroom (current setting)
163840 23.1 GB works, ~1.4 GB headroom
204800 — OOM allocating the 1920 MiB compute buffer

128k fits at all only because of the model's hybrid attention (full_attention_interval = 4 — only every 4th of its 65 layers is full-attention). Dense KV at that depth would be ~17 GB.

If you raise -c, raise CLAUDE_CODE_MAX_CONTEXT_TOKENS in the settings file below to match. Claude Code assumes a 200k window for models it doesn't recognize, so without that variable auto-compact aims past what the server accepts and you hit a hard wall mid-session instead of a graceful compact.

Patched chat template

llama-swap/qwen38-claudecode.jinja is the model's own chat template with one change. The stock template raises System message must be at the beginning. for any system/developer message after the leading block, and Claude Code sends one mid-conversation, so every request failed. The patched copy renders it as an inline system turn instead:

{%- if message.role == "system" or message.role == "developer" %}
    {{- '<|im_start|>system\n' + content + '<|im_end|>' + '\n' }}

Do not work around this with --chat-template chatml. That parses, but the model stops seeing tool definitions entirely and <think> tags leak into visible output, which makes it useless for Claude Code.

Integration with Claude Code

Authentication matters here. Setting ANTHROPIC_AUTH_TOKEN flips Claude Code into third-party "API Usage Billing" mode, which disables claude.ai account features — connectors are explicitly disabled, and Remote Control depends on the same account auth (it connects to a hardcoded wss://bridge.claudeusercontent.com, independent of ANTHROPIC_BASE_URL).

Since llama-server requires no authentication at all, the token is unnecessary. Leaving it out keeps your normal OAuth login active, so --remote-control and connectors keep working while inference goes to the local model.

mkdir -p ~/.claude && cat > ~/.claude/local-qwen.json <<'EOF'
{
  "env": {
    "ANTHROPIC_BASE_URL": "http://localhost:18082",
    "ANTHROPIC_MODEL": "qwen3.8",
    "ANTHROPIC_SMALL_FAST_MODEL": "qwen3.8",
    "CLAUDE_CODE_MAX_CONTEXT_TOKENS": "131072",
    "ANTHROPIC_CUSTOM_MODEL_OPTION": "qwen3.8",
    "ANTHROPIC_CUSTOM_MODEL_OPTION_NAME": "Qwen3.8-27B (local, GPU2)",
    "ANTHROPIC_CUSTOM_MODEL_OPTION_DESCRIPTION": "Local llama.cpp on GPU2, 128k ctx"
  }
}
EOF
alias claude-local='claude --settings ~/.claude/local-qwen.json'

Deliberately omitted:

  • ANTHROPIC_AUTH_TOKEN — see above; adding it back disables Remote Control and connectors.
  • CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC — in OAuth mode the calls to Anthropic are the account/control-plane traffic Remote Control needs.

ANTHROPIC_SMALL_FAST_MODEL is required, not optional: without it Claude Code reaches out to the real Anthropic API for background tasks like title generation.

Plain claude is unaffected and still uses your Anthropic account. claude-local is fully local.

Known limits

  • /model cannot list local and cloud models side by side. Nothing in Claude Code attaches a base URL to a model — ANTHROPIC_CUSTOM_MODEL_OPTION only adds a label to the picker, and modelOverrides remaps names, not endpoints. Mixing the two in one session requires a gateway (e.g. LiteLLM) fronting both, with ANTHROPIC_BASE_URL pointed at it.
  • First request after an idle period pays a ~40 s model load. ttl on the qwen3.8 entry is 1800 rather than 300 so this doesn't fire every time you stop to read something.
  • A 30B-class Q4 local model is a real step down from a frontier model on long agentic loops — expect weaker multi-step tool sequencing. Solid for single-file edits and Q&A.

About

Locally Hosted Sub-Agent AI

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages