Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 2 additions & 1 deletion NOTICE
Original file line number Diff line number Diff line change
Expand Up @@ -9,7 +9,8 @@ regenerate this file by running:
The following third-party licenses are included in this repository:

deploy/helm/container-cache/NOTICE
deploy/helm/llm-routing/spark/NOTICE
deploy/helm/llm-routing/recipes/glm-5.3/NOTICE
deploy/helm/llm-routing/recipes/NOTICE
deploy/helm/nats/NOTICE
deploy/helm/openbao/NOTICE
infra/openbao/plugins/vault-plugin-secrets-jwt/NOTICE
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -47,7 +47,7 @@ clusterCredential:
# sha256: <output of: printf '%s' "$API_KEY" | shasum -a 256>
apiKeys: []

# Optional plaintext UI key, supplied from a private file by the Spark recipe.
# Optional plaintext UI key, supplied from a private file by the routing recipe.
# Creates Secret demo-ui-api-key (field api-key) in this release's namespace.
# apiKeys must also contain its matching hash under id demo-ui.
demoUiApiKey: ""
Expand Down
4 changes: 2 additions & 2 deletions deploy/helm/llm-routing/AGENTS.md
Original file line number Diff line number Diff line change
@@ -1,7 +1,7 @@
# LLM routing stack

This directory deploys the LLM routing stack and GLM on DGX Spark. Read `README.md` and `spark/AGENTS.md`. Edit runtime code in its owning source directories in the current checkout. Builds include local edits. Commit IDs are informational. Image updates compare routing chart contents with the fingerprint recorded in the installed stack.
This directory deploys the LLM routing stack and a model recipe on ARM64 NVIDIA GPU clusters, such as DGX Spark and GB300. Read `README.md` and `recipes/AGENTS.md`. Edit runtime code in its owning source directories in the current checkout. Builds include local edits. Commit IDs are informational. Image updates compare routing chart contents with the fingerprint recorded in the installed stack.

Run the Python tests and offline Helm render documented in the README. Always specify the Kubernetes context. Packaging validation must not change a live model deployment.

Keep mocks under `spark/tests` and out of the default deployment. Keep credentials, private targets, kubeconfigs and evidence in an external work directory. Do not publish private configuration overlays.
Keep mocks under `recipes/tests` and out of the default deployment. Keep credentials, private targets, kubeconfigs and evidence in an external work directory. Do not publish private configuration overlays.
171 changes: 107 additions & 64 deletions deploy/helm/llm-routing/README.md

Large diffs are not rendered by default.

11 changes: 11 additions & 0 deletions deploy/helm/llm-routing/recipes/AGENTS.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,11 @@
# LLM routing recipe tool

This entry point uses the checkout containing `recipe.py`. Builds include local edits and record the current commit ID for debugging. Image updates and rollback require the routing chart fingerprint recorded at installation. Helm dependencies are built automatically. `--source-dir` selects another existing checkout. Keep test mocks under `tests` and out of the normal model deployment.

Keep environment-specific values, credentials, kubeconfigs, generated TLS material and runtime evidence outside this repository. Pass an explicit context on every Kubernetes and Helm command. Do not change a live cluster while testing packaging.

Keep model-specific settings in the recipe folder (`recipe.json`, `model.lock.json`, `NOTICE`), not in `recipe.py` or the charts. Keep `sizing.py` free of cluster access. Placement must work for one model node and for a split across nodes; cover both in `tests/test_topology.py`.

Run `python3 -m unittest discover -s tests -v` and the chart tests under `charts/gguf-backend/tests`. Run `python3 recipe.py --config config.example.json --work-dir /tmp/recipe-render render` for offline chart validation. A render does not establish a fresh-cluster deployment.

Land runtime and chart fixes in their owning source directories with generated API files and tests. Never hand-edit generated CRDs or deepcopy code.
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
# Build application images

Build gateway, router, Pylon and operator images for `linux/arm64` from your checkout, including local edits. Run these commands from `deploy/helm/llm-routing/spark`.
Build gateway, router, Pylon and operator images for `linux/arm64` from your checkout, including local edits. Run these commands from `deploy/helm/llm-routing/recipes`.

## Requirements

Expand All @@ -14,21 +14,21 @@ Router and Pylon builds use Cargo profile `integration`.
1. Build all four images.

```bash
python3 spark.py build-images
python3 recipe.py build-images
```

2. Choose the distribution method.

- Registry: set `images.pullPolicy` to `IfNotPresent`, authenticate Docker with registry write credentials, then push. Every node where Pylon can schedule needs registry pull access.

```bash
python3 spark.py push-images
python3 recipe.py push-images
```

- Node preload: set `images.pullPolicy` to `Never`, export the images, then [import the archive](#import-an-archive).

```bash
python3 spark.py export-images
python3 recipe.py export-images
```

3. Continue with [Deploy in order](../README.md#deploy-in-order).
Expand All @@ -40,21 +40,21 @@ Router and Pylon builds use Cargo profile `integration`.
```bash
COMPONENT=gateway # Or router.
NEW_TAG=dev-$(date -u +%Y%m%d%H%M%S)
python3 spark.py build-images --component "$COMPONENT" --tag "$NEW_TAG"
python3 recipe.py build-images --component "$COMPONENT" --tag "$NEW_TAG"
```

2. Distribute the image using your existing method.

- Registry:

```bash
python3 spark.py push-images --component "$COMPONENT" --tag "$NEW_TAG"
python3 recipe.py push-images --component "$COMPONENT" --tag "$NEW_TAG"
```

- Node preload: export the image, then [import the archive](#import-an-archive).

```bash
python3 spark.py export-images --component "$COMPONENT" --tag "$NEW_TAG"
python3 recipe.py export-images --component "$COMPONENT" --tag "$NEW_TAG"
```

3. Keep `COMPONENT` and `NEW_TAG` set and continue with [Update only gateway or router](../README.md#update-only-gateway-or-router).
Expand All @@ -66,13 +66,13 @@ After `export-images`, upload and import the archive through Kubernetes. Keep ea
For all four images:

```bash
python3 spark.py import-images --allow-containerd-import
python3 recipe.py import-images --allow-containerd-import
```

For a gateway/router rebuild, import only that component on the control node:

```bash
python3 spark.py import-images \
python3 recipe.py import-images \
--component "$COMPONENT" --tag "$NEW_TAG" --allow-containerd-import
```

Expand Down
7 changes: 7 additions & 0 deletions deploy/helm/llm-routing/recipes/NOTICE
Original file line number Diff line number Diff line change
@@ -0,0 +1,7 @@
# External artifacts

The recipe tool downloads and builds llama.cpp from a pinned public revision. llama.cpp is MIT licensed. Its LICENSE is retained in the generated runtime archive. Source: https://github.com/ggml-org/llama.cpp.

Model weights are external downloads and are not distributed in this repository. Each recipe folder has a NOTICE for its model terms.

NVIDIA container images retain their upstream terms. Access to a public image reference does not grant credentials or image redistribution rights. No third-party runtime binaries, images or model weights are embedded in this tool.
3 changes: 3 additions & 0 deletions deploy/helm/llm-routing/recipes/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,3 @@
# Recipe tool implementation

Use the [LLM routing recipes runbook](../README.md). This directory contains the `recipe.py` tool, its model charts, client and regression tests. Each subfolder with a `recipe.json`, such as `glm-5.3`, is one recipe.
Original file line number Diff line number Diff line change
Expand Up @@ -6,8 +6,8 @@
import subprocess
import time

endpoints = os.environ['RPC_ENDPOINTS'].split(',')
assert len(endpoints) == 2
endpoints = [item for item in os.environ['RPC_ENDPOINTS'].split(',') if item]
assert endpoints, 'At least one RPC endpoint is required'
for endpoint in endpoints:
host, port = endpoint.rsplit(':', 1)
for attempt in range(120):
Expand All @@ -19,6 +19,10 @@
if attempt == 119:
raise
time.sleep(2)
subprocess.run(['/artifacts/runtime/test-rpc-multi-server', *endpoints], check=True, timeout=60)
tests = []
if len(endpoints) > 1:
subprocess.run(['/artifacts/runtime/test-rpc-multi-server', *endpoints], check=True, timeout=60)
tests.append('upstream-rpc-buffer-isolation')
subprocess.run(['/artifacts/runtime/rpc-gpu-check', *endpoints], check=True, timeout=120)
print(json.dumps({'result': 'PASS', 'transport': 'TCP', 'tests': ['upstream-rpc-buffer-isolation', 'two-gpu-f32-matmul-three-repeats']}), flush=True)
tests.append(str(len(endpoints)) + '-gpu-f32-matmul-three-repeats')
print(json.dumps({'result': 'PASS', 'transport': 'TCP', 'gpus': len(endpoints), 'tests': tests}), flush=True)
Original file line number Diff line number Diff line change
Expand Up @@ -12,7 +12,7 @@
#include <vector>

int main(int argc, char ** argv) {
if (argc != 3) throw std::runtime_error("Two RPC endpoints are required");
if (argc < 2) throw std::runtime_error("At least one RPC endpoint is required");
constexpr int k = 256, m = 128, n = 64;
std::vector<float> a(k*m), b(k*n), expected(m*n), actual(m*n);
for (int i = 0; i < k*m; ++i) a[i] = float(i % 17 - 8) / 8;
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -19,18 +19,23 @@
manifest = json.loads((root / 'runtime/build-manifest.json').read_text())
binary = root / 'runtime/llama-server'
assert hashlib.file_digest(binary.open('rb'), 'sha256').hexdigest() == manifest['binaries']['llama-server']
host, port = os.environ['RPC_ENDPOINT'].rsplit(':', 1)
for attempt in range(120):
try:
with socket.create_connection((host, int(port)), timeout=2):
pass
break
except OSError:
if attempt == 119:
raise
time.sleep(2)
# Empty for a single model node; otherwise one RPC worker endpoint per additional node.
endpoints = [item for item in os.environ.get('RPC_ENDPOINTS', '').split(',') if item]
for endpoint in endpoints:
host, port = endpoint.rsplit(':', 1)
for attempt in range(120):
try:
with socket.create_connection((host, int(port)), timeout=2):
pass
break
except OSError:
if attempt == 119:
raise
time.sleep(2)
args = [str(binary), '--model', str(root / 'model' / os.environ['FIRST_SHARD']),
'--alias', os.environ['SERVED_MODEL'], '--host', '0.0.0.0', '--port', '8000',
'--rpc', os.environ['RPC_ENDPOINT'], *json.loads(os.environ['SERVER_ARGS'])]
'--alias', os.environ['SERVED_MODEL'], '--host', '0.0.0.0', '--port', '8000']
if endpoints:
args += ['--rpc', ','.join(endpoints)]
args += json.loads(os.environ['SERVER_ARGS'])
print(json.dumps({'verifiedModel': verified, 'runtimeRevision': manifest['revision'], 'command': args}), flush=True)
raise SystemExit(supervise(args))
Original file line number Diff line number Diff line change
Expand Up @@ -29,7 +29,8 @@ data:
apiVersion: batch/v1
kind: Job
metadata:
name: {{ .Release.Name }}-build-{{ .Files.Get "files/build.py" | sha256sum | trunc 8 }}
# Job templates are immutable, so any change to the build inputs needs a new Job name.
name: {{ .Release.Name }}-build-{{ list (.Files.Get "files/build.py") .Values.build.revision .Values.build.cudaArchitectures | join "\n" | sha256sum | trunc 8 }}
spec:
backoffLimit: 0
activeDeadlineSeconds: 7200
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -47,7 +47,7 @@ spec:
- {name: GGML_RPC_NO_RDMA, value: '1'}
- {name: GGML_SCHED_DEBUG, value: '1'}
- {name: LLAMA_REVISION, value: {{ .Values.build.revision | quote }}}
- {name: RPC_ENDPOINTS, value: "{{ .Values.chain.runtimeRelease }}-rpc-leader:50052,{{ .Values.chain.runtimeRelease }}-rpc-worker:50052"}
- {name: RPC_ENDPOINTS, value: "{{ range $i, $target := .Values.targets }}{{ if $i }},{{ end }}{{ $.Values.chain.runtimeRelease }}-rpc-{{ $target.id }}:50052{{ end }}"}
resources:
requests: {cpu: '1', memory: 512Mi}
limits: {cpu: '2', memory: 2Gi}
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -41,7 +41,7 @@ spec:
- {name: NVIDIA_VISIBLE_DEVICES, value: none}
- {name: NVIDIA_DRIVER_CAPABILITIES, value: 'compute,utility'}
- {name: GGML_RPC_NO_RDMA, value: '1'}
- {name: RPC_ENDPOINTS, value: "{{ .Release.Name }}-rpc-leader:50052,{{ .Release.Name }}-rpc-worker:50052"}
- {name: RPC_ENDPOINTS, value: "{{ range $i, $target := .Values.targets }}{{ if $i }},{{ end }}{{ $.Release.Name }}-rpc-{{ $target.id }}:50052{{ end }}"}
resources:
requests: {cpu: '1', memory: 512Mi}
limits: {cpu: '2', memory: 2Gi}
Expand Down
Original file line number Diff line number Diff line change
@@ -1,21 +1,24 @@
{{/* SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved. */}}
{{/* SPDX-License-Identifier: Apache-2.0 */}}
{{- if has .Values.phase (list "qualify" "download" "serve") }}
{{- $leader := (index (required "targets are required" .Values.targets) 0).id }}
{{- if .Values.rpc.cache.enabled }}
{{- range $target := rest .Values.targets }}
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: {{ .Release.Name }}-rpc-cache
name: {{ $.Release.Name }}-rpc-cache-{{ $target.id }}
annotations:
helm.sh/resource-policy: keep
spec:
accessModes: [ReadWriteOnce]
storageClassName: {{ .Values.rpc.cache.storageClassName | quote }}
storageClassName: {{ $.Values.rpc.cache.storageClassName | quote }}
resources:
requests:
storage: {{ .Values.rpc.cache.size }}
storage: {{ $.Values.rpc.cache.size }}
---
{{- end }}
{{- end }}
apiVersion: apps/v1
kind: Deployment
metadata:
Expand Down Expand Up @@ -79,7 +82,7 @@ spec:
podSelector:
matchLabels: {app.kubernetes.io/instance: {{ .Release.Name }}}
matchExpressions:
- {key: app.kubernetes.io/component, operator: In, values: [artifacts, rpc-leader, rpc-worker]}
- {key: app.kubernetes.io/component, operator: In, values: [artifacts{{ range .Values.targets }}, rpc-{{ .id }}{{ end }}]}
policyTypes: [Ingress]
ingress:
- from:
Expand All @@ -101,7 +104,7 @@ data:
fetch-runtime.py: |
{{ .Files.Get "files/fetch-runtime.py" | indent 4 }}
{{- range $target := .Values.targets }}
{{- if or (ne $.Values.phase "serve") (eq $target.id "worker") }}
{{- if or (ne $.Values.phase "serve") (ne $target.id $leader) }}
---
apiVersion: apps/v1
kind: Deployment
Expand Down Expand Up @@ -163,12 +166,12 @@ spec:
- '50052'
- --device
- CUDA0
{{- if and $.Values.rpc.cache.enabled (eq $target.id "worker") }}
{{- if and $.Values.rpc.cache.enabled (ne $target.id $leader) }}
- --cache
{{- end }}
env:
- {name: HOME, value: /tmp}
{{- if and $.Values.rpc.cache.enabled (eq $target.id "worker") }}
{{- if and $.Values.rpc.cache.enabled (ne $target.id $leader) }}
- {name: LLAMA_CACHE, value: /rpc-cache}
{{- end }}
{{- range $name, $value := $.Values.runtime.env }}
Expand All @@ -191,7 +194,7 @@ spec:
- {name: runtime, mountPath: /work, readOnly: true}
- {name: tmp, mountPath: /tmp}
- {name: checks, mountPath: /checks, readOnly: true}
{{- if and $.Values.rpc.cache.enabled (eq $target.id "worker") }}
{{- if and $.Values.rpc.cache.enabled (ne $target.id $leader) }}
- {name: rpc-cache, mountPath: /rpc-cache}
{{- end }}
volumes:
Expand All @@ -201,9 +204,9 @@ spec:
configMap: {name: {{ $.Release.Name }}-runtime}
- name: tmp
emptyDir: {sizeLimit: 2Gi}
{{- if and $.Values.rpc.cache.enabled (eq $target.id "worker") }}
{{- if and $.Values.rpc.cache.enabled (ne $target.id $leader) }}
- name: rpc-cache
persistentVolumeClaim: {claimName: {{ $.Release.Name }}-rpc-cache}
persistentVolumeClaim: {claimName: {{ $.Release.Name }}-rpc-cache-{{ $target.id }}}
{{- end }}
---
apiVersion: v1
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -18,7 +18,7 @@ kind: Deployment
metadata:
name: {{ .Release.Name }}
spec:
replicas: 1
replicas: {{ .Values.model.replicas }}
progressDeadlineSeconds: 3660
strategy: {type: Recreate}
selector:
Expand Down Expand Up @@ -51,7 +51,7 @@ spec:
- {name: PYTHONDONTWRITEBYTECODE, value: '1'}
- {name: FIRST_SHARD, value: {{ required "model.firstShard is required" .Values.model.firstShard | quote }}}
- {name: SERVED_MODEL, value: {{ required "model.servedName is required" .Values.model.servedName | quote }}}
- {name: RPC_ENDPOINT, value: "{{ .Release.Name }}-rpc-worker:50052"}
- {name: RPC_ENDPOINTS, value: "{{ range $i, $target := rest .Values.targets }}{{ if $i }},{{ end }}{{ $.Release.Name }}-rpc-{{ $target.id }}:50052{{ end }}"}
- {name: SERVER_ARGS, value: {{ toJson .Values.model.args | quote }}}
{{- range $name, $value := .Values.runtime.env }}
- {name: {{ $name }}, value: {{ $value | quote }}}
Expand Down Expand Up @@ -103,6 +103,6 @@ spec:
canary: {{ toJson . }}
{{- end }}
maxEngineConcurrency: 1
gpu: {product: NVIDIA-GB10}
gpu: {product: {{ required "gpu.product is required" .Values.gpu.product | quote }}}
{{- end }}
{{- end }}
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,10 @@
phase: preflight
image: ""
runtimeClassName: nvidia
# targets[0] is the leader that runs the model server; later targets run RPC workers.
targets: []
gpu:
product: ""
artifacts:
storageClassName: local-path
size: 400Gi
Expand All @@ -28,6 +31,7 @@ rpc:
qualification:
attempt: 1
model:
replicas: 1
lock: {}
servedName: ""
firstShard: ""
Expand Down
Original file line number Diff line number Diff line change
@@ -1,7 +1,7 @@
#!/usr/bin/env python3
# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
"""Chat with GLM and validate direct or authenticated gateway responses without logging keys."""
"""Chat with the recipe model and validate direct or authenticated gateway responses without logging keys."""
import argparse
import datetime
import http.client
Expand Down Expand Up @@ -168,7 +168,7 @@ def main():
parser.add_argument('--url', default='https://127.0.0.1:18443')
parser.add_argument('--ca-file')
parser.add_argument('--api-key-file')
parser.add_argument('--model', default='GLM-5.3-UD-IQ2_M')
parser.add_argument('--model', required=True, help='Served model name from the recipe definition.')
parser.add_argument('--cluster-id', help='Require a healthy registration from this cluster during gateway discovery checks.')
parser.add_argument('--mode', choices=['chat', 'verify', 'auth', 'discovery'], default='chat')
parser.add_argument('--stream', action='store_true')
Expand Down
Loading