Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 2 additions & 1 deletion deploy/helm/llm-api-gateway/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -69,7 +69,8 @@ Important settings to review before deployment:
- `llmApiGateway.config.nvcfGrpc*` for optional NVCF gRPC auth integration
- `llmApiGateway.metrics.enabled` to expose a metrics port on the Service and Deployment (default: `false`)
- `llmApiGateway.metrics.serviceMonitor.enabled` to create a Prometheus `ServiceMonitor` (requires `metrics.enabled`)
- `llmApiGateway.olric.*` for embedded rate-limit state and peer discovery
- `llmApiGateway.config.lastCluster.*` for session last-cluster hints sent to the LLM Request Router (default: disabled). Sets `STARGATE_LAST_CLUSTER_ENABLED`, `STARGATE_LAST_CLUSTER_TTL` (`10m`), `STARGATE_LAST_CLUSTER_LOOKUP_TIMEOUT` (`20ms`), and `STARGATE_LAST_CLUSTER_LOCAL_MAX_ENTRIES` (`100000`)
- `llmApiGateway.olric.*` for embedded rate-limit and last-cluster state and peer discovery. The Olric node starts whenever `olric.enabled` is `true`, even with rate limiting off
- `llmApiGateway.vault.*` for JWT authentication path, role, and audience values used by the Vault Agent injector

The default values include development-oriented placeholders. Override them before using the chart in any shared or production environment.
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -27,6 +27,10 @@ data:
STARGATE_URL: {{ required "llmApiGateway.config.requestRouterUrl is required" $cfg.requestRouterUrl | quote }}
STARGATE_CONNECT_TIMEOUT: {{ $cfg.requestRouterConnectTimeout | quote }}
STARGATE_REQUEST_TIMEOUT: {{ $cfg.requestRouterRequestTimeout | quote }}
STARGATE_LAST_CLUSTER_ENABLED: {{ $cfg.lastCluster.enabled | toString | quote }}
STARGATE_LAST_CLUSTER_TTL: {{ $cfg.lastCluster.ttl | quote }}
STARGATE_LAST_CLUSTER_LOOKUP_TIMEOUT: {{ $cfg.lastCluster.lookupTimeout | quote }}
STARGATE_LAST_CLUSTER_LOCAL_MAX_ENTRIES: {{ $cfg.lastCluster.localMaxEntries | int64 | toString | quote }}
NVCF_GATEWAY_INFERENCE_WRITE_TIMEOUT: {{ $cfg.inferenceWriteTimeout | quote }}
NVCF_GRPC_ADDR: {{ $cfg.nvcfGrpcAddr | quote }}
NVCF_GRPC_INSECURE: {{ $cfg.nvcfGrpcInsecure | toString | quote }}
Expand Down
22 changes: 17 additions & 5 deletions deploy/helm/llm-api-gateway/llm-api-gateway/values.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -79,11 +79,23 @@ llmApiGateway:
defaultModel: ""
defaultRpm: 120
defaultTpm: 240000

# Rate-limit state is kept in an embedded Olric cluster colocated with the
# gateway pods (no separate Redis deployment). The gateway's server.go picks
# ratelimit.AllowAll when RATE_LIMIT_ENABLED=true AND OLRIC_ENABLED=false, so
# enabling rate limiting without Olric silently turns into a no-op.
# Session last-cluster hint sent to the LLM Request Router as
# x-stargate-last-cluster-id. Entries live in the embedded Olric cluster
# when olric.enabled is true, otherwise in a per-pod LRU capped at
# localMaxEntries. Enable it after the router is upgraded and before any
# model turns on last_cluster_affinity.
lastCluster:
enabled: false
ttl: 10m
lookupTimeout: 20ms
localMaxEntries: 100000

# Rate-limit state and session last-cluster hints are kept in an embedded
# Olric cluster colocated with the gateway pods (no separate Redis
# deployment). The node starts whenever olric.enabled is true. The gateway's
# server.go picks ratelimit.AllowAll when RATE_LIMIT_ENABLED=true AND
# OLRIC_ENABLED=false, so enabling rate limiting without Olric silently turns
# into a no-op.
olric:
enabled: true
# "local" | "lan" | "wan" - memberlist tuning. Use "local" for single-pod k3d
Expand Down
14 changes: 14 additions & 0 deletions deploy/helm/llm-api-gateway/scripts/test-default-ownership.sh
Original file line number Diff line number Diff line change
Expand Up @@ -32,17 +32,31 @@ test "$(yq ea -r 'select(.kind == "Deployment") | [.spec.template.spec.container
fail "generic chart must leave the metrics endpoint disabled by default"
test "$(yq ea -r '[select(.kind == "ServiceMonitor")] | length' "$default_manifest")" = 0 ||
fail "generic chart must not create a ServiceMonitor by default"
test "$(read_config "$default_manifest" STARGATE_LAST_CLUSTER_ENABLED)" = false ||
fail "generic chart must leave last-cluster hints disabled by default"
test "$(read_config "$default_manifest" STARGATE_LAST_CLUSTER_TTL)" = 10m ||
fail "last-cluster TTL default must be 10m"
test "$(read_config "$default_manifest" STARGATE_LAST_CLUSTER_LOOKUP_TIMEOUT)" = 20ms ||
fail "last-cluster lookup timeout default must be 20ms"
test "$(read_config "$default_manifest" STARGATE_LAST_CLUSTER_LOCAL_MAX_ENTRIES)" = 100000 ||
fail "last-cluster local store cap default must be 100000"

helm template llm-api-gateway "$chart_dir" \
--namespace nvcf \
--set-string llmApiGateway.image.repository=example.invalid/llm-api-gateway \
--set llmApiGateway.config.nvcfGrpcInsecure=true \
--set llmApiGateway.metrics.enabled=true \
--set llmApiGateway.metrics.serviceMonitor.enabled=true \
--set llmApiGateway.config.lastCluster.enabled=true \
--set llmApiGateway.config.lastCluster.localMaxEntries=1000000 \
>"$override_manifest"

test "$(read_config "$override_manifest" NVCF_GRPC_INSECURE)" = true ||
fail "explicit plaintext transport override did not reach the ConfigMap"
test "$(read_config "$override_manifest" STARGATE_LAST_CLUSTER_ENABLED)" = true ||
fail "last-cluster opt-in did not reach the ConfigMap"
test "$(read_config "$override_manifest" STARGATE_LAST_CLUSTER_LOCAL_MAX_ENTRIES)" = 1000000 ||
fail "large last-cluster cap must render as a plain integer"
test "$(yq ea -r 'select(.kind == "Deployment") | [.spec.template.spec.containers[0].ports[] | select(.name == "metrics")] | length' "$override_manifest")" = 1 ||
fail "explicit metrics override did not expose the metrics container port"
test "$(yq ea -r '[select(.kind == "ServiceMonitor" and .metadata.name == "llm-api-gateway-metrics")] | length' "$override_manifest")" = 1 ||
Expand Down
3 changes: 3 additions & 0 deletions docs/observability/metrics/llm-api-gateway/metrics.md
Original file line number Diff line number Diff line change
Expand Up @@ -40,3 +40,6 @@ other unbounded request fields as metric labels.
| `llm_api_gateway_rate_limit_synchronizer_queue_wait_seconds` | Histogram | `llm-api-gateway:9464/metrics` | None | Time spent queueing a rate limit event in seconds. |
| `llm_api_gateway_rate_limit_synchronizer_queue_length` | Gauge | `llm-api-gateway:9464/metrics` | None | Current rate limit synchronizer queue length. |
| `llm_api_gateway_rate_limit_synchronizer_events_dropped_total` | Counter | `llm-api-gateway:9464/metrics` | `reason` | Number of rate limit events dropped before publishing. `reason` is a bounded enum such as `old_message`. |
| `llm_api_gateway_last_cluster_lookups_total` | Counter | `llm-api-gateway:9464/metrics` | `result` | Session last-cluster store lookups. `result` is `hit`, `miss`, `error`, `timeout`, or `skipped`. All values start at zero. |
| `llm_api_gateway_last_cluster_writes_total` | Counter | `llm-api-gateway:9464/metrics` | `result` | Session last-cluster store writes. `result` is `ok` or `error`. All values start at zero. |
| `llm_api_gateway_last_cluster_lookup_duration_seconds` | Histogram | `llm-api-gateway:9464/metrics` | `result` | Session last-cluster store lookup duration in seconds, with buckets from 1 ms to 50 ms. Skipped lookups are not timed. |
7 changes: 4 additions & 3 deletions docs/observability/metrics/llm-request-router/metrics.md
Original file line number Diff line number Diff line change
Expand Up @@ -14,9 +14,9 @@ The self-managed stack maps
## Label Boundaries

Use bounded labels only. Keep `routing_key`, `model`, `inference_server_id`,
`algorithm`, `selection`, `status`, `result`, `reason`, and `cache` to bounded
service dimensions. Routing expression content never becomes a
label value. Do not add request IDs, session IDs, function IDs, organization
`algorithm`, `selection`, `session_state`, `status`, `result`, `reason`, and
`cache` to bounded service dimensions. Routing expression content never becomes
a label value. Do not add request IDs, session IDs, function IDs, organization
IDs, project IDs, raw URLs, raw prompts, authorization values, or other
unbounded request fields as metric labels.

Expand All @@ -29,6 +29,7 @@ unbounded request fields as metric labels.
| `stargate_proxy_retries_total` | Counter | `llm-request-router:9090/metrics` | `routing_key`, `model`, `reason` | Total proxy retries by retry reason. |
| `stargate_routing_selections_total` | Counter | `llm-request-router:9090/metrics` | `routing_key`, `model`, `algorithm`, `selection` | Primary and ranked fallback cluster choices used for upstream attempts. |
| `stargate_routing_kv_free_token_fallback_selections_total` | Counter | `llm-request-router:9090/metrics` | `routing_key`, `model`, `algorithm` | Routes selected after a higher-ranked candidate failed the KV free-token check. |
| `stargate_routing_session_selections_total` | Counter | `llm-request-router:9090/metrics` | `routing_key`, `model`, `algorithm`, `session_state`, `selection` | Cluster choices for requests classified by `last_cluster_affinity`. `session_state` is `new`, `returning`, or `stale`. For `returning`, `primary` means the request went to its last cluster; for `new` and `stale`, it means the cluster came from the affinity group. Every state and selection pair starts at zero once a routing key, model, and algorithm records its first classified request. |
| `stargate_routing_expressions_total` | Counter | `llm-request-router:9090/metrics` | `algorithm`, `cache` | Accepted `x-routing-method` values with parameters. `cache` is `hit`, `build`, or `rebuild`. Values without parameters and rejected values are not counted. Rejected `x-routing-method` values write the warn log `invalid routing algorithm header` with their `rejection_reason`. |
| `stargate_routing_expression_cache_entries` | Gauge | `llm-request-router:9090/metrics` | none | Approximate number of routing targets that hold a routing expression configuration, with a capacity of 16384 per process. The router runs cache maintenance once a minute to remove entries idle for 15 minutes. Scrapes read the count without maintenance, so it can lag behind pending insertions or removals until the next maintenance run. |
| `stargate_proxy_retry_exhausted_total` | Counter | `llm-request-router:9090/metrics` | `routing_key`, `model`, `reason` | Total requests that exhausted retry options. |
Expand Down
45 changes: 44 additions & 1 deletion docs/overview/llm-gateway.md
Original file line number Diff line number Diff line change
Expand Up @@ -419,6 +419,46 @@ sequenceDiagram
Sticky routing only affects backend selection when the LLM request router is
configured with a cache-affinity-aware routing method for the target model.

### Last-Cluster Hints

The gateway can remember which LLM request router cluster served each
session's last successful request and send it back on the next request in the
trusted `x-stargate-last-cluster-id` header. The router uses the hint when the
target model's `wait-and-widen` or `pulsar-wait-and-widen` configuration sets
`last_cluster_affinity: true`. A returning session then waits for the cluster
that holds its KV cache, and a new session goes straight to the best cluster
that is free now.

A session is the routing key, model, and cache affinity key. Only sessions
from `prompt_cache_key`, `conversation.id`, or `x-multi-turn-session-id` are
eligible. Sessions that fall back to a messages or input hash change key every
turn, so the gateway never looks them up or stores them.

The gateway records the `x-stargate-cluster-id` of each 2xx router response
when the response headers arrive. Errors never change the stored value. A
client-supplied `x-stargate-last-cluster-id` is always stripped.

| Env var | Helm value | Default | Effect |
| --- | --- | --- | --- |
| `STARGATE_LAST_CLUSTER_ENABLED` | `llmApiGateway.config.lastCluster.enabled` | `false` | Turns lookup, the header, and store writes on. |
| `STARGATE_LAST_CLUSTER_TTL` | `llmApiGateway.config.lastCluster.ttl` | `10m` | Entry lifetime, reset on every write. |
| `STARGATE_LAST_CLUSTER_LOOKUP_TIMEOUT` | `llmApiGateway.config.lastCluster.lookupTimeout` | `20ms` | Lookup deadline on the request path. The request proceeds without the hint after it. |
| `STARGATE_LAST_CLUSTER_LOCAL_MAX_ENTRIES` | `llmApiGateway.config.lastCluster.localMaxEntries` | `100000` | Per-replica LRU cap when Olric is off. |

With `OLRIC_ENABLED=true`, entries live in the `stargate-last-cluster` DMap of
the gateway's embedded Olric cluster and are shared by all replicas. The Olric
node starts whenever Olric is enabled, even with rate limiting off.

Roll out in this order:

1. Upgrade the LLM request router with `last_cluster_affinity` off. It accepts
the header and ignores it.
2. Set `STARGATE_LAST_CLUSTER_ENABLED=true` on the gateway.
3. Turn on `last_cluster_affinity` per model in the router configuration.

Do not turn on `last_cluster_affinity` before step 2. Without the hint, every
request is treated as a new session and skips the affinity wait.

## Metrics

LLM API Gateway request metrics include a `function_id` label. The value is the
Expand All @@ -437,9 +477,12 @@ key, such as health checks, use `function_id="none"`.
| `llm_api_gateway_provider_time_seconds` | Histogram | `endpoint`, `phase`, `stream`, `function_id` | Provider-reported timing phases. |
| `llm_api_gateway_stream_first_token_seconds` | Histogram | `endpoint`, `function_id` | Time from stream request start to the first token. |
| `llm_api_gateway_stream_duration_seconds` | Histogram | `endpoint`, `status`, `function_id` | Total stream duration. |
| `llm_api_gateway_last_cluster_lookups_total` | Counter | `result` | Last-cluster store lookups: `hit`, `miss`, `error`, `timeout`, or `skipped`. |
| `llm_api_gateway_last_cluster_writes_total` | Counter | `result` | Last-cluster store writes: `ok` or `error`. |
| `llm_api_gateway_last_cluster_lookup_duration_seconds` | Histogram | `result` | Last-cluster store lookup latency, 1 ms to 50 ms buckets. |

Infrastructure metrics for authentication, rate limit synchronization, pub/sub,
and the distributed cache do not include `function_id` because they are not
the last-cluster store, and the distributed cache do not include `function_id` because they are not
associated with a single routed function.

Use the label to calculate request rate by function:
Expand Down
54 changes: 53 additions & 1 deletion src/invocation-plane-services/llm-api-gateway/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -72,6 +72,37 @@ conversation. The gateway preserves the raw body field for the model backend.
It forwards only a SHA-256-derived value in the internal
`x-cache-affinity-key` header to Stargate.

### Last-cluster hints

With `STARGATE_LAST_CLUSTER_ENABLED=true`, the gateway remembers which
Stargate cluster served each session's last successful request and sends it
back on the session's next request as `x-stargate-last-cluster-id`. Stargate
uses the hint only for models whose routing algorithm sets
`last_cluster_affinity`.

- A session is the tuple of routing key, model, and cache affinity key.
- Only sessions from `prompt_cache_key`, `conversation_id`, or the
`x-multi-turn-session-id` header are eligible. Payload-derived sessions hash
the full message list, which changes every turn, so they are never looked up
or stored. Sessions from `x-claude-code-session-id` are also skipped.
- The gateway writes the `x-stargate-cluster-id` of a 2xx Stargate response
when the response headers arrive, before a stream ends. Non-2xx responses and
transport errors never change the stored value. Writes run in the background
and never delay or fail the client response.
- The lookup is bounded by `STARGATE_LAST_CLUSTER_LOOKUP_TIMEOUT`. On a miss,
error, or timeout the request proceeds without the hint.
- Store keys are `lc:v1:` plus the hex SHA-256 of the length-prefixed routing
key, model, and cache affinity key. Raw routing keys never appear in store
keys. Affinity keys and session IDs never appear in store keys or logs.
Cluster IDs over 256 bytes are not stored.
- With `OLRIC_ENABLED=true`, entries live in the `stargate-last-cluster` DMap
of the embedded Olric cluster, so every replica sees them. With Olric off,
each replica keeps its own LRU capped at
`STARGATE_LAST_CLUSTER_LOCAL_MAX_ENTRIES`. Every write resets the entry TTL.
- The gateway always strips a client-supplied `x-stargate-last-cluster-id`,
whether or not the feature is enabled, and keeps stripping `x-stargate-*`
response headers from client responses.

When `NVCF_GRPC_ADDR` is configured, the gateway authenticates each request
through the NVCF LLM gRPC auth service, derives the per-caller rate-limit key
from `authContext["ncaId"]`, optionally scopes it further by project, and keeps
Expand Down Expand Up @@ -146,6 +177,14 @@ Useful overrides:
413 (default `0`, no limit)
- `STARGATE_CONNECT_TIMEOUT` to control Stargate dial timeout
- `STARGATE_REQUEST_TIMEOUT` to cap end-to-end Stargate request time
- `STARGATE_LAST_CLUSTER_ENABLED=true` to send session last-cluster hints to
Stargate (default `false`)
- `STARGATE_LAST_CLUSTER_TTL` for the last-cluster entry lifetime, reset on
every write (default `10m`)
- `STARGATE_LAST_CLUSTER_LOOKUP_TIMEOUT` for the lookup deadline on the
request path (default `20ms`)
- `STARGATE_LAST_CLUSTER_LOCAL_MAX_ENTRIES` for the per-replica store size
when Olric is off (default `100000`)
- `NVCF_GATEWAY_INFERENCE_WRITE_TIMEOUT` to cap how long one response write
may stall on a client that stopped reading (default `60s`, `0s` disables).
It applies only while a write is in progress, so long streams, long
Expand All @@ -160,7 +199,9 @@ Useful overrides:
- `NVCF_GRPC_TIMEOUT` to cap each gRPC auth or policy call
- `RATE_LIMIT_ENABLED=false` to disable rate limiting locally
- `RATE_LIMIT_FAIL_OPEN=false` to make Olric or limiter failures fatal
- `OLRIC_ENABLED=false` to skip starting the embedded Olric node
- `OLRIC_ENABLED=false` to skip starting the embedded Olric node. When it is
`true`, the node starts even if rate limiting is off, and it also holds
last-cluster hints
- `OLRIC_BIND_PORT`, `OLRIC_MEMBERLIST_BIND_PORT`, and `OLRIC_PEERS` for
multi-instance Olric clustering
- `OTEL_SERVICE_NAME` to override the emitted service name
Expand All @@ -178,6 +219,17 @@ time, first-token time, and stream duration metrics. Infrastructure metrics for
authentication, pub/sub, rate-limit synchronization, and Olric remain
function-independent.

Last-cluster store metrics use only a `result` label:

- `llm_api_gateway_last_cluster_lookups_total`: `hit`, `miss`, `error`,
`timeout`, or `skipped` (enabled, but the session is not eligible)
- `llm_api_gateway_last_cluster_writes_total`: `ok` or `error`
- `llm_api_gateway_last_cluster_lookup_duration_seconds`: lookup latency, with
buckets from 1 ms to 50 ms

Lookups and writes also create the `llm-api-gateway.last_cluster_lookup` and
`llm-api-gateway.last_cluster_write` spans.

Example request-rate query:

```promql
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -89,6 +89,7 @@ go_test(
"body_limit_test.go",
"embeddings_handler_test.go",
"info_test.go",
"last_cluster_test.go",
"middleware_telemetry_test.go",
"messages_handler_test.go",
"middleware_test.go",
Expand All @@ -113,6 +114,7 @@ go_test(
"//src/invocation-plane-services/llm-api-gateway/api/adapters/openairesponses",
"//src/invocation-plane-services/llm-api-gateway/config",
"//src/invocation-plane-services/llm-api-gateway/internal/ptr",
"//src/invocation-plane-services/llm-api-gateway/lastcluster",
"//src/invocation-plane-services/llm-api-gateway/models",
"//src/invocation-plane-services/llm-api-gateway/nvcf",
"//src/invocation-plane-services/llm-api-gateway/provider",
Expand Down
Loading
Loading