You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
feat(stargate,llm-api-gateway): route returning sessions to their last cluster in wait-and-widen and pulsar-wait-and-widen #2289
Add a session-aware mode to wait-and-widen and pulsar-wait-and-widen, and
the matching LLM API gateway support. The gateway remembers which cluster
served each session's last request and sends it back to Stargate on the next
request. Stargate uses that hint to tell first requests apart from follow-up
requests:
A new session (no hint) skips the affinity wait and the affinity prefill
discount. It goes to the best-ranked cluster that is selectable now and does
not hold for its rank-1 cluster.
A returning session (hint present) moves the hinted cluster to the front of
its affinity order. It then runs the configured affinity wait, discount, and
widening unchanged.
It also removes the pulsar-wait-and-widen shortcut that sends a request to
its rank-1 cluster without any load check when no queue-SLO fields are set.
This keeps the stable affinity order and lets a new session overflow quickly
and deterministically. It also stops a session from moving back to rank 1 once
that cluster frees up. That move throws away the KV cache the session built on
the overflow cluster.
Dependencies
The Stargate work builds on two open Pull Requests. Start it after both merge.
The gateway work does not depend on them.
With an affinity wait, a request waits for its affinity group before
overflowing. That suits a returning session, whose KV cache is on the
affinity group. It hurts a new session, which has no cache anywhere and gains
nothing from waiting.
Overflow also causes cache thrash. When the rank-1 cluster is busy, a session
overflows to a lower-ranked cluster and builds KV cache there. When rank 1
recovers, the next request routes back to rank 1 and misses the cache. Under
steady load this repeats.
Stargate cannot tell new from returning sessions alone. It keeps no
per-session state and does not share routing state between pods. The gateway
already sees x-stargate-cluster-id on every Stargate response and runs an
embedded Olric cluster shared by its replicas, so it can carry the last
cluster forward.
Terms
Session: one (x-routing-key, x-model, x-cache-affinity-key) tuple. In
the gateway, these are RoutingKey, Model, and CacheAffinityKey on the
request context.
Last cluster: the x-stargate-cluster-id value from the session's most
recent 2xx response.
Affinity order: the ordered cluster list the algorithm derives from the
affinity key.
wait-and-widen: the hash ring's selected clusters, in ring-walk order.
pulsar-wait-and-widen: the full Pulsar ranking.
Affinity group: the first cache_affinity_backend_selection_count (k)
clusters of the affinity order.
The x-stargate-cluster-id from the session's last 2xx response.
Use the cluster ID, not a rank index. A cluster ID is valid for both the hash
ring and the Pulsar ranking and survives candidate-set changes.
Stargate consumes the header and does not forward it to Pylon. Response
headers do not change.
Remove the eligible-primary shortcut
Today pulsar-wait-and-widen returns the rank-1 cluster immediately when the
config has no queue-SLO fields (max_queue_time_floor_ms, max_queue_time_ceil_ms) and the cluster passes retry exclusion and the KV
free-token check. That check ignores load, so a full rank-1 cluster keeps
taking requests with no limit. A hot key then slows every other key whose
rank-1 cluster it shares.
Remove the shortcut for every request, regardless of last_cluster_affinity.
The rank-1 cluster then goes through the affinity phase like any other
request: queue admission with max_queued, the affinity wait, and the
affinity discount. Deployments that want unlimited stickiness use pulsar.
Deployments that want bounded stickiness raise the affinity group's max_queued and set cache_affinity_wait_ms.
Behavior change: a no-SLO config left at the defaults (max_queued: 0, cache_affinity_wait_ms: 0) overflows as soon as its rank-1 cluster has no
free engine slot.
Stargate config
Add a boolean algorithm field, last_cluster_affinity, default false, to wait-and-widen and pulsar-wait-and-widen. Reject it at startup for every
other algorithm.
Session classification
When last_cluster_affinity is true, classify each request once, before
the first routing attempt. Retries keep the first classification.
Condition
Session state
No affinity group: x-cache-affinity-key absent, or wait-and-widen without k
Not applicable. Existing behavior. Header ignored.
Header absent, blank, or not valid UTF-8
new
Header names a cluster not in the current candidate set
stale (routed as new)
Header names a cluster in the current candidate set
returning
Selection rules
new and stale, both algorithms:
Treat X as 0 and s as 1.0 for this request.
Everything else runs as configured: affinity order, k, TTFT buckets, queue
admission, band_widen_interval_ms, and fallback_max_queued.
returning, both algorithms:
Move the last cluster to position 1 of the affinity order. Keep the relative
order of every other cluster.
The affinity group is the first k clusters of the reordered order. Its size
stays k. If the last cluster was outside the original group, the original
group's last member drops out.
pulsar-wait-and-widen widening bands come from the reordered ranking. wait-and-widen global selection is unchanged because it already covers
every cluster.
X, s, and every other setting apply as configured.
Retry exclusion and KV free-token checks apply to the last cluster the same
way they apply to any other candidate. Per-key ring-selection and ranking
caches keep storing the original order. The promotion is per request.
Gateway changes
Code lives in src/invocation-plane-services/llm-api-gateway. Chart wiring
lives in deploy/helm/llm-api-gateway.
Configuration, all off or unset by default:
Env var
Default
Effect
STARGATE_LAST_CLUSTER_ENABLED
false
Turns lookup, header, and store writes on.
STARGATE_LAST_CLUSTER_TTL
10m
Entry lifetime, reset on every write.
STARGATE_LAST_CLUSTER_LOOKUP_TIMEOUT
20ms
Lookup deadline on the request path.
STARGATE_LAST_CLUSTER_LOCAL_MAX_ENTRIES
100000
Size cap for the in-process store when Olric is off.
Eligible sessions:
Only sessions whose SessionSource is prompt_cache_key, conversation_id, or header.
Sessions with SessionSourcepayload are skipped for lookup and writes.
Their key is a hash of the full message list, which changes every turn, so
they never return.
Store:
With OLRIC_ENABLED=true, use a dedicated DMap named stargate-last-cluster on the gateway's embedded Olric node. Start the
Olric node whenever Olric is enabled, not only when the rate limiter is on.
Today newRateLimiter in server/server.go owns the node.
With Olric off, use an in-process LRU capped at STARGATE_LAST_CLUSTER_LOCAL_MAX_ENTRIES, with the same TTL.
Key: lc:v1: plus the hex SHA-256 of the length-prefixed routing key,
model, and cache affinity key. Raw routing keys never appear in store keys.
Value: the cluster ID. Values over 256 bytes are not stored.
Request path, for every Stargate call (Complete, Stream, and Proxy in provider/stargate.go):
Delete any inbound x-stargate-last-cluster-id before building the
outbound request. Proxy clones inbound headers, so the delete must run
after the clone.
For eligible sessions, look up the key with the lookup timeout. On a hit,
set x-stargate-last-cluster-id. On a miss, error, or timeout, omit it and
continue.
Response path:
On a 2xx Stargate response with a non-empty x-stargate-cluster-id, write key -> cluster ID with the TTL. Write when response headers arrive, not
at stream end. Prefill has already started on that cluster.
Writes run off the request path. They never delay or fail the client
response.
Do not write on non-2xx responses or transport errors.
Concurrent requests in one session may both miss. Last write wins.
Keep stripping x-stargate-* response headers from client responses.
Rollout
Order:
Deploy Stargate with this change and last_cluster_affinity off. It
consumes the header and ignores it.
Deploy the gateway with STARGATE_LAST_CLUSTER_ENABLED=true.
Turn on last_cluster_affinity per model.
Before upgrading Stargate, review pulsar-wait-and-widen configs without queue-SLO
fields. Set max_queued and cache_affinity_wait_ms, or switch to pulsar,
to keep requests on the rank-1 cluster under load.
Do not turn on last_cluster_affinity before step 2. Otherwise every request
is classified new and loses its affinity wait.
Acceptance criteria
Shortcut removal
The eligible-primary shortcut is gone from pulsar-wait-and-widen.
No queue-SLO fields, max_queued: 0, rank-1 cluster at its engine
concurrency limit, rank-2 cluster with a free slot: the request is not
sent to rank 1. Replace without_queue_slo_keeps_full_primary with this
test.
Same setup with max_queued: 4 and 1 request queued on rank 1: the
request goes to rank 1.
Same setup with cache_affinity_wait_ms: 300 and max_queued: 0: the
decision is a timed wait until 300 ms, then the request widens.
Compatibility
With last_cluster_affinity unset or false, routing decisions for wait-and-widen and pulsar-wait-and-widen do not change, with or
without the header, except for the shortcut removal above. Existing
load-balancer tests pass with no changes to their assertions, except
the tests that encode the shortcut.
Setting last_cluster_affinity on power-of-n, round-robin, random, or pulsar fails startup with an error that names the field
and the algorithm.
x-stargate-last-cluster-id is never forwarded to Pylon, for any
algorithm or flag value.
New sessions
Run each criterion for both algorithms, with X = 300 ms and s = 0.1.
Rank-1 cluster A is full and rank-2 cluster B has a free slot: the
request goes to B at elapsed 0 with no routing wait.
Every cluster has a free slot: the request goes to the cluster that the
flag-off algorithm picks at X = 0 and s = 1.0.
The TTFT used for bucket placement and comparison uses s = 1.0, not 0.1.
No candidate is selectable: the decision equals the flag-off decision at
X = 0 and s = 1.0 for the same snapshot and elapsed time.
stale decisions equal new decisions for the same snapshot.
Returning sessions
k = 1, original order A, B, C, last cluster B: during X, the request
waits for B and is never sent to A, even when A has a free slot and B is
full.
k = 1, last cluster B: s applies to B, not to A.
k = 2, original group (A, B), last cluster C: the affinity group is
(C, A). B is selectable only after X.
k = 2, original group (A, B), last cluster B: the affinity group members
stay A and B.
pulsar-wait-and-widen, k = 1, band_widen_interval_ms > 0, ranking
A, B, C, D, E, last cluster D: after X the open set is D, A, B. C opens
with the next band.
Last cluster retry-excluded or failing the KV free-token check: selection
continues as if it were an ineligible rank-1 cluster.
After a promoted request, a request for the same key with no header uses
the original order. The promotion is not cached.
Two Stargate instances with the same seed, config, and candidate
snapshot make the same choice for the same request and header.
Cache-thrash regression test
Integration test in crates/stargate/tests/suite/load_balancing.rs, run
for both algorithms with k = 1:
Session S has rank 1 on A and rank 2 on B. A is full.
The first request with no header goes to B.
A drains and has free slots.
The next 10 requests send x-stargate-last-cluster-id: B.
All 10 go to B.
With the flag off, the same test sends the step 4 requests to A, which
shows the test covers the behavior change.
Header parsing
Blank or non-UTF-8 header values are treated as absent and do not return 400.
The header has no effect when x-cache-affinity-key is absent, or when wait-and-widen has no cache_affinity_backend_selection_count.
The header has no effect when x-routing-method selects an algorithm
without last_cluster_affinity enabled.
Observability
New counter stargate_routing_session_selections_total with labels routing_key, model, algorithm, session_state (new, returning, stale), and selection (primary, fallback). For returning, primary means the request went to the last cluster.
Counters start at zero for every session_state and selection pair.
The proxy request span records routing.session_state.
A debug log line records the session state and whether the last
cluster was promoted, with request ID, model, and routing key.
Gateway
Tests live under src/invocation-plane-services/llm-api-gateway.
With STARGATE_LAST_CLUSTER_ENABLED=false, outbound Stargate requests
and store contents do not change, except that inbound x-stargate-last-cluster-id is always stripped.
A client-supplied x-stargate-last-cluster-id never reaches Stargate on
the Complete, Stream, or Proxy paths, enabled or not.
Two requests with the same prompt_cache_key: the first carries no
header. The first response returns x-stargate-cluster-id: B. The second
request carries x-stargate-last-cluster-id: B.
Same flow with conversation_id and with x-multi-turn-session-id.
Payload-derived sessions never trigger a lookup or a write.
Same affinity key with a different model or routing key: no header.
Non-2xx responses and transport errors do not change the stored value.
A 2xx streaming response writes the value before the stream ends.
An entry older than the TTL is a miss. A write resets the TTL.
Lookup slower than the timeout: the request proceeds without the header
within the timeout plus 5 ms.
Store errors on lookup or write do not change the client status code or
body.
Two gateway replicas in one Olric cluster: a write through replica 1 is
visible to a lookup through replica 2.
Olric off: the in-process store holds at most STARGATE_LAST_CLUSTER_LOCAL_MAX_ENTRIES entries and evicts least
recently used first.
Olric on and rate limiter off: the Olric node starts and the store
works.
x-stargate-cluster-id is still stripped from client responses.
Counter llm_api_gateway_last_cluster_lookups_total with label result
(hit, miss, error, timeout, skipped) and counter llm_api_gateway_last_cluster_writes_total with label result (ok, error). Both start at zero for every label value.
Histogram llm_api_gateway_last_cluster_lookup_duration_seconds with
buckets from 1 ms to 50 ms.
Lookup and write each create a child span (llm-api-gateway.last_cluster_lookup, llm-api-gateway.last_cluster_write) with a result attribute and
error status on failure.
Store errors log at warn with request ID, model, and routing key. Do
not log the affinity key or session ID.
deploy/helm/llm-api-gateway exposes the four env vars through values,
with the defaults above.
End-to-end
A test that runs the gateway against a real Stargate with last_cluster_affinity on reproduces the cache-thrash scenario above:
after the first request lands on B and A drains, the next 10 requests in
the session go to B.
Documentation
docs/api-gateway-contract.md: add the header to the optional trusted
headers table, add the gateway rules above, and update the checklist.
docs/load-balancer-configuration.md: document last_cluster_affinity,
session classification, the selection rules for both algorithms, and the
rollout order. Remove the statement that an eligible primary wins
immediately without queue-SLO fields, and add the migration guidance
from the rollout section.
The request-router metrics reference lists the new counter.
src/invocation-plane-services/llm-api-gateway/README.md documents the
env vars, eligible session sources, and store behavior.
docs/overview/llm-gateway.md lists the env vars and the rollout
order.
Validation
Run a multi-turn workload with a saturated rank-1 cluster through stargate-bench against mock-dynamo, or the routing simulator from test(stargate): add discrete-event routing simulator #2223. Cover both algorithms, flag off and on. Post KV cache reuse, TTFT
p50 and p99, goodput, and the share of returning requests that left
their last cluster in the Pull Request.
Add a pulsar-wait-and-widen case to the load-balancer microbenchmark
in stargate-bench. Post per-decision latency with and without the
shortcut.
Out of scope
pulsar, power-of-n, round-robin, and random.
Sharing the last-cluster store across gateway deployments or regions.
changed the title [-]feat(stargate): route returning sessions to their last cluster in wait-and-widen and pulsar-wait-and-widen[/-][+]feat(stargate,llm-api-gateway): route returning sessions to their last cluster in wait-and-widen and pulsar-wait-and-widen[/+]on Oct 5, 2026
Summary
Add a session-aware mode to
wait-and-widenandpulsar-wait-and-widen, andthe matching LLM API gateway support. The gateway remembers which cluster
served each session's last request and sends it back to Stargate on the next
request. Stargate uses that hint to tell first requests apart from follow-up
requests:
discount. It goes to the best-ranked cluster that is selectable now and does
not hold for its rank-1 cluster.
its affinity order. It then runs the configured affinity wait, discount, and
widening unchanged.
It also removes the
pulsar-wait-and-widenshortcut that sends a request toits rank-1 cluster without any load check when no queue-SLO fields are set.
This keeps the stable affinity order and lets a new session overflow quickly
and deterministically. It also stops a session from moving back to rank 1 once
that cluster frees up. That move throws away the KV cache the session built on
the overflow cluster.
Dependencies
The Stargate work builds on two open Pull Requests. Start it after both merge.
The gateway work does not depend on them.
pulsar-wait-and-widenthe same affinity group,affinity wait (
cache_affinity_wait_ms), affinity prefill discount(
cache_affinity_input_tokens_scale), and timed widening thatwait-and-widenhas. Both algorithms then share one model: an affinityphase, followed by a widening phase. This issue changes how a request
enters that model and removes the eligible-primary shortcut that feat(stargate): make pulsar-wait-and-widen wait for affinity and widen gradually #2288
keeps.
TPS instead of live mean TPS. The ranking no longer shifts with load, so a
new session's rank order is stable, and a promoted last cluster is not
undone by ranking churn.
Problem
With an affinity wait, a request waits for its affinity group before
overflowing. That suits a returning session, whose KV cache is on the
affinity group. It hurts a new session, which has no cache anywhere and gains
nothing from waiting.
Overflow also causes cache thrash. When the rank-1 cluster is busy, a session
overflows to a lower-ranked cluster and builds KV cache there. When rank 1
recovers, the next request routes back to rank 1 and misses the cache. Under
steady load this repeats.
Stargate cannot tell new from returning sessions alone. It keeps no
per-session state and does not share routing state between pods. The gateway
already sees
x-stargate-cluster-idon every Stargate response and runs anembedded Olric cluster shared by its replicas, so it can carry the last
cluster forward.
Terms
(x-routing-key, x-model, x-cache-affinity-key)tuple. Inthe gateway, these are
RoutingKey,Model, andCacheAffinityKeyon therequest context.
x-stargate-cluster-idvalue from the session's mostrecent 2xx response.
affinity key.
wait-and-widen: the hash ring's selected clusters, in ring-walk order.pulsar-wait-and-widen: the full Pulsar ranking.cache_affinity_backend_selection_count(k)clusters of the affinity order.
cache_affinity_wait_ms.cache_affinity_input_tokens_scale.Proposed design
Request contract
Add one trusted request header:
x-stargate-last-cluster-idx-stargate-cluster-idfrom the session's last 2xx response.Use the cluster ID, not a rank index. A cluster ID is valid for both the hash
ring and the Pulsar ranking and survives candidate-set changes.
Stargate consumes the header and does not forward it to Pylon. Response
headers do not change.
Remove the eligible-primary shortcut
Today
pulsar-wait-and-widenreturns the rank-1 cluster immediately when theconfig has no queue-SLO fields (
max_queue_time_floor_ms,max_queue_time_ceil_ms) and the cluster passes retry exclusion and the KVfree-token check. That check ignores load, so a full rank-1 cluster keeps
taking requests with no limit. A hot key then slows every other key whose
rank-1 cluster it shares.
Remove the shortcut for every request, regardless of
last_cluster_affinity.The rank-1 cluster then goes through the affinity phase like any other
request: queue admission with
max_queued, the affinity wait, and theaffinity discount. Deployments that want unlimited stickiness use
pulsar.Deployments that want bounded stickiness raise the affinity group's
max_queuedand setcache_affinity_wait_ms.Behavior change: a no-SLO config left at the defaults (
max_queued: 0,cache_affinity_wait_ms: 0) overflows as soon as its rank-1 cluster has nofree engine slot.
Stargate config
Add a boolean algorithm field,
last_cluster_affinity, defaultfalse, towait-and-widenandpulsar-wait-and-widen. Reject it at startup for everyother algorithm.
Session classification
When
last_cluster_affinityistrue, classify each request once, beforethe first routing attempt. Retries keep the first classification.
x-cache-affinity-keyabsent, orwait-and-widenwithout knewstale(routed asnew)returningSelection rules
newandstale, both algorithms:0and s as1.0for this request.admission,
band_widen_interval_ms, andfallback_max_queued.returning, both algorithms:order of every other cluster.
stays k. If the last cluster was outside the original group, the original
group's last member drops out.
pulsar-wait-and-widenwidening bands come from the reordered ranking.wait-and-widenglobal selection is unchanged because it already coversevery cluster.
Retry exclusion and KV free-token checks apply to the last cluster the same
way they apply to any other candidate. Per-key ring-selection and ranking
caches keep storing the original order. The promotion is per request.
Gateway changes
Code lives in
src/invocation-plane-services/llm-api-gateway. Chart wiringlives in
deploy/helm/llm-api-gateway.Configuration, all off or unset by default:
STARGATE_LAST_CLUSTER_ENABLEDfalseSTARGATE_LAST_CLUSTER_TTL10mSTARGATE_LAST_CLUSTER_LOOKUP_TIMEOUT20msSTARGATE_LAST_CLUSTER_LOCAL_MAX_ENTRIES100000Eligible sessions:
SessionSourceisprompt_cache_key,conversation_id, orheader.SessionSourcepayloadare skipped for lookup and writes.Their key is a hash of the full message list, which changes every turn, so
they never return.
Store:
OLRIC_ENABLED=true, use a dedicated DMap namedstargate-last-clusteron the gateway's embedded Olric node. Start theOlric node whenever Olric is enabled, not only when the rate limiter is on.
Today
newRateLimiterinserver/server.goowns the node.STARGATE_LAST_CLUSTER_LOCAL_MAX_ENTRIES, with the same TTL.lc:v1:plus the hex SHA-256 of the length-prefixed routing key,model, and cache affinity key. Raw routing keys never appear in store keys.
Request path, for every Stargate call (
Complete,Stream, andProxyinprovider/stargate.go):x-stargate-last-cluster-idbefore building theoutbound request.
Proxyclones inbound headers, so the delete must runafter the clone.
set
x-stargate-last-cluster-id. On a miss, error, or timeout, omit it andcontinue.
Response path:
x-stargate-cluster-id, writekey -> cluster IDwith the TTL. Write when response headers arrive, notat stream end. Prefill has already started on that cluster.
response.
x-stargate-*response headers from client responses.Rollout
Order:
last_cluster_affinityoff. Itconsumes the header and ignores it.
STARGATE_LAST_CLUSTER_ENABLED=true.last_cluster_affinityper model.Before upgrading Stargate, review
pulsar-wait-and-widenconfigs without queue-SLOfields. Set
max_queuedandcache_affinity_wait_ms, or switch topulsar,to keep requests on the rank-1 cluster under load.
Do not turn on
last_cluster_affinitybefore step 2. Otherwise every requestis classified
newand loses its affinity wait.Acceptance criteria
Shortcut removal
pulsar-wait-and-widen.max_queued: 0, rank-1 cluster at its engineconcurrency limit, rank-2 cluster with a free slot: the request is not
sent to rank 1. Replace
without_queue_slo_keeps_full_primarywith thistest.
max_queued: 4and 1 request queued on rank 1: therequest goes to rank 1.
cache_affinity_wait_ms: 300andmax_queued: 0: thedecision is a timed wait until 300 ms, then the request widens.
Compatibility
last_cluster_affinityunset orfalse, routing decisions forwait-and-widenandpulsar-wait-and-widendo not change, with orwithout the header, except for the shortcut removal above. Existing
load-balancer tests pass with no changes to their assertions, except
the tests that encode the shortcut.
last_cluster_affinityonpower-of-n,round-robin,random, orpulsarfails startup with an error that names the fieldand the algorithm.
x-stargate-last-cluster-idis never forwarded to Pylon, for anyalgorithm or flag value.
New sessions
Run each criterion for both algorithms, with X = 300 ms and s = 0.1.
request goes to B at elapsed 0 with no routing wait.
flag-off algorithm picks at X = 0 and s = 1.0.
X = 0 and s = 1.0 for the same snapshot and elapsed time.
staledecisions equalnewdecisions for the same snapshot.Returning sessions
waits for B and is never sent to A, even when A has a free slot and B is
full.
(C, A). B is selectable only after X.
stay A and B.
pulsar-wait-and-widen, k = 1,band_widen_interval_ms> 0, rankingA, B, C, D, E, last cluster D: after X the open set is D, A, B. C opens
with the next band.
continues as if it were an ineligible rank-1 cluster.
the original order. The promotion is not cached.
snapshot make the same choice for the same request and header.
Cache-thrash regression test
crates/stargate/tests/suite/load_balancing.rs, runfor both algorithms with k = 1:
x-stargate-last-cluster-id: B.shows the test covers the behavior change.
Header parsing
400.x-cache-affinity-keyis absent, or whenwait-and-widenhas nocache_affinity_backend_selection_count.x-routing-methodselects an algorithmwithout
last_cluster_affinityenabled.Observability
stargate_routing_session_selections_totalwith labelsrouting_key,model,algorithm,session_state(new,returning,stale), andselection(primary,fallback). Forreturning,primarymeans the request went to the last cluster.Counters start at zero for every
session_stateandselectionpair.routing.session_state.debuglog line records the session state and whether the lastcluster was promoted, with request ID, model, and routing key.
Gateway
Tests live under
src/invocation-plane-services/llm-api-gateway.STARGATE_LAST_CLUSTER_ENABLED=false, outbound Stargate requestsand store contents do not change, except that inbound
x-stargate-last-cluster-idis always stripped.x-stargate-last-cluster-idnever reaches Stargate onthe
Complete,Stream, orProxypaths, enabled or not.prompt_cache_key: the first carries noheader. The first response returns
x-stargate-cluster-id: B. The secondrequest carries
x-stargate-last-cluster-id: B.conversation_idand withx-multi-turn-session-id.within the timeout plus 5 ms.
body.
visible to a lookup through replica 2.
STARGATE_LAST_CLUSTER_LOCAL_MAX_ENTRIESentries and evicts leastrecently used first.
works.
x-stargate-cluster-idis still stripped from client responses.llm_api_gateway_last_cluster_lookups_totalwith labelresult(
hit,miss,error,timeout,skipped) and counterllm_api_gateway_last_cluster_writes_totalwith labelresult(ok,error). Both start at zero for every label value.llm_api_gateway_last_cluster_lookup_duration_secondswithbuckets from 1 ms to 50 ms.
llm-api-gateway.last_cluster_lookup,llm-api-gateway.last_cluster_write) with aresultattribute anderror status on failure.
warnwith request ID, model, and routing key. Donot log the affinity key or session ID.
deploy/helm/llm-api-gatewayexposes the four env vars through values,with the defaults above.
End-to-end
last_cluster_affinityon reproduces the cache-thrash scenario above:after the first request lands on B and A drains, the next 10 requests in
the session go to B.
Documentation
docs/api-gateway-contract.md: add the header to the optional trustedheaders table, add the gateway rules above, and update the checklist.
docs/load-balancer-configuration.md: documentlast_cluster_affinity,session classification, the selection rules for both algorithms, and the
rollout order. Remove the statement that an eligible primary wins
immediately without queue-SLO fields, and add the migration guidance
from the rollout section.
src/invocation-plane-services/llm-api-gateway/README.mddocuments theenv vars, eligible session sources, and store behavior.
docs/overview/llm-gateway.mdlists the env vars and the rolloutorder.
Validation
stargate-benchagainstmock-dynamo, or the routing simulator fromtest(stargate): add discrete-event routing simulator #2223. Cover both algorithms, flag off and on. Post KV cache reuse, TTFT
p50 and p99, goodput, and the share of returning requests that left
their last cluster in the Pull Request.
pulsar-wait-and-widencase to the load-balancer microbenchmarkin
stargate-bench. Post per-decision latency with and without theshortcut.
Out of scope
pulsar,power-of-n,round-robin, andrandom.Open decisions
These defaults apply unless someone changes them before implementation starts:
x-stargate-last-cluster-id.last_cluster_affinity.