Skip to content

feat(stargate,llm-api-gateway): route returning sessions to their last cluster in wait-and-widen and pulsar-wait-and-widen #2289

Description

@barrygreengus

Summary

Add a session-aware mode to wait-and-widen and pulsar-wait-and-widen, and
the matching LLM API gateway support. The gateway remembers which cluster
served each session's last request and sends it back to Stargate on the next
request. Stargate uses that hint to tell first requests apart from follow-up
requests:

  • A new session (no hint) skips the affinity wait and the affinity prefill
    discount. It goes to the best-ranked cluster that is selectable now and does
    not hold for its rank-1 cluster.
  • A returning session (hint present) moves the hinted cluster to the front of
    its affinity order. It then runs the configured affinity wait, discount, and
    widening unchanged.

It also removes the pulsar-wait-and-widen shortcut that sends a request to
its rank-1 cluster without any load check when no queue-SLO fields are set.

This keeps the stable affinity order and lets a new session overflow quickly
and deterministically. It also stops a session from moving back to rank 1 once
that cluster frees up. That move throws away the KV cache the session built on
the overflow cluster.

Dependencies

The Stargate work builds on two open Pull Requests. Start it after both merge.
The gateway work does not depend on them.

Problem

With an affinity wait, a request waits for its affinity group before
overflowing. That suits a returning session, whose KV cache is on the
affinity group. It hurts a new session, which has no cache anywhere and gains
nothing from waiting.

Overflow also causes cache thrash. When the rank-1 cluster is busy, a session
overflows to a lower-ranked cluster and builds KV cache there. When rank 1
recovers, the next request routes back to rank 1 and misses the cache. Under
steady load this repeats.

Stargate cannot tell new from returning sessions alone. It keeps no
per-session state and does not share routing state between pods. The gateway
already sees x-stargate-cluster-id on every Stargate response and runs an
embedded Olric cluster shared by its replicas, so it can carry the last
cluster forward.

Terms

  • Session: one (x-routing-key, x-model, x-cache-affinity-key) tuple. In
    the gateway, these are RoutingKey, Model, and CacheAffinityKey on the
    request context.
  • Last cluster: the x-stargate-cluster-id value from the session's most
    recent 2xx response.
  • Affinity order: the ordered cluster list the algorithm derives from the
    affinity key.
    • wait-and-widen: the hash ring's selected clusters, in ring-walk order.
    • pulsar-wait-and-widen: the full Pulsar ranking.
  • Affinity group: the first cache_affinity_backend_selection_count (k)
    clusters of the affinity order.
  • Affinity wait (X): cache_affinity_wait_ms.
  • Affinity discount (s): cache_affinity_input_tokens_scale.

Proposed design

Request contract

Add one trusted request header:

Header Set by Value
x-stargate-last-cluster-id Gateway The x-stargate-cluster-id from the session's last 2xx response.

Use the cluster ID, not a rank index. A cluster ID is valid for both the hash
ring and the Pulsar ranking and survives candidate-set changes.

Stargate consumes the header and does not forward it to Pylon. Response
headers do not change.

Remove the eligible-primary shortcut

Today pulsar-wait-and-widen returns the rank-1 cluster immediately when the
config has no queue-SLO fields (max_queue_time_floor_ms,
max_queue_time_ceil_ms) and the cluster passes retry exclusion and the KV
free-token check. That check ignores load, so a full rank-1 cluster keeps
taking requests with no limit. A hot key then slows every other key whose
rank-1 cluster it shares.

Remove the shortcut for every request, regardless of last_cluster_affinity.
The rank-1 cluster then goes through the affinity phase like any other
request: queue admission with max_queued, the affinity wait, and the
affinity discount. Deployments that want unlimited stickiness use pulsar.
Deployments that want bounded stickiness raise the affinity group's
max_queued and set cache_affinity_wait_ms.

Behavior change: a no-SLO config left at the defaults (max_queued: 0,
cache_affinity_wait_ms: 0) overflows as soon as its rank-1 cluster has no
free engine slot.

Stargate config

Add a boolean algorithm field, last_cluster_affinity, default false, to
wait-and-widen and pulsar-wait-and-widen. Reject it at startup for every
other algorithm.

Session classification

When last_cluster_affinity is true, classify each request once, before
the first routing attempt. Retries keep the first classification.

Condition Session state
No affinity group: x-cache-affinity-key absent, or wait-and-widen without k Not applicable. Existing behavior. Header ignored.
Header absent, blank, or not valid UTF-8 new
Header names a cluster not in the current candidate set stale (routed as new)
Header names a cluster in the current candidate set returning

Selection rules

new and stale, both algorithms:

  • Treat X as 0 and s as 1.0 for this request.
  • Everything else runs as configured: affinity order, k, TTFT buckets, queue
    admission, band_widen_interval_ms, and fallback_max_queued.

returning, both algorithms:

  • Move the last cluster to position 1 of the affinity order. Keep the relative
    order of every other cluster.
  • The affinity group is the first k clusters of the reordered order. Its size
    stays k. If the last cluster was outside the original group, the original
    group's last member drops out.
  • pulsar-wait-and-widen widening bands come from the reordered ranking.
    wait-and-widen global selection is unchanged because it already covers
    every cluster.
  • X, s, and every other setting apply as configured.

Retry exclusion and KV free-token checks apply to the last cluster the same
way they apply to any other candidate. Per-key ring-selection and ranking
caches keep storing the original order. The promotion is per request.

Gateway changes

Code lives in src/invocation-plane-services/llm-api-gateway. Chart wiring
lives in deploy/helm/llm-api-gateway.

Configuration, all off or unset by default:

Env var Default Effect
STARGATE_LAST_CLUSTER_ENABLED false Turns lookup, header, and store writes on.
STARGATE_LAST_CLUSTER_TTL 10m Entry lifetime, reset on every write.
STARGATE_LAST_CLUSTER_LOOKUP_TIMEOUT 20ms Lookup deadline on the request path.
STARGATE_LAST_CLUSTER_LOCAL_MAX_ENTRIES 100000 Size cap for the in-process store when Olric is off.

Eligible sessions:

  • Only sessions whose SessionSource is prompt_cache_key,
    conversation_id, or header.
  • Sessions with SessionSource payload are skipped for lookup and writes.
    Their key is a hash of the full message list, which changes every turn, so
    they never return.

Store:

  • With OLRIC_ENABLED=true, use a dedicated DMap named
    stargate-last-cluster on the gateway's embedded Olric node. Start the
    Olric node whenever Olric is enabled, not only when the rate limiter is on.
    Today newRateLimiter in server/server.go owns the node.
  • With Olric off, use an in-process LRU capped at
    STARGATE_LAST_CLUSTER_LOCAL_MAX_ENTRIES, with the same TTL.
  • Key: lc:v1: plus the hex SHA-256 of the length-prefixed routing key,
    model, and cache affinity key. Raw routing keys never appear in store keys.
  • Value: the cluster ID. Values over 256 bytes are not stored.

Request path, for every Stargate call (Complete, Stream, and Proxy in
provider/stargate.go):

  1. Delete any inbound x-stargate-last-cluster-id before building the
    outbound request. Proxy clones inbound headers, so the delete must run
    after the clone.
  2. For eligible sessions, look up the key with the lookup timeout. On a hit,
    set x-stargate-last-cluster-id. On a miss, error, or timeout, omit it and
    continue.

Response path:

  1. On a 2xx Stargate response with a non-empty x-stargate-cluster-id, write
    key -> cluster ID with the TTL. Write when response headers arrive, not
    at stream end. Prefill has already started on that cluster.
  2. Writes run off the request path. They never delay or fail the client
    response.
  3. Do not write on non-2xx responses or transport errors.
  4. Concurrent requests in one session may both miss. Last write wins.
  5. Keep stripping x-stargate-* response headers from client responses.

Rollout

Order:

  1. Deploy Stargate with this change and last_cluster_affinity off. It
    consumes the header and ignores it.
  2. Deploy the gateway with STARGATE_LAST_CLUSTER_ENABLED=true.
  3. Turn on last_cluster_affinity per model.

Before upgrading Stargate, review pulsar-wait-and-widen configs without queue-SLO
fields. Set max_queued and cache_affinity_wait_ms, or switch to pulsar,
to keep requests on the rank-1 cluster under load.

Do not turn on last_cluster_affinity before step 2. Otherwise every request
is classified new and loses its affinity wait.

Acceptance criteria

Shortcut removal

  • The eligible-primary shortcut is gone from pulsar-wait-and-widen.
  • No queue-SLO fields, max_queued: 0, rank-1 cluster at its engine
    concurrency limit, rank-2 cluster with a free slot: the request is not
    sent to rank 1. Replace without_queue_slo_keeps_full_primary with this
    test.
  • Same setup with max_queued: 4 and 1 request queued on rank 1: the
    request goes to rank 1.
  • Same setup with cache_affinity_wait_ms: 300 and max_queued: 0: the
    decision is a timed wait until 300 ms, then the request widens.

Compatibility

  • With last_cluster_affinity unset or false, routing decisions for
    wait-and-widen and pulsar-wait-and-widen do not change, with or
    without the header, except for the shortcut removal above. Existing
    load-balancer tests pass with no changes to their assertions, except
    the tests that encode the shortcut.
  • Setting last_cluster_affinity on power-of-n, round-robin,
    random, or pulsar fails startup with an error that names the field
    and the algorithm.
  • x-stargate-last-cluster-id is never forwarded to Pylon, for any
    algorithm or flag value.

New sessions

Run each criterion for both algorithms, with X = 300 ms and s = 0.1.

  • Rank-1 cluster A is full and rank-2 cluster B has a free slot: the
    request goes to B at elapsed 0 with no routing wait.
  • Every cluster has a free slot: the request goes to the cluster that the
    flag-off algorithm picks at X = 0 and s = 1.0.
  • The TTFT used for bucket placement and comparison uses s = 1.0, not 0.1.
  • No candidate is selectable: the decision equals the flag-off decision at
    X = 0 and s = 1.0 for the same snapshot and elapsed time.
  • stale decisions equal new decisions for the same snapshot.

Returning sessions

  • k = 1, original order A, B, C, last cluster B: during X, the request
    waits for B and is never sent to A, even when A has a free slot and B is
    full.
  • k = 1, last cluster B: s applies to B, not to A.
  • k = 2, original group (A, B), last cluster C: the affinity group is
    (C, A). B is selectable only after X.
  • k = 2, original group (A, B), last cluster B: the affinity group members
    stay A and B.
  • pulsar-wait-and-widen, k = 1, band_widen_interval_ms > 0, ranking
    A, B, C, D, E, last cluster D: after X the open set is D, A, B. C opens
    with the next band.
  • Last cluster retry-excluded or failing the KV free-token check: selection
    continues as if it were an ineligible rank-1 cluster.
  • After a promoted request, a request for the same key with no header uses
    the original order. The promotion is not cached.
  • Two Stargate instances with the same seed, config, and candidate
    snapshot make the same choice for the same request and header.

Cache-thrash regression test

  • Integration test in crates/stargate/tests/suite/load_balancing.rs, run
    for both algorithms with k = 1:
    1. Session S has rank 1 on A and rank 2 on B. A is full.
    2. The first request with no header goes to B.
    3. A drains and has free slots.
    4. The next 10 requests send x-stargate-last-cluster-id: B.
    5. All 10 go to B.
  • With the flag off, the same test sends the step 4 requests to A, which
    shows the test covers the behavior change.

Header parsing

  • Blank or non-UTF-8 header values are treated as absent and do not return
    400.
  • The header has no effect when x-cache-affinity-key is absent, or when
    wait-and-widen has no cache_affinity_backend_selection_count.
  • The header has no effect when x-routing-method selects an algorithm
    without last_cluster_affinity enabled.

Observability

  • New counter stargate_routing_session_selections_total with labels
    routing_key, model, algorithm, session_state (new,
    returning, stale), and selection (primary, fallback). For
    returning, primary means the request went to the last cluster.
    Counters start at zero for every session_state and selection pair.
  • The proxy request span records routing.session_state.
  • A debug log line records the session state and whether the last
    cluster was promoted, with request ID, model, and routing key.

Gateway

Tests live under src/invocation-plane-services/llm-api-gateway.

  • With STARGATE_LAST_CLUSTER_ENABLED=false, outbound Stargate requests
    and store contents do not change, except that inbound
    x-stargate-last-cluster-id is always stripped.
  • A client-supplied x-stargate-last-cluster-id never reaches Stargate on
    the Complete, Stream, or Proxy paths, enabled or not.
  • Two requests with the same prompt_cache_key: the first carries no
    header. The first response returns x-stargate-cluster-id: B. The second
    request carries x-stargate-last-cluster-id: B.
  • Same flow with conversation_id and with x-multi-turn-session-id.
  • Payload-derived sessions never trigger a lookup or a write.
  • Same affinity key with a different model or routing key: no header.
  • Non-2xx responses and transport errors do not change the stored value.
  • A 2xx streaming response writes the value before the stream ends.
  • An entry older than the TTL is a miss. A write resets the TTL.
  • Lookup slower than the timeout: the request proceeds without the header
    within the timeout plus 5 ms.
  • Store errors on lookup or write do not change the client status code or
    body.
  • Two gateway replicas in one Olric cluster: a write through replica 1 is
    visible to a lookup through replica 2.
  • Olric off: the in-process store holds at most
    STARGATE_LAST_CLUSTER_LOCAL_MAX_ENTRIES entries and evicts least
    recently used first.
  • Olric on and rate limiter off: the Olric node starts and the store
    works.
  • x-stargate-cluster-id is still stripped from client responses.
  • Counter llm_api_gateway_last_cluster_lookups_total with label result
    (hit, miss, error, timeout, skipped) and counter
    llm_api_gateway_last_cluster_writes_total with label result (ok,
    error). Both start at zero for every label value.
  • Histogram llm_api_gateway_last_cluster_lookup_duration_seconds with
    buckets from 1 ms to 50 ms.
  • Lookup and write each create a child span (llm-api-gateway.last_cluster_lookup,
    llm-api-gateway.last_cluster_write) with a result attribute and
    error status on failure.
  • Store errors log at warn with request ID, model, and routing key. Do
    not log the affinity key or session ID.
  • deploy/helm/llm-api-gateway exposes the four env vars through values,
    with the defaults above.

End-to-end

  • A test that runs the gateway against a real Stargate with
    last_cluster_affinity on reproduces the cache-thrash scenario above:
    after the first request lands on B and A drains, the next 10 requests in
    the session go to B.

Documentation

  • docs/api-gateway-contract.md: add the header to the optional trusted
    headers table, add the gateway rules above, and update the checklist.
  • docs/load-balancer-configuration.md: document last_cluster_affinity,
    session classification, the selection rules for both algorithms, and the
    rollout order. Remove the statement that an eligible primary wins
    immediately without queue-SLO fields, and add the migration guidance
    from the rollout section.
  • The request-router metrics reference lists the new counter.
  • src/invocation-plane-services/llm-api-gateway/README.md documents the
    env vars, eligible session sources, and store behavior.
  • docs/overview/llm-gateway.md lists the env vars and the rollout
    order.

Validation

  • Run a multi-turn workload with a saturated rank-1 cluster through
    stargate-bench against mock-dynamo, or the routing simulator from
    test(stargate): add discrete-event routing simulator #2223. Cover both algorithms, flag off and on. Post KV cache reuse, TTFT
    p50 and p99, goodput, and the share of returning requests that left
    their last cluster in the Pull Request.
  • Add a pulsar-wait-and-widen case to the load-balancer microbenchmark
    in stargate-bench. Post per-decision latency with and without the
    shortcut.

Out of scope

Open decisions

These defaults apply unless someone changes them before implementation starts:

  1. Header name: x-stargate-last-cluster-id.
  2. Config field name: last_cluster_affinity.
  3. Gateway TTL default: 10 minutes.
  4. Gateway lookup timeout default: 20 ms.
  5. Gateway in-process store cap: 100000 entries.
  6. Gateway updates only on 2xx responses, at response-header time.

Activity

  1. changed the title [-]feat(stargate): route returning sessions to their last cluster in wait-and-widen and pulsar-wait-and-widen[/-] [+]feat(stargate,llm-api-gateway): route returning sessions to their last cluster in wait-and-widen and pulsar-wait-and-widen[/+] on Oct 5, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions