Skip to content

llm-routing: preflight and download Jobs cannot be rerun after a failure #2335

Description

@kristinapathak

Describe the bug

The model backend chart (deploy/helm/llm-routing/spark/charts/gguf-backend; #2331 moves it under recipes/) gives the preflight and download Jobs fixed names: <release>-preflight-<target> and <release>-download. After one of these Jobs fails, rerunning the phase applies the same Job name.

  • If the Job spec is unchanged, Helm keeps the failed Job and --wait-for-jobs fails again immediately.
  • If the spec changed, for example after a runtime image change, the upgrade fails because a Job's pod template is immutable.

Either way, the Job has to be deleted by hand before the phase can run again. The qualify and chain Jobs avoid this with attempt counters and --retry. #2331 fixes the same problem for the build Job by hashing its inputs into the name.

Steps or code to reproduce bug

  1. Make preflight or download fail, for example with a node that lacks free memory or a temporary download failure.
  2. Fix the cause and rerun the phase.

Expected behavior

A rerun after a failure creates a new Job, for example by recording an attempt number per phase as qualify does, or by deleting the previous failed Job owned by the release before applying.

Additional context

Introduced in #2322. Found during review of #2331.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't workingllm-stackLLM Gateway + Stargate Request Router + Vanity Gatewayneeds-triageIssue or PR awaiting maintainer triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions