Skip to content

llm-routing: an interrupted model download resume can fail permanently and leave an oversized partial file #2333

Description

@kristinapathak

Describe the bug

In charts/gguf-backend/files/download.py (deploy/helm/llm-routing/spark; #2331 moves it under recipes/ without changing it), a resumed download retries only on OSError. The resume checks raise other errors:

  • A response other than 206 raises AssertionError (Server refused resume).
  • A Content-Range that does not start at the offset raises AssertionError, and a missing Content-Range raises AttributeError on None.
  • The size check runs after a chunk is written, so an oversize response leaves a .partial file larger than the pinned size before raising AssertionError.

None of these are retried, so the download Job fails at once. After an oversize response, every later run fails at Partial file exceeds pinned size until someone deletes the partial file from the artifacts volume by hand.

Steps or code to reproduce bug

Serve a ranged request with a 200 response, with no Content-Range, or with more bytes than the pinned size remaining, then run the download phase twice.

Expected behavior

Discard the partial file and restart that file from offset 0, within the existing retry budget. Check the size before writing a chunk, so an oversize response never leaves an oversized partial file.

Additional context

Introduced in #2322. Found during review of #2331.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't workingllm-stackLLM Gateway + Stargate Request Router + Vanity Gatewayneeds-triageIssue or PR awaiting maintainer triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions