Skip to content

feat(nvsnap): checkpoint and restore tensor-parallel workloads through gpushare - #2318

Merged
balajinvda merged 14 commits into
nvsnap/gpushare-importfrom
nvsnap/gpushare-agent
Oct 7, 2026
Merged

balajinvda merged 14 commits into
nvsnap/gpushare-importfrom
nvsnap/gpushare-agent

Conversation

@balajinvda

@balajinvda balajinvda commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Why

#2300 imported gpushare (libnvsnap_gpushare.so and nvsnap-gpu-suspend), which lets processes that share GPU memory (NCCL P2P and NVLS, CUDA IPC, page-locked host memory) be checkpointed with their default settings. Nothing in nvsnap used it yet. This wires it into the agent's criu-v2 capture and restore and the webhook, so a tensor-parallel vLLM pod can be captured and restored by nvsnap itself. A manual rehearsal of exactly this path on GB300 (7B, TP=4, cross-node) passed before this change.

What changed

  • Webhook: pods annotated nvsnap.io/gpushare: "true" get, on every container requesting GPUs, the node bundle at /nvsnap (read-only), the shim appended to LD_PRELOAD (keeping the pod's value; containers whose LD_PRELOAD comes from a reference are skipped with a warning), and a per-pod chunk store from the node-local checkpoint disk at /var/run/nvsnap/gpushare (subPathExpr on the pod uid).
  • Capture (checkpoint_v2.go, gpushare.go): selects the dumped session's processes that load the shim and refuses a GPU process without it; records the GPU map; suspend and stop inside the pod's namespaces; dumps without the CUDA plugin and with --image-io-mode direct when the bundled CRIU supports it; resumes the source on --leave-running or after a failed dump; moves the store into <checkpoint>/gpushare; records pids, store and shim paths in the metadata. Multi-GPU CRIU is accepted when the shim is loaded. One capture per container at a time (ErrCaptureInProgress).
  • Restore (restore_v2.go): checks the placeholder provides the shim and the store, restores without the CUDA plugin, then resume --gpu-map limited to the placeholder's allocated GPUs (NVIDIA_VISIBLE_DEVICES).
  • Placeholders: the agent-generated one and the e2e templates mount the shim and the checkpoint's store. Also fixes the agent-generated placeholder naming the agent's in-container checkpoint path as its hostPath (that path does not exist on the node) and requests the checkpoint's GPU count.
  • e2e: vllm-small-gpushare (TP=1) and vllm-tp4-gpushare (TP=4); annotated workloads stay on criu-v2 at any GPU count (harness and conformance test).
  • Docs: docs/GPUSHARE.md gains "Use through nvsnap" with the flow and the integration's limits.

Customer Release Notes

Not customer visible yet (opt-in annotation, not enabled anywhere).

Plan Summary

No chart changes. New hostPath volumes on opted-in pods and placeholders only: <checkpoint root>/gpushare-pods (DirectoryOrCreate) and the node bundle.

Usage

Annotate the workload pod nvsnap.io/gpushare: "true"; capture and restore as for any criu-v2 workload. See docs/GPUSHARE.md.

Testing

  • Unit tests: webhook placement through applied JSON patches (store, bundle, LD_PRELOAD append and idempotence, reference skip, bundle reuse, through Mutate); session process selection on a fake procfs (leader first, helpers ignored, GPU-without-shim refused, no-shim plain); dump and restore argv with and without gpushare and direct I/O; generated placeholder for gpushare and plain checkpoints (node paths, GPU count); GPU UUID parsing from environ.
  • go test ./..., golangci-lint --new-from-rev clean, check-gazelle clean.
  • Manual rehearsal of this path on GB300, 7B TP=4 cross-node: CRIU restore 24 GB in 6.4 s with direct I/O, resume 9.9 s, output matches.
  • e2e runs with an image from this branch (TP=1 regression, TP=4 on GB300) to follow on this PR.

Notes

Stacked on #2304 (#2300, #2261, #2106, #2104). Known limits listed in the doc: an image-only LD_PRELOAD (Dockerfile ENV) is replaced; per-pod store directories are emptied on capture, not deleted with the pod; single-pod workloads only.

While wiring this, Bazel drift across the stack was found and fixed on each branch (#2104, #2106, #2261): BUILD files had not been regenerated, and #2261 left a pierrec/lz4/v4 repo import in MODULE.bazel after its last user was removed. Those checks only run on PRs to main.

References

Relates to #2299

Related Pull Requests

#2300, #2304, #2261

Dependencies

None.

🤖 Generated with Claude Code

Summary by CodeRabbit

  • New Features
    • Added GPU-sharing support for checkpointing and restoring GPU workloads, including workloads using multiple GPUs.
    • Added sample vLLM workloads for single-GPU and four-GPU serving, with matching restore configurations.
  • Documentation
    • Added guidance on GPU-sharing setup, checkpoint and restore requirements, and configuration limitations.

The nvsnap Bazel row covers its Go code only, so libnvsnap_gpushare.so,
nvsnap-gpu-suspend and the gpushare GPU tests were first compiled when the
agent base image was built.

- scripts/check-gpushare-build.sh compiles all of them with warnings as
  errors, once per platform, in the CUDA image Dockerfile.base's
  gpushare-builder stage uses (read from Dockerfile.base, so the two cannot
  drift), and fails if a binary needs a newer glibc than the documented
  2.35. Needs Docker; no GPU.
- The nvsnap gpushare workflow runs it for linux/amd64 and linux/arm64
  (QEMU) when the C code, its build or the workflow changes.
- make check-gpushare runs the same check locally.

Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
Signed-off-by: Balaji Ganesan <bganesan@nvidia.com>
… outputs present

cp -r copies any binaries a developer built in the source tree, and make would
treat them as up to date and skip the -Werror compile. make -B rebuilds every
target regardless of timestamps.

Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
Signed-off-by: Balaji Ganesan <bganesan@nvidia.com>
…h gpushare

Pods annotated nvsnap.io/gpushare=true run their GPU containers under
libnvsnap_gpushare.so, and the agent drives nvsnap-gpu-suspend around its
criu-v2 dump and restore, so multi-GPU workloads checkpoint with their
default NCCL and all-reduce settings.

- webhook: for annotated pods, mounts the node bundle at /nvsnap, appends
  the shim to LD_PRELOAD (keeping the pod's value; skips containers whose
  LD_PRELOAD comes from a reference), and mounts a per-pod chunk store from
  the node-local checkpoint disk at /var/run/nvsnap/gpushare.
- capture: finds the dumped session's shim processes (refusing a GPU
  process without the shim), records the GPU map, suspends and stops them,
  dumps without the CUDA plugin and with --image-io-mode direct when the
  bundled CRIU supports it, resumes the source on --leave-running or after
  a failed dump, and moves the store into the checkpoint. Multi-GPU CRIU is
  accepted when the shim is loaded. One capture per container at a time.
- restore: checks the placeholder provides the shim and the store, restores
  without the CUDA plugin, then resumes with the GPU map limited to the
  placeholder's allocated GPUs.
- placeholders (agent-generated and e2e templates) mount the shim and the
  checkpoint's store; the agent-generated one now names the node path of
  the checkpoint directory instead of the agent's in-container path, and
  requests the checkpoint's GPU count.
- e2e: vllm-small-gpushare (TP=1) and vllm-tp4-gpushare (TP=4) workloads;
  annotated workloads stay on criu-v2 at any GPU count.

Validated by hand on GB300 (7B, TP=4, cross-node) through the same path
before this change; the e2e runs follow.

Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
Signed-off-by: Balaji Ganesan <bganesan@nvidia.com>
Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
Signed-off-by: Balaji Ganesan <bganesan@nvidia.com>
@balajinvda
balajinvda requested a review from a team as a code owner October 6, 2026 14:25
@balajinvda
balajinvda requested a review from apartha-nv October 6, 2026 14:25
@balajinvda

Copy link
Copy Markdown
Contributor Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Oct 6, 2026 •

Copy link
Copy Markdown
⚠️ Action not completed

Pull request base or head changed.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@coderabbitai

coderabbitai Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

🗂️ Base branches to auto review (1)
  • main

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration
  • Configuration used: Repository: NVIDIA/nvcf/.coderabbit.yaml
  • Review profile: CHILL
  • Plan: Enterprise
  • Run ID: aac81634-6c86-4660-a4a7-6cf39f89f03e

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
✨ Finishing Touches
📝 Generate docstrings
  • Commit to this branch
  • Create a new PR
🧪 Generate unit tests (beta)
  • Commit to this branch
  • Create a new PR
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Comment @coderabbitai help to get the list of available commands.

Base automatically changed from nvsnap/gpushare-ci to nvsnap/gpushare-import October 6, 2026 14:30
@FamousDirector
FamousDirector requested a review from a team as a code owner October 6, 2026 14:30
… workload's context

- The chunk store moves from /var/run/nvsnap/gpushare to /nvsnap-gpushare.
  The agent reaches it through /proc/<pid>/root, and in the vLLM image
  /var/run is an absolute symlink to /run, which resolves against the
  agent's root there; the first TP=4 e2e capture reported the store
  missing although the webhook had mounted it. A test pins the path to a
  single top-level directory.
- vllm-small-gpushare asked TinyLlama for a 4096-token context; its
  maximum is 2048.

Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
Signed-off-by: Balaji Ganesan <bganesan@nvidia.com>
The capture wrote the GPU map into the workload's store through
/proc/<pid>/root, and creating a directory there fails with ENOENT (the
store is a bind mount in the workload's mount namespace; reading through
that path works, creating does not). The map now stays in the agent and is
written into the checkpoint's gpushare directory after the store is
collected; the restore finds it at <store>/ckpt/gpus because the
placeholder mounts that directory at the store path. Covered by
TestGPUShareCollectStore.

Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
Signed-off-by: Balaji Ganesan <bganesan@nvidia.com>
The capture passed --ckpt-dir <store>/ckpt, which the tool does not create
and the agent cannot create inside the workload's mount namespace; every
release failed with "open .../ckpt/gpu-<pid>.chunks: No such file or
directory". The chunk lists now go in the store root, which the mount
guarantees, and the GPU map sits beside them (<store>/gpus).

Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
Signed-off-by: Balaji Ganesan <bganesan@nvidia.com>
…eanups; suspend only GPU processes

- scripts/checkpoint.sh wiped the node's checkpoint root before every
  capture (rm -rf .../checkpoints/*), which also deleted gpushare-pods/<uid>,
  the live store of the pod about to be captured: its mount then pointed at
  an unlinked directory and every release failed with ENOENT. The wipe now
  keeps gpushare-pods/, and the doc tells operators to do the same.
- The capture suspended every session process that maps the shim. The API
  server and multiprocessing helpers inherit the preload but have no CUDA
  context, and the driver checkpoint API refuses them. Only processes that
  load the shim and use the GPU are suspended now; CRIU dumps the rest.

Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
Signed-off-by: Balaji Ganesan <bganesan@nvidia.com>
…ontext

The app Dockerfile copied the cuda-checkpoint wrapper from the root of a
build context that only scripts/build-agent.sh assembled. Building with
the nvsnap directory as the context, as a release pipeline does, failed
on that COPY. Reference the wrapper by its path in the tree and stage it
at the same path in build-agent.sh.

Also bring the Dockerfile's default base image up to the version
versions.sh pins (v0.0.23), so a build without build args uses the
current base.

Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
Signed-off-by: Balaji Ganesan <bganesan@nvidia.com>
versions.sh moved to base v0.0.24 with the gpushare review fixes. Keep the
Dockerfile's BASE_IMAGE default on the same base, so a build without
build args gets those fixes.

Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
Signed-off-by: Balaji Ganesan <bganesan@nvidia.com>
@balajinvda
balajinvda merged commit 477cccc into nvsnap/gpushare-import Oct 7, 2026
3 checks passed
@balajinvda
balajinvda deleted the nvsnap/gpushare-agent branch October 7, 2026 18:23
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant