Repository navigation
feat(nvsnap): checkpoint and restore tensor-parallel workloads through gpushare - #2318
Merged
Merged
Conversation
The nvsnap Bazel row covers its Go code only, so libnvsnap_gpushare.so, nvsnap-gpu-suspend and the gpushare GPU tests were first compiled when the agent base image was built. - scripts/check-gpushare-build.sh compiles all of them with warnings as errors, once per platform, in the CUDA image Dockerfile.base's gpushare-builder stage uses (read from Dockerfile.base, so the two cannot drift), and fails if a binary needs a newer glibc than the documented 2.35. Needs Docker; no GPU. - The nvsnap gpushare workflow runs it for linux/amd64 and linux/arm64 (QEMU) when the C code, its build or the workflow changes. - make check-gpushare runs the same check locally. Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com> Signed-off-by: Balaji Ganesan <bganesan@nvidia.com>
… outputs present cp -r copies any binaries a developer built in the source tree, and make would treat them as up to date and skip the -Werror compile. make -B rebuilds every target regardless of timestamps. Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com> Signed-off-by: Balaji Ganesan <bganesan@nvidia.com>
…h gpushare Pods annotated nvsnap.io/gpushare=true run their GPU containers under libnvsnap_gpushare.so, and the agent drives nvsnap-gpu-suspend around its criu-v2 dump and restore, so multi-GPU workloads checkpoint with their default NCCL and all-reduce settings. - webhook: for annotated pods, mounts the node bundle at /nvsnap, appends the shim to LD_PRELOAD (keeping the pod's value; skips containers whose LD_PRELOAD comes from a reference), and mounts a per-pod chunk store from the node-local checkpoint disk at /var/run/nvsnap/gpushare. - capture: finds the dumped session's shim processes (refusing a GPU process without the shim), records the GPU map, suspends and stops them, dumps without the CUDA plugin and with --image-io-mode direct when the bundled CRIU supports it, resumes the source on --leave-running or after a failed dump, and moves the store into the checkpoint. Multi-GPU CRIU is accepted when the shim is loaded. One capture per container at a time. - restore: checks the placeholder provides the shim and the store, restores without the CUDA plugin, then resumes with the GPU map limited to the placeholder's allocated GPUs. - placeholders (agent-generated and e2e templates) mount the shim and the checkpoint's store; the agent-generated one now names the node path of the checkpoint directory instead of the agent's in-container path, and requests the checkpoint's GPU count. - e2e: vllm-small-gpushare (TP=1) and vllm-tp4-gpushare (TP=4) workloads; annotated workloads stay on criu-v2 at any GPU count. Validated by hand on GB300 (7B, TP=4, cross-node) through the same path before this change; the e2e runs follow. Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com> Signed-off-by: Balaji Ganesan <bganesan@nvidia.com>
Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com> Signed-off-by: Balaji Ganesan <bganesan@nvidia.com>
Contributor
Author
|
@coderabbitai review |
|
|
Important Review skippedAuto reviews are disabled on base/target branches other than the default branch. 🗂️ Base branches to auto review (1)
Please check the settings in the CodeRabbit UI or the ⚙️ Run configuration
You can disable this status message by setting the Use the checkbox below for a quick retry:
✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
Comment |
… workload's context - The chunk store moves from /var/run/nvsnap/gpushare to /nvsnap-gpushare. The agent reaches it through /proc/<pid>/root, and in the vLLM image /var/run is an absolute symlink to /run, which resolves against the agent's root there; the first TP=4 e2e capture reported the store missing although the webhook had mounted it. A test pins the path to a single top-level directory. - vllm-small-gpushare asked TinyLlama for a 4096-token context; its maximum is 2048. Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com> Signed-off-by: Balaji Ganesan <bganesan@nvidia.com>
The capture wrote the GPU map into the workload's store through /proc/<pid>/root, and creating a directory there fails with ENOENT (the store is a bind mount in the workload's mount namespace; reading through that path works, creating does not). The map now stays in the agent and is written into the checkpoint's gpushare directory after the store is collected; the restore finds it at <store>/ckpt/gpus because the placeholder mounts that directory at the store path. Covered by TestGPUShareCollectStore. Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com> Signed-off-by: Balaji Ganesan <bganesan@nvidia.com>
The capture passed --ckpt-dir <store>/ckpt, which the tool does not create and the agent cannot create inside the workload's mount namespace; every release failed with "open .../ckpt/gpu-<pid>.chunks: No such file or directory". The chunk lists now go in the store root, which the mount guarantees, and the GPU map sits beside them (<store>/gpus). Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com> Signed-off-by: Balaji Ganesan <bganesan@nvidia.com>
Open
4 tasks done
…eanups; suspend only GPU processes - scripts/checkpoint.sh wiped the node's checkpoint root before every capture (rm -rf .../checkpoints/*), which also deleted gpushare-pods/<uid>, the live store of the pod about to be captured: its mount then pointed at an unlinked directory and every release failed with ENOENT. The wipe now keeps gpushare-pods/, and the doc tells operators to do the same. - The capture suspended every session process that maps the shim. The API server and multiprocessing helpers inherit the preload but have no CUDA context, and the driver checkpoint API refuses them. Only processes that load the shim and use the GPU are suspended now; CRIU dumps the rest. Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com> Signed-off-by: Balaji Ganesan <bganesan@nvidia.com>
…ontext The app Dockerfile copied the cuda-checkpoint wrapper from the root of a build context that only scripts/build-agent.sh assembled. Building with the nvsnap directory as the context, as a release pipeline does, failed on that COPY. Reference the wrapper by its path in the tree and stage it at the same path in build-agent.sh. Also bring the Dockerfile's default base image up to the version versions.sh pins (v0.0.23), so a build without build args uses the current base. Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com> Signed-off-by: Balaji Ganesan <bganesan@nvidia.com>
…nap/gpushare-agent
versions.sh moved to base v0.0.24 with the gpushare review fixes. Keep the Dockerfile's BASE_IMAGE default on the same base, so a build without build args gets those fixes. Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com> Signed-off-by: Balaji Ganesan <bganesan@nvidia.com>
…nap/gpushare-agent
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
#2300 imported gpushare (libnvsnap_gpushare.so and nvsnap-gpu-suspend), which lets processes that share GPU memory (NCCL P2P and NVLS, CUDA IPC, page-locked host memory) be checkpointed with their default settings. Nothing in nvsnap used it yet. This wires it into the agent's criu-v2 capture and restore and the webhook, so a tensor-parallel vLLM pod can be captured and restored by nvsnap itself. A manual rehearsal of exactly this path on GB300 (7B, TP=4, cross-node) passed before this change.
What changed
nvsnap.io/gpushare: "true"get, on every container requesting GPUs, the node bundle at/nvsnap(read-only), the shim appended toLD_PRELOAD(keeping the pod's value; containers whoseLD_PRELOADcomes from a reference are skipped with a warning), and a per-pod chunk store from the node-local checkpoint disk at/var/run/nvsnap/gpushare(subPathExpr on the pod uid).checkpoint_v2.go,gpushare.go): selects the dumped session's processes that load the shim and refuses a GPU process without it; records the GPU map;suspendandstopinside the pod's namespaces; dumps without the CUDA plugin and with--image-io-mode directwhen the bundled CRIU supports it; resumes the source on--leave-runningor after a failed dump; moves the store into<checkpoint>/gpushare; records pids, store and shim paths in the metadata. Multi-GPU CRIU is accepted when the shim is loaded. One capture per container at a time (ErrCaptureInProgress).restore_v2.go): checks the placeholder provides the shim and the store, restores without the CUDA plugin, thenresume --gpu-maplimited to the placeholder's allocated GPUs (NVIDIA_VISIBLE_DEVICES).vllm-small-gpushare(TP=1) andvllm-tp4-gpushare(TP=4); annotated workloads stay on criu-v2 at any GPU count (harness and conformance test).docs/GPUSHARE.mdgains "Use through nvsnap" with the flow and the integration's limits.Customer Release Notes
Not customer visible yet (opt-in annotation, not enabled anywhere).
Plan Summary
No chart changes. New hostPath volumes on opted-in pods and placeholders only:
<checkpoint root>/gpushare-pods(DirectoryOrCreate) and the node bundle.Usage
Annotate the workload pod
nvsnap.io/gpushare: "true"; capture and restore as for any criu-v2 workload. Seedocs/GPUSHARE.md.Testing
go test ./...,golangci-lint --new-from-revclean,check-gazelleclean.Notes
Stacked on #2304 (#2300, #2261, #2106, #2104). Known limits listed in the doc: an image-only
LD_PRELOAD(DockerfileENV) is replaced; per-pod store directories are emptied on capture, not deleted with the pod; single-pod workloads only.While wiring this, Bazel drift across the stack was found and fixed on each branch (#2104, #2106, #2261): BUILD files had not been regenerated, and #2261 left a
pierrec/lz4/v4repo import inMODULE.bazelafter its last user was removed. Those checks only run on PRs to main.References
Relates to #2299
Related Pull Requests
#2300, #2304, #2261
Dependencies
None.
🤖 Generated with Claude Code
Summary by CodeRabbit