feat(snapshot): checkpoint POSIX CUDA multicast - #13249
Closed
galletas1712 wants to merge 1 commit into
Closed
galletas1712 wants to merge 1 commit into
galletas1712 wants to merge 1 commit into
Conversation
Contributor
|
galletas1712
force-pushed
the
feat/snapshot-cuda-vmm-multicast-posix
branch
from
August 14, 2026 10:05
2cbdc9f to
bdc5140
Compare
galletas1712
force-pushed
the
feat/snapshot-cuda-vmm-multicast-posix
branch
2 times, most recently
from
August 14, 2026 21:29
ee23bef to
bb284cc
Compare
galletas1712
force-pushed
the
feat/snapshot-cuda-vmm-multicast-posix
branch
from
August 14, 2026 23:47
bb284cc to
9bfee56
Compare
Signed-off-by: Schwinn Saereesitthipitak <schwinns@nvidia.com>
galletas1712
force-pushed
the
feat/snapshot-cuda-vmm-multicast-posix
branch
from
August 15, 2026 21:02
9bfee56 to
cc46e5f
Compare
This was referenced Aug 15, 2026
Contributor
Author
|
Closing: |
galletas1712
added a commit
to ai-dynamo/snapshot
that referenced
this pull request
Aug 25, 2026
Port of ai-dynamo/dynamo#13249. Extend the VMM interposer to capture and restore complete same-node CUDA multicast groups. Signed-off-by: Schwinn Saereesitthipitak <schwinns@nvidia.com>
galletas1712
added a commit
to ai-dynamo/snapshot
that referenced
this pull request
Aug 25, 2026
Port of ai-dynamo/dynamo#13249. Extend the VMM interposer to capture and restore complete same-node CUDA multicast groups. Signed-off-by: Schwinn Saereesitthipitak <schwinns@nvidia.com>
galletas1712
added a commit
to ai-dynamo/snapshot
that referenced
this pull request
Aug 27, 2026
Port of ai-dynamo/dynamo#13249. Extend the VMM interposer to capture and restore complete same-node CUDA multicast groups. Signed-off-by: Schwinn Saereesitthipitak <schwinns@nvidia.com>
galletas1712
added a commit
to ai-dynamo/snapshot
that referenced
this pull request
Aug 27, 2026
Port of ai-dynamo/dynamo#13249. Extend the VMM interposer to capture and restore complete same-node CUDA multicast groups. Signed-off-by: Schwinn Saereesitthipitak <schwinns@nvidia.com>
galletas1712
added a commit
to ai-dynamo/snapshot
that referenced
this pull request
Aug 27, 2026
Port of ai-dynamo/dynamo#13249. Extend the VMM interposer to capture and restore complete same-node CUDA multicast groups. Signed-off-by: Schwinn Saereesitthipitak <schwinns@nvidia.com>
galletas1712
added a commit
to ai-dynamo/snapshot
that referenced
this pull request
Aug 27, 2026
Port of ai-dynamo/dynamo#13249. Extend the VMM interposer to capture and restore complete same-node CUDA multicast groups. Signed-off-by: Schwinn Saereesitthipitak <schwinns@nvidia.com>
galletas1712
added a commit
to ai-dynamo/snapshot
that referenced
this pull request
Aug 27, 2026
Port of ai-dynamo/dynamo#13249. Extend the VMM interposer to capture and restore complete same-node CUDA multicast groups. Signed-off-by: Schwinn Saereesitthipitak <schwinns@nvidia.com>
galletas1712
added a commit
to ai-dynamo/snapshot
that referenced
this pull request
Aug 27, 2026
Port of ai-dynamo/dynamo#13249. Extend the VMM interposer to capture and restore complete same-node CUDA multicast groups. Signed-off-by: Schwinn Saereesitthipitak <schwinns@nvidia.com>
galletas1712
added a commit
to ai-dynamo/snapshot
that referenced
this pull request
Aug 27, 2026
Port of ai-dynamo/dynamo#13249. Extend the VMM interposer to capture and restore complete same-node CUDA multicast groups. Signed-off-by: Schwinn Saereesitthipitak <schwinns@nvidia.com>
galletas1712
added a commit
to ai-dynamo/snapshot
that referenced
this pull request
Aug 27, 2026
Port of ai-dynamo/dynamo#13249. Extend the VMM interposer to capture and restore complete same-node CUDA multicast groups. Signed-off-by: Schwinn Saereesitthipitak <schwinns@nvidia.com>
galletas1712
added a commit
to ai-dynamo/snapshot
that referenced
this pull request
Sep 1, 2026
Port of ai-dynamo/dynamo#13249. Extend the VMM interposer to capture and restore complete same-node CUDA multicast groups. Signed-off-by: Schwinn Saereesitthipitak <schwinns@nvidia.com>
galletas1712
added a commit
to ai-dynamo/snapshot
that referenced
this pull request
Sep 1, 2026
Port of ai-dynamo/dynamo#13249. Extend the VMM interposer to capture and restore complete same-node CUDA multicast groups. Signed-off-by: Schwinn Saereesitthipitak <schwinns@nvidia.com>
galletas1712
added a commit
to ai-dynamo/snapshot
that referenced
this pull request
Sep 1, 2026
Port of ai-dynamo/dynamo#13249. Extend the VMM interposer to capture and restore complete same-node CUDA multicast groups. Signed-off-by: Schwinn Saereesitthipitak <schwinns@nvidia.com>
galletas1712
added a commit
to ai-dynamo/snapshot
that referenced
this pull request
Sep 1, 2026
Port of ai-dynamo/dynamo#13249. Extend the VMM interposer to capture and restore complete same-node CUDA multicast groups. Signed-off-by: Schwinn Saereesitthipitak <schwinns@nvidia.com>
galletas1712
added a commit
to ai-dynamo/snapshot
that referenced
this pull request
Sep 1, 2026
Port of ai-dynamo/dynamo#13249. Extend the VMM interposer to capture and restore complete same-node CUDA multicast groups. Signed-off-by: Schwinn Saereesitthipitak <schwinns@nvidia.com>
galletas1712
added a commit
to ai-dynamo/snapshot
that referenced
this pull request
Sep 1, 2026
Port of ai-dynamo/dynamo#13249. Extend the VMM interposer to capture and restore complete same-node CUDA multicast groups. Signed-off-by: Schwinn Saereesitthipitak <schwinns@nvidia.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Add transparent checkpoint and restore for complete same-node CUDA multicast
groups backed by
CU_MEM_HANDLE_TYPE_POSIX_FILE_DESCRIPTOR, stacked on thePOSIX CUDA VMM interposer in #13129.
This PR covers the common PyTorch symmetric-memory and NCCL NVLS topology on one
node. It does not add FABRIC/IMEX/MNNVL, cross-node rendezvous, Redis, legacy
CUDA IPC wrappers, operator API, or Go-side CUDA topology knowledge.
Design
Resource model
allocations.
handles.
and each member process's local bindings, mappings, access, and logical-handle
liveness.
processes. Processes outside that object are phase no-ops.
property validation covers every member and property of each multicast
object without requiring unrelated processes to own it.
cuCheckpoint; multicast doesnot introduce another byte-saving path.
Ownership boundaries
interpose.cposix.cSCM_RIGHTSmulticast.cBindMem/BindAddr, mappings/access, detach, reconstruction, and validationcoordinator.cGo transports the opaque native sidecar and orders the existing native
checkpoint lifecycle. It does not enumerate or serialize CUDA multicast
topology.
Interception surface
Track and translate:
cuMulticastCreatecuMulticastAddDevicecuMulticastBindMemand CUDA 13.1_v2cuMulticastBindAddrand CUDA 13.1_v2cuMulticastUnbindcuMulticastGetGranularityaccess consumers
Cover direct Driver API calls plus:
dlsym()lookups fromlibcuda.soandlibcudart.so;cuGetProcAddress,cuGetProcAddress_v2, and_ptsz; andcudaGetDriverEntryPointandcudaGetDriverEntryPointByVersion, includingCUDA 13
_ptszexports.Checkpoint topology teardown
The coordinator first validates the complete object-local participant topology
and the existing unicast creator-anchor invariant. Every discovered process
enters the ordered phases, but a process that is not a member of any multicast
object is a multicast phase no-op.
Multicast is detached before unicast normalization:
importer handles; and
cuCheckpointsave unicast allocation bytes.The shim commits each multicast record's detached/checkpointed bookkeeping only
after the corresponding CUDA call succeeds. A failed CUDA teardown therefore
does not falsely advance native state.
Restore topology replay
After native CUDA restore and driver unlock, the application remains externally
quiesced while the shim reconstructs:
BindMemorBindAddrbinding;Nonmember processes remain no-ops throughout multicast reconstruction. CUDA
multicast bind may block until the full object-local team joins. The shim
therefore releases its bookkeeping mutex only around the native bind call so
creator endpoints remain serviceable, then reacquires it before committing
local state.
The coordinator writes opaque
snapshot-cuda-posix-v2state. It contains stableparticipant/resource identities and topology only—no raw FDs, authorization
tokens, creator endpoints, or allocation bytes.
Runtime contract
snapshot-cuda-vmm-launchorsnapshotctl --cuda-vmm-interpose.cuda-checkpoint --launch-jobpath. Manualsnapshotctlusers pass both--cuda-checkpoint-wrapand--cuda-vmm-interpose.preparation.
multicast detach, restore, and validation phases are no-ops in nonmembers.
each multicast object.
Meaningful CUDA tests
The colocated self-contained pytest project contains two real two-GPU
cuCheckpointscenarios:The multicast case:
BindMemduring symmetric-memory setup;cuMulticastBindAddr;coordinator restore;
restore.
No CRIU or PodSnapshot is required for this component gate; the tests call the
native
cuCheckpointAPI directly and run undercuda-checkpoint --launch-job.Earlier two-B200 qualification passed both component scenarios:
An earlier composed acceptance run also exercised DeepSeek V4 Flash TP4 with
multicast enabled, disjoint restore from GPUs 0-3 to GPUs 4-7, nonmember process
phase no-ops, three multicast objects, and coherent post-restore inference. That
campaign used independent composed-test fixes not included in this product PR.
GPU campaigns were not rerun for this restack.
Supported scope and limitations
Supported:
discovered processes;
NCCL NVLS;
BindMemandBindAddrcommon paths; andExplicitly excluded:
Potentially silent raw-FD alias, unshimmed broker, uncovered resolver,
raw-handle collision, and pass-through
BindAddrlimitations are documented inthe interposer README.
Validation
The multicast changes were replayed from only the two canonical commits:
git range-diffshows the second patch is unchanged. The first differs onlywhere it inherits #13129's simplified identity model and fixed-structure
reserved-byte layout.
Validated on the final restacked tree:
Additional passing checks:
-std=gnu11 -O2 -Wall -Wextra -Werror;libcudaorlibcudartdependency;pyproject.tomlparsing;git diff --check; anddeploy/snapshot/cmd/cuda-vmm-interpose/changed above feat(snapshot): checkpoint POSIX CUDA VMM sharing #13129.Live composed-stack acceptance
Build on Demand run
31908730437produced the exact composed SHAad5d5ebc6f4f30da9bc65ac7355a836702622328. That composition passed:three multicast objects, nonmember-process phase no-ops, restore from pinned
GPUs 0-3 to disjoint pinned GPUs 4-7 on
s2877, and three exactDSV4_RESTORE_OKresponses; andeight-GPU set on
s2877(not a migration test), with three deterministic,coherent HTTP 200 responses ending in
GLM52_RESTORE_OKafter restore.DeepSeek V4 is the direct live qualification of raw CUDA multicast/NCCL NVLS
replay. GLM does not make that claim: its accepted snapshot lifecycle policy
runtime-injected
NCCL_CUMEM_ENABLE=0,NCCL_NVLS_ENABLE=0, andNCCL_IB_DISABLE=1. GLM instead validates the wider composed checkpoint,restore, GMS, wake, and inference path around this stack.
No container image was built locally. GPU tests were not rerun during this
restack.
Stack
rewrite/snapshot-cuda-vmm-posix-native)