Skip to content

feat(cuinterpose): host-carrier storage for shared allocation bytes - #330

Draft
galletas1712 wants to merge 1 commit into
schwinns/cuinterpose-rust-memory-ipcfrom
schwinns/cuinterpose-rust-05-host-carrier
Draft

galletas1712 wants to merge 1 commit into
schwinns/cuinterpose-rust-memory-ipcfrom
schwinns/cuinterpose-rust-05-host-carrier

Conversation

@galletas1712

@galletas1712 galletas1712 commented Sep 16, 2026

Copy link
Copy Markdown
Contributor

Summary

Refs #295 (approved).

Save shared creator allocation bytes in a registered host arena that CRIU captures with process memory. Recreate backing and copy the bytes back during restore.

Review boundary

Host carriers are the only shim content implementation. Preserve the existing context grouping, copy cleanup, and resource handling. No storage abstraction, selector, or PageBroker client.

This PR contains one signed-off commit, b6da73194f2218f1ab18a11c75aaf8ce967ca4eb. Diff: 1 file changed, 552 insertions(+).

Stack and compatibility

The stack starts directly on main; it does not depend on the PageBroker GPU-transfer branch or #323. The shim only saves shared creator bytes through host carriers. Never-shared allocations remain native CUDA state. The stack removes launch-job/jobfile support, and older jobfile-dependent or draft shim artifacts are rejected rather than migrated.

PageBroker and native CustomStorage changes are separate. No PR in this stack adds that implementation.

Order PR Scope
1 #326 C preload frontend and private ABI
2 #327 typed identities and peer transport
3 #328 process identity and peer export service
4 #329 VMM ownership and allocation lifecycle
5 #342 implement CUDA memory IPC over tracked VMM
6 #330 host-carrier storage for shared allocation bytes
7 #331 multicast tracking and reconstruction
8 #332 Rust lifecycle coordinator
9 #333 package frontend, Rust backend, and coordinator
10 #334 deliver CUDA tools and enable opt-in preload
11 #335 coordinate cuinterpose capture and restore without jobfiles
12 #336 deliver CUDA tools without rewriting workload commands
13 #337 explain host carriers, memory IPC, and restore ordering
14 #338 unit, integration, and GPU lifecycle coverage

Validation

Validation of the assembled implementation:

  • Full API, agent, and operator Go tests; key agent race tests.
  • Rust workspace tests, strict Clippy, GNU/musl builds, ABI/ELF checks, packaged frontend/fake-driver tests, and memory-IPC regressions.
  • Helm tests/lint, Python manifest/report tests, repository lint, and make check in a clean disposable worktree.
  • A small real-agent CRIU cross-node test passed on two GPU nodes, covering memory IPC, private/shared VMM, multicast/graph replay, bytes, and original addresses. It preceded only the final manifest-format guard; that guard has local regression coverage.

GLM testing without CustomStorage was cancelled at the user's request and is not a pass. GLM qualification uses a separate composition with CustomStorage; previous experimental GLM results are not qualification of this rebuilt stack. The installed test driver is not claimed to be a stock-driver qualification.

Tests above were run on the assembled implementation, not claimed independently for every source-only intermediate PR. The standard-library cleanup was validated with Rust workspace tests and strict Clippy, GNU/musl builds and packaged fake-driver tests, plus targeted pod-contract and CUDA-agent Go tests. Physical-GPU tests were not rerun for this cleanup; the earlier GPU results above remain attributed to the pre-cleanup implementation.

@coderabbitai

coderabbitai Bot commented Sep 16, 2026

Copy link
Copy Markdown

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
  auto_review:
    drafts: true

Comment @coderabbitai help to get the list of available commands.

@copy-pr-bot

copy-pr-bot Bot commented Sep 17, 2026

Copy link
Copy Markdown

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

@galletas1712
galletas1712 removed this pull request from stack #339 September 17, 2026 08:53
@galletas1712
galletas1712 changed the base branch from schwinns/cuinterpose-rust-04-vmm to schwinns/cuinterpose-rust-memory-ipc September 17, 2026 08:53
@galletas1712
galletas1712 added this pull request to stack #343 September 17, 2026 08:53
@galletas1712 galletas1712 changed the title feat(cuinterpose): isolated host-carrier allocation storage feat(cuinterpose): host-carrier storage for shared allocation bytes Sep 17, 2026
@galletas1712
galletas1712 force-pushed the schwinns/cuinterpose-rust-05-host-carrier branch from c73b79b to b6da731 Compare September 17, 2026 09:30
@galletas1712
galletas1712 force-pushed the schwinns/cuinterpose-rust-05-host-carrier branch from b6da731 to 316d51d Compare September 17, 2026 20:53
@galletas1712
galletas1712 force-pushed the schwinns/cuinterpose-rust-05-host-carrier branch from 316d51d to 40b469d Compare September 17, 2026 22:39
@galletas1712
galletas1712 force-pushed the schwinns/cuinterpose-rust-05-host-carrier branch from 40b469d to 7da42d8 Compare September 18, 2026 01:13
@galletas1712
galletas1712 force-pushed the schwinns/cuinterpose-rust-05-host-carrier branch 2 times, most recently from 0d7a3c3 to 62f026b Compare September 18, 2026 03:41
Refs #295.

Signed-off-by: Schwinn Saereesitthipitak <schwinns@nvidia.com>
@galletas1712
galletas1712 force-pushed the schwinns/cuinterpose-rust-05-host-carrier branch from 62f026b to d90992c Compare September 18, 2026 18:55
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant