Skip to content

Run model inference on remote GPU workers (bdmap2-4), off by default - #270

Merged
aperson30 merged 1 commit into
BodyMaps:mainfrom
aperson30:feat/remote-gpu-workers
Oct 3, 2026
Merged

aperson30 merged 1 commit into
BodyMaps:mainfrom
aperson30:feat/remote-gpu-workers

Conversation

@aperson30

@aperson30 aperson30 commented Oct 3, 2026 •

Copy link
Copy Markdown
Collaborator

Lets the website run model inference on bdmap2-4 so bdmap1's GPU stays free (bdmap1 rebooting takes the site offline; bdmap2-4 rebooting does not).

Off by default (GPU_WORKERS_ENABLED=false): merging changes no behavior.

How it works

services/gpu_workers.py, hooked into _tracked_run in auto_segmentor.py. Every model command already uses absolute /home/visitor/... paths, so each one runs unchanged on a worker:

  1. health-check all workers in parallel and pick the least loaded healthy one;
  2. refuse it if it shares bdmap1's session filesystem, then rsync the session dir (and flask-server/scripts) to the same path there;
  3. preflight, then run the command over SSH in its own process group;
  4. fetch results back, delete the worker copy.

Staging, post-processing, zipping, status and ETA stay on bdmap1, unchanged. ShapeKit and MedIA-Agentic always stay local. Jobs still run one at a time under the existing GPU lock.

Routing (health check per job)

A worker is used only if it is:

  • reachable, with a working GPU;
  • not hot (below 85 °C) and not in a hardware or thermal slowdown;
  • GPU use below 20%, which catches jobs whose processes are hidden, e.g. in containers;
  • CPU load below 0.75 per core;
  • at least 40 GB free unified memory;
  • running no GPU process except our own *_warm_server.py;
  • not the web host;
  • not in a cooldown: 60 s after being unreachable, 10 min after failing a job.

All thresholds are configurable.

Fail-safe behavior

Rule: the remote path never fails a job that a local run would complete.

  • Worker problem before or during the run (busy, unreachable, missing files, reboot): next worker, then local fallback.
  • Remote run failed, hung past GPU_WORKER_MAX_RUN_SECONDS, or its results can't be fetched: re-run once locally. If local succeeds, the worker was at fault and cools down. If local also fails, the job fails exactly as today.
  • Unexpected error in the dispatch code: the remote run is killed and the job runs locally.
  • Cancel at any stage: kills the remote process tree; never falls back. It is only recorded for the run holding the GPU slot.
  • Shared filesystem: refused by a probe file plus a network-fs check. A refused worker is never cleaned up.
  • Result fetch: never copies back the files the web app writes itself (job.json, .owner, auto_masks.zip), and no cross-host mtime comparison.
  • Leftover worker copies (from a web restart mid-job) are removed after 2 days.
  • Secrets: model env vars are forwarded; anything that looks like a secret is not.

Latency

Measured from bdmap1:

  • SSH connections are reused per worker: 0.007 s per call, against 0.36 s for a new one.
  • Choosing a worker takes 0.10 s, or 0.7 s the first time while the connections open.
  • The session copy over the 100 Gbit/s LAN is typically 1-2 s.
  • Model time is unchanged: same GB10 hardware, and bdmap1's warm predictor is not running today, so both take the cold path.

The shared connection is started separately with its output streams closed. A connection started implicitly by a client would keep that client's output pipes open, making the caller wait until it exits.

Maintenance

  • flask-server/scripts/sync_gpu_workers.sh [--check|--undo] [host]: mirrors the model runtimes (~43 GB) at the lowest CPU/IO priority on both ends, capped at 40 MB/s.
    • It never writes into a path it did not create. On bdmap2 that protects an existing nnUNet checkout that has research edits.
    • It records what it creates, so --undo removes exactly that.
    • --check reports missing, stale or differing files.
    • It refuses bdmap1, a shared /home, and a worker with less than 100 GB free.
  • flask-server/deploy/GPU_WORKERS.md: setup, enable/disable, troubleshooting. Rollback is setting GPU_WORKERS_ENABLED=false and restarting.

Testing

  • 47 unit tests (tests/unit/test_gpu_workers.py, added to CI) cover: default-off behavior, routing (heat, throttling, load, cooldowns, least-loaded choice), fallback, failover, connection loss, local re-run after remote failure/hang/lost results, cancel at each stage, shared-filesystem refusal, cleanup guards, secret filtering and raw api/../.. session paths.
  • Read-only checks on the real nodes:
    • bdmap2 and bdmap4 pass the health check, tools/disk check and filesystem check (local btrfs);
    • bdmap1 is rejected as the web host and bdmap3 as busy;
    • the remote shell snippets pass bash -n.
  • Two rounds of independent adversarial review, both addressed.
  • Not yet run end-to-end on a worker; that needs the model runtimes mirrored first (see GPU_WORKERS.md).

Next step (optional)

Starting warm predictors on the workers would make jobs faster than today.

@aperson30
aperson30 force-pushed the feat/remote-gpu-workers branch 2 times, most recently from 3a4db75 to 0299e71 Compare October 3, 2026 01:44
Keeps bdmap1's GPU free so the website host stays low-load: a bdmap2-4
reboot never affects the site, a bdmap1 reboot takes it offline.

Off by default (GPU_WORKERS_ENABLED). When enabled, each GPU model command
from _tracked_run runs on the least loaded healthy worker in
GPU_WORKER_HOSTS. Workers are health-checked in parallel (GPU present, not
hot or throttled, GPU/CPU not busy, free unified memory, no foreign GPU
process, not the web host, not cooling down after a failure). Then:
shared-filesystem refusal (probe + network fs type), session upload,
preflight (model files, cwd, tools, disk), run over SSH in its own process
group with a max run time, results fetched back, worker copy deleted. SSH
connections are reused per worker, so overhead is about the LAN copy time.

Failure policy: the remote path never fails a job a local run would
complete. Worker problems before or during the run fall back to the next
worker, then to local; a remote run that fails, hangs or loses its results
is re-run once locally, and the worker cools down if local succeeds. A user
cancel at any stage kills the remote tree and never falls back. ShapeKit and
MedIA-Agentic stay local.

Adds scripts/sync_gpu_workers.sh (additive, low-priority runtime mirror that
never writes into paths it did not create, with --check and --undo),
deploy/GPU_WORKERS.md, .env.example entries and unit tests run in CI.
@aperson30
aperson30 force-pushed the feat/remote-gpu-workers branch from 0299e71 to 4108db7 Compare October 3, 2026 01:52
@aperson30
aperson30 merged commit 6a1ee81 into BodyMaps:main Oct 3, 2026
8 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant