Repository navigation
Run model inference on remote GPU workers (bdmap2-4), off by default - #270
Merged
Merged
Conversation
aperson30
force-pushed
the
feat/remote-gpu-workers
branch
2 times, most recently
from
October 3, 2026 01:44
3a4db75 to
0299e71
Compare
Keeps bdmap1's GPU free so the website host stays low-load: a bdmap2-4 reboot never affects the site, a bdmap1 reboot takes it offline. Off by default (GPU_WORKERS_ENABLED). When enabled, each GPU model command from _tracked_run runs on the least loaded healthy worker in GPU_WORKER_HOSTS. Workers are health-checked in parallel (GPU present, not hot or throttled, GPU/CPU not busy, free unified memory, no foreign GPU process, not the web host, not cooling down after a failure). Then: shared-filesystem refusal (probe + network fs type), session upload, preflight (model files, cwd, tools, disk), run over SSH in its own process group with a max run time, results fetched back, worker copy deleted. SSH connections are reused per worker, so overhead is about the LAN copy time. Failure policy: the remote path never fails a job a local run would complete. Worker problems before or during the run fall back to the next worker, then to local; a remote run that fails, hangs or loses its results is re-run once locally, and the worker cools down if local succeeds. A user cancel at any stage kills the remote tree and never falls back. ShapeKit and MedIA-Agentic stay local. Adds scripts/sync_gpu_workers.sh (additive, low-priority runtime mirror that never writes into paths it did not create, with --check and --undo), deploy/GPU_WORKERS.md, .env.example entries and unit tests run in CI.
aperson30
force-pushed
the
feat/remote-gpu-workers
branch
from
October 3, 2026 01:52
0299e71 to
4108db7
Compare
This was referenced Oct 3, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Lets the website run model inference on bdmap2-4 so bdmap1's GPU stays free (bdmap1 rebooting takes the site offline; bdmap2-4 rebooting does not).
Off by default (
GPU_WORKERS_ENABLED=false): merging changes no behavior.How it works
services/gpu_workers.py, hooked into_tracked_runinauto_segmentor.py. Every model command already uses absolute/home/visitor/...paths, so each one runs unchanged on a worker:flask-server/scripts) to the same path there;Staging, post-processing, zipping, status and ETA stay on bdmap1, unchanged. ShapeKit and MedIA-Agentic always stay local. Jobs still run one at a time under the existing GPU lock.
Routing (health check per job)
A worker is used only if it is:
*_warm_server.py;All thresholds are configurable.
Fail-safe behavior
Rule: the remote path never fails a job that a local run would complete.
GPU_WORKER_MAX_RUN_SECONDS, or its results can't be fetched: re-run once locally. If local succeeds, the worker was at fault and cools down. If local also fails, the job fails exactly as today.job.json,.owner,auto_masks.zip), and no cross-host mtime comparison.Latency
Measured from bdmap1:
The shared connection is started separately with its output streams closed. A connection started implicitly by a client would keep that client's output pipes open, making the caller wait until it exits.
Maintenance
flask-server/scripts/sync_gpu_workers.sh [--check|--undo] [host]: mirrors the model runtimes (~43 GB) at the lowest CPU/IO priority on both ends, capped at 40 MB/s.nnUNetcheckout that has research edits.--undoremoves exactly that.--checkreports missing, stale or differing files./home, and a worker with less than 100 GB free.flask-server/deploy/GPU_WORKERS.md: setup, enable/disable, troubleshooting. Rollback is settingGPU_WORKERS_ENABLED=falseand restarting.Testing
tests/unit/test_gpu_workers.py, added to CI) cover: default-off behavior, routing (heat, throttling, load, cooldowns, least-loaded choice), fallback, failover, connection loss, local re-run after remote failure/hang/lost results, cancel at each stage, shared-filesystem refusal, cleanup guards, secret filtering and rawapi/../..session paths.bash -n.GPU_WORKERS.md).Next step (optional)
Starting warm predictors on the workers would make jobs faster than today.