Skip to content

Run several website model jobs at once, one per GPU worker (off by default) - #272

Merged
aperson30 merged 1 commit into
BodyMaps:mainfrom
aperson30:feat/parallel-gpu-jobs
Oct 3, 2026
Merged

aperson30 merged 1 commit into
BodyMaps:mainfrom
aperson30:feat/parallel-gpu-jobs

Conversation

@aperson30

@aperson30 aperson30 commented Oct 3, 2026 •

Copy link
Copy Markdown
Collaborator

Follow-up to #270/#271. Today one global lock makes every job wait, even when other GPUs are idle. This lets several jobs run at once, one per GPU worker.

Off by default. GPU_WORKER_PARALLEL defaults to 1, which is exactly today's behavior (one job at a time), so merging changes nothing until it is raised. It is capped at the number of workers and ignored (1) while workers are disabled or EPAI_REMOTE_ENABLED is on. Rollback: unset it or set it to 1, and reload.

What changes

  • auto_segmentor: the single _gpu_lock becomes _job_slots, a semaphore with max_parallel_jobs() slots (1 = the old lock).
  • bdmap1's own GPU gets its own lock (_local_gpu_lock), so it still runs at most one model at a time however many jobs are in flight. It is only taken when workers are enabled, so the disabled path is unchanged. Waiting for it is cancellable.
  • gpu_workers.acquire_host: reservation is now safe for concurrent jobs (see below), and a job that finds every usable worker busy with our own jobs waits (up to GPU_WORKER_MAX_WAIT_SECONDS, default 300, cancellable) instead of overflowing to bdmap1. bdmap1 is used only when no worker is up at all, or after that wait.
  • MedIA-Agentic (organs/vertebrae) runs on bdmap1's GPU directly, not through the worker path; it now takes the same bdmap1 GPU lock, since the old global lock used to cover it by accident.
  • The SSH connection master per worker is created under a lock.
  • The queue cap INFERENCE_MAX_PENDING defaults to 3 + N (4 at N=1, as before).

A bug this fixes before it can bite

The merged acquire_host reserved every idle worker while it health-checked them. With one job at a time that was harmless, but with two simultaneous jobs and two idle workers the second one saw "no worker free" and would have run on bdmap1's GPU. Replaying that on the merged code: ['None', 'w1']. Now workers are checked without being reserved and the winner is reserved atomically, so two jobs never get the same worker and never see a false "none free".

Independent review (5 rounds, 8 findings, all fixed)

  • P1 MedIA-Agentic bypassed the new bdmap1 GPU lock: now guarded.
  • P2 the busy-worker snapshot predated the health checks, so a worker taken meanwhile could be mistaken for "nothing to wait for": reservations are re-read after the checks.
  • P2 cancellation was only checked when the lock wait timed out: now also checked after the lock is acquired.
  • P2 (round 2) a repeat request for the same session could overlap the first (the old single lock used to serialize them): a session now never runs against itself, and a waiting duplicate does not use up a job slot.
  • P2 (round 2) the older EPAI_REMOTE_ENABLED single-host ePAI mode sends every ePAI job to one fixed GPU outside the pool: it now keeps jobs strictly serial. (It is not set on the live server.)
  • P2 (round 3) the serial-mode guard for EPAI_REMOTE_ENABLED used a yes/no parser that did not accept y, which the ePAI runner does: so EPAI_REMOTE_ENABLED=y plus parallel jobs would have sent two ePAI jobs to the same fixed GPU. There is now one parser for every environment flag, so the two cannot disagree.
  • P2 (round 4) the final reservation checked that a worker was unreserved but not that it was out of cooldown, so a worker another job had just failed on could be handed out on a stale health score. The commit step now re-checks both under the lock.
  • P2 (round 5) a worker another job released while this job was health-checking was not reconsidered: the job concluded "nothing to wait for" and fell back to bdmap1 although a worker had just become free. It now goes around again and checks the newly free worker first.
    Each has a regression test that fails on the code without its fix.

Testing

  • 90 unit tests in test_gpu_workers.py (42 new, stable over repeated runs): a stress test (8 jobs over 3 workers x 40 rounds: no worker ever held twice), two simultaneous jobs get two workers, waiting beats overflowing, the wait is bounded and cancellable, nothing-to-wait-for returns at once, an unhealthy free worker waits for the busy healthy one, max_parallel_jobs limits and bad values, slots bound concurrent models (1 = old behavior, 2 lets two in), bdmap1's GPU runs one model at a time even with parallel jobs, its wait is cancellable, and the disabled path never takes the lock.
  • Real run (3 ePAI jobs, N=2, bdmap2 + bdmap4, local fallback disabled): jobs 1 and 2 started 0.2 s apart on different workers and finished in about 49 s each, job 3 took the first freed worker; 92 s total against about 140 s one at a time (re-run on the final reviewed code). All three completed, no leftovers on any host, and bdmap1's GPU stayed empty.

Assumptions (documented in GPU_WORKERS.md)

Each worker has one GPU (placement is per worker); one gunicorn process (reservations are in-process); no strict first-come-first-served order between waiting jobs; reload only when no job is running.

Rollout plan

  1. Merge and deploy with GPU_WORKER_PARALLEL unset (no change).
  2. Set GPU_WORKER_PARALLEL=2 (bdmap2 + bdmap4; bdmap3 stays a spare) and reload.
  3. Test with two simultaneous uploads.

@aperson30
aperson30 force-pushed the feat/parallel-gpu-jobs branch 4 times, most recently from 4f73fc8 to e8c1058 Compare October 3, 2026 04:27
Today one global lock makes every job wait, even when other GPUs are idle.
GPU_WORKER_PARALLEL=N (default 1, capped by the number of workers, ignored
while workers are disabled or the older EPAI_REMOTE_ENABLED single-host mode
is on) lets up to N jobs run at once. With the default, behavior is unchanged.

Safety:
- Worker reservation: workers are health-checked without being reserved and
  the winner is reserved atomically afterwards. Previously the check reserved
  every idle worker, so with two simultaneous jobs the second saw "no worker
  free" and would have run on the web host's GPU. Reservations are re-read
  after the health checks, so a worker another job took meanwhile is waited
  for rather than mistaken for "nothing to wait for".
  The winner is committed only if, under the lock, it is still unreserved and
  not in cooldown, so a worker another job just failed on is never handed on a
  stale health score.
  A worker released during the health checks is reconsidered before waiting or
  falling back, so a job is not sent to the web host while a worker has just
  become free.
- A job that finds every usable worker busy with our own jobs waits (bounded
  by GPU_WORKER_MAX_WAIT_SECONDS, cancellable) instead of overflowing to the
  web host. It uses the web host only when no worker is up, or after the wait.
- The web host's own GPU runs at most one model at a time however many jobs
  are in flight (a lock of its own). This includes MedIA-Agentic, which runs
  on the web host's GPU directly and used to be covered only by the old global
  lock. Cancellation is checked after the lock is acquired as well as while
  waiting. Not taken when workers are disabled, so that path is unchanged.
- A session never runs against itself: a repeat request for a session whose
  job is queued or running waits for it (as it did behind the old lock)
  without using a job slot, instead of overlapping it in the same workspace.
- The older EPAI_REMOTE_ENABLED mode keeps jobs strictly serial. It is read
  with the same yes/no parser the ePAI runner uses, and there is now a single
  parser for every environment flag, so a spelling such as "y" cannot mean
  yes to one setting and no to another.
- The SSH connection master per worker is created under a lock.
- The queue cap (INFERENCE_MAX_PENDING) defaults to 3 + N (4 as before at 1).

Verified with a stress test (8 jobs over 3 workers: no double booking), unit
tests for waiting/cancel/limits and for every review finding (each fails on
the code without its fix), and a real run of 3 ePAI jobs with N=2 on bdmap2
and bdmap4: jobs 1 and 2 ran side by side, job 3 took the first free worker;
92 s total against about 140 s one at a time.
@aperson30
aperson30 force-pushed the feat/parallel-gpu-jobs branch from e8c1058 to c959b37 Compare October 3, 2026 04:43
@aperson30
aperson30 merged commit 04d8eec into BodyMaps:main Oct 3, 2026
8 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant