Repository navigation
Run several website model jobs at once, one per GPU worker (off by default) - #272
Merged
Merged
Conversation
aperson30
force-pushed
the
feat/parallel-gpu-jobs
branch
4 times, most recently
from
October 3, 2026 04:27
4f73fc8 to
e8c1058
Compare
Today one global lock makes every job wait, even when other GPUs are idle. GPU_WORKER_PARALLEL=N (default 1, capped by the number of workers, ignored while workers are disabled or the older EPAI_REMOTE_ENABLED single-host mode is on) lets up to N jobs run at once. With the default, behavior is unchanged. Safety: - Worker reservation: workers are health-checked without being reserved and the winner is reserved atomically afterwards. Previously the check reserved every idle worker, so with two simultaneous jobs the second saw "no worker free" and would have run on the web host's GPU. Reservations are re-read after the health checks, so a worker another job took meanwhile is waited for rather than mistaken for "nothing to wait for". The winner is committed only if, under the lock, it is still unreserved and not in cooldown, so a worker another job just failed on is never handed on a stale health score. A worker released during the health checks is reconsidered before waiting or falling back, so a job is not sent to the web host while a worker has just become free. - A job that finds every usable worker busy with our own jobs waits (bounded by GPU_WORKER_MAX_WAIT_SECONDS, cancellable) instead of overflowing to the web host. It uses the web host only when no worker is up, or after the wait. - The web host's own GPU runs at most one model at a time however many jobs are in flight (a lock of its own). This includes MedIA-Agentic, which runs on the web host's GPU directly and used to be covered only by the old global lock. Cancellation is checked after the lock is acquired as well as while waiting. Not taken when workers are disabled, so that path is unchanged. - A session never runs against itself: a repeat request for a session whose job is queued or running waits for it (as it did behind the old lock) without using a job slot, instead of overlapping it in the same workspace. - The older EPAI_REMOTE_ENABLED mode keeps jobs strictly serial. It is read with the same yes/no parser the ePAI runner uses, and there is now a single parser for every environment flag, so a spelling such as "y" cannot mean yes to one setting and no to another. - The SSH connection master per worker is created under a lock. - The queue cap (INFERENCE_MAX_PENDING) defaults to 3 + N (4 as before at 1). Verified with a stress test (8 jobs over 3 workers: no double booking), unit tests for waiting/cancel/limits and for every review finding (each fails on the code without its fix), and a real run of 3 ePAI jobs with N=2 on bdmap2 and bdmap4: jobs 1 and 2 ran side by side, job 3 took the first free worker; 92 s total against about 140 s one at a time.
aperson30
force-pushed
the
feat/parallel-gpu-jobs
branch
from
October 3, 2026 04:43
e8c1058 to
c959b37
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Follow-up to #270/#271. Today one global lock makes every job wait, even when other GPUs are idle. This lets several jobs run at once, one per GPU worker.
Off by default.
GPU_WORKER_PARALLELdefaults to 1, which is exactly today's behavior (one job at a time), so merging changes nothing until it is raised. It is capped at the number of workers and ignored (1) while workers are disabled orEPAI_REMOTE_ENABLEDis on. Rollback: unset it or set it to 1, and reload.What changes
auto_segmentor: the single_gpu_lockbecomes_job_slots, a semaphore withmax_parallel_jobs()slots (1 = the old lock).bdmap1's own GPU gets its own lock (_local_gpu_lock), so it still runs at most one model at a time however many jobs are in flight. It is only taken when workers are enabled, so the disabled path is unchanged. Waiting for it is cancellable.gpu_workers.acquire_host: reservation is now safe for concurrent jobs (see below), and a job that finds every usable worker busy with our own jobs waits (up toGPU_WORKER_MAX_WAIT_SECONDS, default 300, cancellable) instead of overflowing to bdmap1. bdmap1 is used only when no worker is up at all, or after that wait.INFERENCE_MAX_PENDINGdefaults to3 + N(4 at N=1, as before).A bug this fixes before it can bite
The merged
acquire_hostreserved every idle worker while it health-checked them. With one job at a time that was harmless, but with two simultaneous jobs and two idle workers the second one saw "no worker free" and would have run on bdmap1's GPU. Replaying that on the merged code:['None', 'w1']. Now workers are checked without being reserved and the winner is reserved atomically, so two jobs never get the same worker and never see a false "none free".Independent review (5 rounds, 8 findings, all fixed)
EPAI_REMOTE_ENABLEDsingle-host ePAI mode sends every ePAI job to one fixed GPU outside the pool: it now keeps jobs strictly serial. (It is not set on the live server.)EPAI_REMOTE_ENABLEDused a yes/no parser that did not accepty, which the ePAI runner does: soEPAI_REMOTE_ENABLED=yplus parallel jobs would have sent two ePAI jobs to the same fixed GPU. There is now one parser for every environment flag, so the two cannot disagree.Each has a regression test that fails on the code without its fix.
Testing
test_gpu_workers.py(42 new, stable over repeated runs): a stress test (8 jobs over 3 workers x 40 rounds: no worker ever held twice), two simultaneous jobs get two workers, waiting beats overflowing, the wait is bounded and cancellable, nothing-to-wait-for returns at once, an unhealthy free worker waits for the busy healthy one,max_parallel_jobslimits and bad values, slots bound concurrent models (1 = old behavior, 2 lets two in), bdmap1's GPU runs one model at a time even with parallel jobs, its wait is cancellable, and the disabled path never takes the lock.Assumptions (documented in
GPU_WORKERS.md)Each worker has one GPU (placement is per worker); one gunicorn process (reservations are in-process); no strict first-come-first-served order between waiting jobs; reload only when no job is running.
Rollout plan
GPU_WORKER_PARALLELunset (no change).GPU_WORKER_PARALLEL=2(bdmap2 + bdmap4; bdmap3 stays a spare) and reload.