Skip to content

Record which machine ran each model job - #273

Merged
aperson30 merged 1 commit into
BodyMaps:mainfrom
aperson30:feat/record-job-machine
Oct 3, 2026
Merged

aperson30 merged 1 commit into
BodyMaps:mainfrom
aperson30:feat/record-job-machine

Conversation

@aperson30

Copy link
Copy Markdown
Collaborator

Follow-up to #270/#271/#272. Nothing records which machine ran a job or whether it fell back to bdmap1; the only trace is a buffered log line that is lost on reboot. That makes "how often do workers fail?" and "is bdmap1's GPU staying clear?" impossible to answer from data. This adds a small, separate record.

What it adds

Each finished job (completed, failed or cancelled) appends one line to <sessions dir>/job_runs.jsonl:
model, status, ran_on (the machine(s)), fell_back (to the web host), duration_seconds, input_size_bytes, a short error, session_id, recorded_at. No user id and no IP address. The newest 5000 are kept.

python -m services.job_run_log <file> [days] summarizes jobs by outcome, model and machine, and the share that fell back to bdmap1.

Why it is safe

  • It is a separate append-only file, not a job.json field or a database column, so nothing the site reads changes and there is no schema change.
  • Writing is best-effort and never raises: a problem there cannot affect a user's job.
  • ran_on bookkeeping in auto_segmentor only notes where each model command ran (a worker; this host; a fallback; a re-run after a worker failure). It never changes how a job runs.
  • The in-memory entry is always dropped (in the job thread's finally) and is capped, so it cannot leak.

Changes

  • services/job_run_log.py (new, stdlib only): the file, a loader and the summary.
  • auto_segmentor: _note_run / pop_run_info, called where a model command finishes (worker, local, fallback, out-of-memory), and for MedIA-Agentic.
  • gpu_workers.run attaches the worker that ran the command.
  • api_blueprint: the job thread logs every way a job can end and always releases the entry.
  • deploy/GPU_WORKERS.md documents it; CI runs the new tests.

Testing

  • 17 new unit tests (file format and no user data, trimming, loader and summary, every routing outcome: worker, plain local, fallback, re-run after a worker failure, out-of-memory, non-model commands) plus a structural test that every exit of the job thread is logged and the entry is always released (api_blueprint cannot be imported without the app's dependencies in the unit-test environment). 107 pass with the existing GPU worker suite.
  • Real run on the server (throwaway copy, real LesionSegmenter job): recorded ran_on: ["bdmap4"], 30 s, no fallback; the summary command read it back.

Rollout

Deploys with the next reload. The file is created on the first finished job.

Nothing recorded which machine ran a job or whether it fell back to the web
host; the only trace was a buffered log line lost on reboot. That made it
impossible to answer "how often do workers fail?" or "is bdmap1's GPU staying
clear?" from data.

Each finished job (completed, failed or cancelled) now adds one line to
<sessions dir>/job_runs.jsonl: model, outcome, the machine(s) it ran on,
whether it fell back to the web host, duration, input size and a short error.
It holds no user id or IP. The newest 5000 are kept.

It is a separate append-only file, not a job.json field or database column,
so nothing the site reads changes, and writing it is best-effort: a problem
there can never affect a user's job. `python -m services.job_run_log <file>
[days]` summarizes jobs by outcome, model and machine and the share that fell
back.

Bookkeeping in auto_segmentor notes where each model command ran (a worker, or
this host, flagged when it was a fallback or a re-run after a worker failure);
gpu_workers.run says which worker ran it; the job thread in api_blueprint logs
every way a job can end and always drops the in-memory entry.

Tested with unit tests (file format, trimming, summary, every routing outcome,
and a structural check that every job exit is logged) and a real LesionSegmenter
job on a worker: it recorded bdmap4, 30 s, no fallback.
@aperson30
aperson30 merged commit 3bf69f5 into BodyMaps:main Oct 3, 2026
8 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant