Repository navigation
Record which machine ran each model job - #273
Merged
Merged
Conversation
Nothing recorded which machine ran a job or whether it fell back to the web host; the only trace was a buffered log line lost on reboot. That made it impossible to answer "how often do workers fail?" or "is bdmap1's GPU staying clear?" from data. Each finished job (completed, failed or cancelled) now adds one line to <sessions dir>/job_runs.jsonl: model, outcome, the machine(s) it ran on, whether it fell back to the web host, duration, input size and a short error. It holds no user id or IP. The newest 5000 are kept. It is a separate append-only file, not a job.json field or database column, so nothing the site reads changes, and writing it is best-effort: a problem there can never affect a user's job. `python -m services.job_run_log <file> [days]` summarizes jobs by outcome, model and machine and the share that fell back. Bookkeeping in auto_segmentor notes where each model command ran (a worker, or this host, flagged when it was a fallback or a re-run after a worker failure); gpu_workers.run says which worker ran it; the job thread in api_blueprint logs every way a job can end and always drops the in-memory entry. Tested with unit tests (file format, trimming, summary, every routing outcome, and a structural check that every job exit is logged) and a real LesionSegmenter job on a worker: it recorded bdmap4, 30 s, no fallback.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Follow-up to #270/#271/#272. Nothing records which machine ran a job or whether it fell back to bdmap1; the only trace is a buffered log line that is lost on reboot. That makes "how often do workers fail?" and "is bdmap1's GPU staying clear?" impossible to answer from data. This adds a small, separate record.
What it adds
Each finished job (completed, failed or cancelled) appends one line to
<sessions dir>/job_runs.jsonl:model,status,ran_on(the machine(s)),fell_back(to the web host),duration_seconds,input_size_bytes, a shorterror,session_id,recorded_at. No user id and no IP address. The newest 5000 are kept.python -m services.job_run_log <file> [days]summarizes jobs by outcome, model and machine, and the share that fell back to bdmap1.Why it is safe
job.jsonfield or a database column, so nothing the site reads changes and there is no schema change.ran_onbookkeeping inauto_segmentoronly notes where each model command ran (a worker; this host; a fallback; a re-run after a worker failure). It never changes how a job runs.finally) and is capped, so it cannot leak.Changes
services/job_run_log.py(new, stdlib only): the file, a loader and the summary.auto_segmentor:_note_run/pop_run_info, called where a model command finishes (worker, local, fallback, out-of-memory), and for MedIA-Agentic.gpu_workers.runattaches the worker that ran the command.api_blueprint: the job thread logs every way a job can end and always releases the entry.deploy/GPU_WORKERS.mddocuments it; CI runs the new tests.Testing
api_blueprintcannot be imported without the app's dependencies in the unit-test environment). 107 pass with the existing GPU worker suite.ran_on: ["bdmap4"], 30 s, no fallback; the summary command read it back.Rollout
Deploys with the next reload. The file is created on the first finished job.