Context
PR #58 / #60's supervised path is single-worker only — when `workers > 1` the run() function falls back to the unsupervised `ProcessPoolExecutor` path which has no hard timeout. For non-MPS workloads (embedding-only re-runs, future Docling parallelism, CPU-bound extractors) proper multi-worker supervision would unlock real parallelism.
Not urgent because MPS contention makes parallel marker workers strictly bad anyway (#64).
Proposal
Extend `_run_supervised_single_worker` into `_run_supervised_workers(n)`:
- Spawn N `_subprocess_worker` children, each with its own UUID-based `claimed_by`.
- Single supervisor loop polls `_find_stuck_extracting` per-worker and SIGKILLs only the offending child.
- Respawn killed children individually.
- Same process-group kill semantics for grandchild reaping.
`run()` dispatches to this when `workers > 1` AND `extract_timeout_s > 0`.
Acceptance criteria
Out of scope
This issue does NOT advocate using multiple MPS workers — see #64. The motivating use case is CPU-bound or non-MPS extractors.
🤖 Generated with Claude Code
Context
PR #58 / #60's supervised path is single-worker only — when `workers > 1` the run() function falls back to the unsupervised `ProcessPoolExecutor` path which has no hard timeout. For non-MPS workloads (embedding-only re-runs, future Docling parallelism, CPU-bound extractors) proper multi-worker supervision would unlock real parallelism.
Not urgent because MPS contention makes parallel marker workers strictly bad anyway (#64).
Proposal
Extend `_run_supervised_single_worker` into `_run_supervised_workers(n)`:
`run()` dispatches to this when `workers > 1` AND `extract_timeout_s > 0`.
Acceptance criteria
Out of scope
This issue does NOT advocate using multiple MPS workers — see #64. The motivating use case is CPU-bound or non-MPS extractors.
🤖 Generated with Claude Code