Skip to content

feat: concurrency pool + fail-fast cancellation (F6a) - #20

Merged
halaprix merged 2 commits into
mainfrom
feat/f6-concurrency-pool
Jul 23, 2026
Merged

halaprix merged 2 commits into
mainfrom
feat/f6-concurrency-pool

Conversation

@halaprix

Copy link
Copy Markdown
Owner

Summary

Implements the first half of spec F6 (1.2.0 execution controls — all opt-in):

  • maxConcurrentBatches (default 1): concurrency-limited pool per step (src/core/pool.ts); steps remain hard barriers; routing is index-based and completion-order-independent (fuzz-tested with random delays). Default serial path is bit-identical to 1.1 — including error-object rethrow identity and read-at-settlement result semantics (each batch's results are snapshotted at settlement so buffer-reusing executors stay legal).
  • Fail-fast cancellation for run per spec (a)–(d): queued batches never dispatched after a failure; in-flight batches settle inside their workers' try/catch (structurally no unhandled rejections — a global vitest unhandledRejection guard now enforces this suite-wide); thrown error selected by lowest (batchIndex, callIndex) among discovered terminals (deterministic for exactly one failing batch — pinned over 100 shuffled runs; multi-failure identity explicitly unasserted); in-flight results discarded. runSettled never cancels — its policy resolves transport failures as per-call kind: 'batch' errors.
  • Runner unification (ledger debt from F5): one shared step engine (src/core/engine.ts) with a 3-hook per-runner policy capturing the run/runSettled divergence; the duplicated FSM loops are gone.
  • Validation: batchSize/maxConcurrentBatches/maxBatchAttempts safe integers ≥1, pre-consumption; maxBatchAttempts defaults to 2·⌈log₂(batchSize)⌉+1, stored for bisection (next PR). adaptiveBatching/dedupe/pinBlock deliberately NOT in the type yet — each lands with its implementing feature.
  • Review hardening: result snapshotting (above), async rejection-guard timing, concurrency tests assert observed max-in-flight + relative wall-clock (self-calibrating, no absolute CI-hostile deadlines).

Test evidence

227/227 (18 new concurrency tests), compat-vs-dist 38/38, snippets clean, full gate build-first. Bundle 11.5KB gzip (15KB budget).

Spec: spec.md § 1.2.0 F6 (pool + cancellation; bisection follows).

halaprix added 2 commits July 23, 2026 18:44
Unify runMultistepTasks/runSettled's duplicated per-step loop (F5 ledger
debt) into one shared engine (src/core/engine.ts) driven by a concurrency-
limited batch pool (src/core/pool.ts). Adds maxConcurrentBatches (default
1, genuinely sequential and bit-identical to 1.1) and maxBatchAttempts
(validated + default-computed, not yet consumed — bisection lands later)
to BatchOptions.

run() gains fail-fast cancellation: on the first irrecoverable batch
rejection, queued batches are not dispatched, in-flight batches settle
without leaking, and the run rejects with the lowest-index discovered
terminal error. runSettled() never cancels — its executeBatch policy
converts transport rejections into recorded kind:'batch' failures instead
of ever rejecting into the pool.

Adds a global unhandledRejection test guard (vitest setupFiles, both
configs) and src/__tests__/concurrency.test.ts covering wall-clock
parallelism, completion-order-independent routing (fuzzed), fail-fast
(a)-(d), error-identity determinism over 100 runs, runSettled's never-
cancel behavior, numeric validation, and the step barrier under
concurrency.
…ng, use observed concurrency in timing tests

External review (3 findings), same branch:

- pool.ts: clone each batch's result array/wrappers at the moment its own
  await resolves (results[i] = batchResults.map(r => ({...r}))) instead of
  retaining the raw executor-owned objects. Pre-F6a sequential code read
  and converted results immediately after each await; deferring routing
  until the whole step's pool settles let a later batch's mutation of a
  reused array/wrapper corrupt an earlier, already-settled batch's entry
  (reproducible at maxConcurrentBatches: 1 too). New test proves a
  mutating-buffer executor routes identically to a well-behaved one at
  conc=1 and conc=2.

- unhandled-rejection guard: afterEach now awaits one macrotask
  (setImmediate) before asserting/clearing, since Node's unhandledRejection
  emission can land on a later tick than the offending test's own body;
  afterAll deregisters the listener so re-evaluation can't stack them.

- concurrency.test.ts: replaced the two absolute wall-clock-bound tests
  with instrumented max-in-flight/dispatch-order assertions (deterministic,
  no timing assumption) plus a single self-calibrated relative wall-clock
  check (parallel < 0.6x serial for the same fixture, in the same test) —
  removes the absolute-deadline flake risk on loaded CI.
@halaprix
halaprix merged commit 48ce4b7 into main Jul 23, 2026
2 checks passed
@halaprix
halaprix deleted the feat/f6-concurrency-pool branch July 23, 2026 16:59
halaprix added a commit that referenced this pull request Jul 24, 2026
* feat: concurrency pool + fail-fast cancellation (F6a)

Unify runMultistepTasks/runSettled's duplicated per-step loop (F5 ledger
debt) into one shared engine (src/core/engine.ts) driven by a concurrency-
limited batch pool (src/core/pool.ts). Adds maxConcurrentBatches (default
1, genuinely sequential and bit-identical to 1.1) and maxBatchAttempts
(validated + default-computed, not yet consumed — bisection lands later)
to BatchOptions.

run() gains fail-fast cancellation: on the first irrecoverable batch
rejection, queued batches are not dispatched, in-flight batches settle
without leaking, and the run rejects with the lowest-index discovered
terminal error. runSettled() never cancels — its executeBatch policy
converts transport rejections into recorded kind:'batch' failures instead
of ever rejecting into the pool.

Adds a global unhandledRejection test guard (vitest setupFiles, both
configs) and src/__tests__/concurrency.test.ts covering wall-clock
parallelism, completion-order-independent routing (fuzzed), fail-fast
(a)-(d), error-identity determinism over 100 runs, runSettled's never-
cancel behavior, numeric validation, and the step barrier under
concurrency.

* fix: snapshot pool results at settlement, harden rejection-guard timing, use observed concurrency in timing tests

External review (3 findings), same branch:

- pool.ts: clone each batch's result array/wrappers at the moment its own
  await resolves (results[i] = batchResults.map(r => ({...r}))) instead of
  retaining the raw executor-owned objects. Pre-F6a sequential code read
  and converted results immediately after each await; deferring routing
  until the whole step's pool settles let a later batch's mutation of a
  reused array/wrapper corrupt an earlier, already-settled batch's entry
  (reproducible at maxConcurrentBatches: 1 too). New test proves a
  mutating-buffer executor routes identically to a well-behaved one at
  conc=1 and conc=2.

- unhandled-rejection guard: afterEach now awaits one macrotask
  (setImmediate) before asserting/clearing, since Node's unhandledRejection
  emission can land on a later tick than the offending test's own body;
  afterAll deregisters the listener so re-evaluation can't stack them.

- concurrency.test.ts: replaced the two absolute wall-clock-bound tests
  with instrumented max-in-flight/dispatch-order assertions (deterministic,
  no timing assumption) plus a single self-calibrated relative wall-clock
  check (parallel < 0.6x serial for the same fixture, in the same test) —
  removes the absolute-deadline flake risk on loaded CI.

---------

Co-authored-by: halaprix <halaprix@users.noreply.github.com>
halaprix added a commit that referenced this pull request Jul 24, 2026
* feat: concurrency pool + fail-fast cancellation (F6a)

Unify runMultistepTasks/runSettled's duplicated per-step loop (F5 ledger
debt) into one shared engine (src/core/engine.ts) driven by a concurrency-
limited batch pool (src/core/pool.ts). Adds maxConcurrentBatches (default
1, genuinely sequential and bit-identical to 1.1) and maxBatchAttempts
(validated + default-computed, not yet consumed — bisection lands later)
to BatchOptions.

run() gains fail-fast cancellation: on the first irrecoverable batch
rejection, queued batches are not dispatched, in-flight batches settle
without leaking, and the run rejects with the lowest-index discovered
terminal error. runSettled() never cancels — its executeBatch policy
converts transport rejections into recorded kind:'batch' failures instead
of ever rejecting into the pool.

Adds a global unhandledRejection test guard (vitest setupFiles, both
configs) and src/__tests__/concurrency.test.ts covering wall-clock
parallelism, completion-order-independent routing (fuzzed), fail-fast
(a)-(d), error-identity determinism over 100 runs, runSettled's never-
cancel behavior, numeric validation, and the step barrier under
concurrency.

* fix: snapshot pool results at settlement, harden rejection-guard timing, use observed concurrency in timing tests

External review (3 findings), same branch:

- pool.ts: clone each batch's result array/wrappers at the moment its own
  await resolves (results[i] = batchResults.map(r => ({...r}))) instead of
  retaining the raw executor-owned objects. Pre-F6a sequential code read
  and converted results immediately after each await; deferring routing
  until the whole step's pool settles let a later batch's mutation of a
  reused array/wrapper corrupt an earlier, already-settled batch's entry
  (reproducible at maxConcurrentBatches: 1 too). New test proves a
  mutating-buffer executor routes identically to a well-behaved one at
  conc=1 and conc=2.

- unhandled-rejection guard: afterEach now awaits one macrotask
  (setImmediate) before asserting/clearing, since Node's unhandledRejection
  emission can land on a later tick than the offending test's own body;
  afterAll deregisters the listener so re-evaluation can't stack them.

- concurrency.test.ts: replaced the two absolute wall-clock-bound tests
  with instrumented max-in-flight/dispatch-order assertions (deterministic,
  no timing assumption) plus a single self-calibrated relative wall-clock
  check (parallel < 0.6x serial for the same fixture, in the same test) —
  removes the absolute-deadline flake risk on loaded CI.

---------

Co-authored-by: halaprix <halaprix@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant