engine: shared schema gate and the write critical section (RFC 2026-09-18) - #783
Conversation
|
Following up on @azimafroozeh's review, which came in chat. 1. "The benchmark mixes copy-on-write and the schema gate." It doesn't mix in copy-on-write. The The underlying point was right, though: this PR bundles two effects of its own.
A diagnostic build (this branch plus one process-wide publication lock, so captures overlap but publications still serialize) separates them.
2. "With one writer we're bottlenecked by the writer: rejections, more IO, and IO latency dominates." IO latency dominates, but a single writer never rejects: 0 failed revalidations with one writer.
Rejections come from several writers on one branch: 6.6 per commit locally and 2.9 at +30 ms, each burning about 300 ms of revalidation IO under the gate. Rejecting stale attempts from memory before any IO was prototyped:
It also can't beat one writer, because every publication invalidates every other prepared write on the branch. So there are two levers:
|
The mechanism for RFC 2026-09-18-shared-schema-gate, with zero callers: SchemaGateSlot (tokio RwLock + release epoch), SchemaSharedPermit / SchemaExclusivePermit whose drops follow the QueueGuard protocol (turn first, release inside it, epoch bump last), scheduled shared/exclusive acquisition that keeps scheduled_lock's three DST invariants (the contender never leaves the turn loop, installed mode attempts with try_* only, yield every iteration), HeldWriteGates (field order = drop order, schema first, matching the former guards[0] release sequence), and the acquire_schema_shared/exclusive manager methods. Unit tests: shared permits concurrent; exclusive excludes shared in both directions; the plain-mode fairness pin (a shared permit requested after a queued exclusive parks behind it — reds if a tokio upgrade changes the RwLock policy); one gate per canonical root across handles.
The atomic half of RFC 2026-09-18-shared-schema-gate: all 22 acquisition
sites move to the permit API in one commit so no mixed-mechanism window
exists, and schema_apply_serial_queue_key() is deleted with its synthetic
("__schema_apply__", None) entry.
Exclusive (contract-lifecycle passes, unchanged mutual exclusion):
schema apply, the system-column upgrade, open, sync_branch, refresh,
settle_pending_schema_install, reload_schema_if_source_changed.
Shared (readers of the accepted view — they stop excluding each other
and keep excluding every contract-lifecycle pass): mutation/load
commit_all, the no-op conditional-mutation CAS, branch merge,
ensure_indices, optimize, cleanup, repair, orphaned-branch
reconciliation, branch create/create-from/delete,
ensure_no_pending_recovery, and the three read-view captures (reads no
longer serialize behind writers' publish holds).
The staged write returns HeldWriteGates (shared permit + branch + sorted
table guards; field order preserves the former release order, so seeded
DST grant sequences keep their event order) in place of the anonymous
guard Vec. The apply-side comments are re-derived to name the permit and
the shared-re-entry parking hazard.
Three regression owners for RFC 2026-09-18-shared-schema-gate: - cross_branch_writers_overlap_inside_schema_gate — the headline: a writer parked inside its full envelope on b1 no longer blocks a b2 writer, which runs to completion while A is parked (impossible under the exclusive mutex). - read_capture_proceeds_while_writer_parked — reads capture their catalog under a shared permit and stop waiting for a writer's publish hold. - parked_writer_blocks_schema_apply — the reverse direction and the mis-classification tripwire: a held shared permit blocks the apply's exclusive acquisition outright; a writer wrongly left off the gate would let the apply proceed mid-write and red this test. The existing mid-apply pin keeps its assertion; its comment now names the shared-behind-exclusive mechanism.
writes.md names the shared permit in the finalization order and the promotion-after-release step; architecture.md's read sentence now states what is true (a capture takes a shared schema permit, no branch or table gate, and does not wait for writers — before this change it took the exclusive gate). The RFC flips to accepted/complete and its decision log records the implementation resolutions (optimize/cleanup shared; RO-open/reload exclusive as recorded conservatism; the tripwire and fairness pins; the DST arm with strict replay).
The coverage RFC 2026-09-18-shared-schema-gate commissions: the concurrent universe gains a schema actor that applies monotone additive schemas against main while the data writers run. dst_schema_apply_racing_writers_first_contact asserts every data write and every apply commits (neither side starves), each apply lands exactly one era commit, at least one seed interleaves writer commits (a non-vacuous concurrency claim), and under the seam scheduler the permits' turn/epoch protocol keeps the schedule seed-ordered (sched_escapes == 0, strict replay). The ignored hunt sweeps the interleaving space.
The DST schema-apply-racing-writers arm found the write capture outside the gate: open_write_txn checked the durable sentinel before any permit, so a writer arriving on another handle while an apply was in flight was refused with a typed conflict and retried hot until the apply ended (256 immediate retries exhausted the arm's OCC budget on its first op), while a writer that had entered before the apply parked behind the exclusive side. The capture now probes the sentinel (schema_apply_sentinel_present, factored out of ensure_schema_apply_not_locked) and, when it stands, parks on the shared side until the apply releases, then recaptures under the promoted contract; a cross-process apply grants the permit at once and the bounded loop ends in the same typed refusal as before. The common path takes no permit: a capture-scoped shared permit was tried first and cost every reprepare attempt two arbiter turns under the DST seam scheduler (dst_seam_scheduler_bite_and_replay went from 0.8 s to 5.2 s, near its escape budget); with the park-on-refusal shape it runs in 0.6 s. Evidence: the arm's seeds commit every write and every apply; the mid-apply pin keeps its assertion (its mutation captures before the apply starts).
Review follow-up. The first-contact pin's non-vacuity check counted only writer alternations, so it could not show that an apply ever contended with the writers — and none had: the schema actor applied back to back the moment the start barrier opened, and the write-preferring lock ran every apply before the first writer commit. The concurrent report now counts era commits that land between two writer commits (era_commits_between_data), the pin requires one somewhere across its seeds, and the schema actor pauses a seeded 2-31 ms before each apply so the applies spread across the writers' lifetime. Every apply of every seed now lands between writer commits. The hunt runs without readers: a reader's read-only open takes the schema gate's exclusive side with no arbiter hook, so its transitions fell outside the turns sched_escapes == 0 certifies. The pin's dead scheduler branch, left from when its scheduler seed moved to the hunt, is removed.
Review follow-up. A read on the writer's own handle still waits while that handle publishes on its bound branch: the publish holds the coordinator lock across the manifest compare-and-swap and the read capture takes it. The architecture guide, the release note, the RFC's behavior section and a test comment now say that, and the remaining wait is listed under the RFC's unresolved questions. The RFC's decision log records the independent review and the step 2 measurement: committing detached effects before the gates is 13-15% of the gate hold and, because revalidation fails on any head move, would turn most same-branch attempts into dead detached commits, so it is not adopted. RFC 0067's plan text gains a dated amendment pointing there, and its step 3 no longer describes a promotion pass that detached-only tables removed.
The in-process stale-attempt check caught every failed revalidation but gained only at object-store latency (0.40 to 0.53 commits/s with eight writers on one branch at +30 ms) and lost locally (25.5 to 23.2), and it left the starvation of #784 unchanged. The measurement bounds the whole family: with branch-wide revalidation, eight writers on one branch cannot exceed one writer alone, so raising same-branch throughput needs group commit's footprint admission rather than more work on the gate.
…laim A diagnostic build with one process-wide publication lock shows the whole eight-branch gain comes from overlapping publications across branches; interleaved reruns show no one-branch difference, so the earlier +36% was drift between sessions.
35926d3 to
bc09eec
Compare
What & why
Implements RFC 2026-09-18: shared schema gate and the write critical section (#754) — step 2 of RFC 0067's throughput path. Based on
main(the RFC merged as #754). Ported onto main after #776 removed table promotion and #779 extractedomnigraph-core.The process-local schema gate becomes a shared/exclusive lock under one rule — readers of the accepted view share; publishers of the accepted view exclude:
sync_branch,refresh,settle_pending_schema_install,reload_schema_if_source_changed.commit_all, the no-op conditional-mutation CAS, branch merge,ensure_indices, optimize, cleanup, repair, orphaned-branch reconciliation, branch create/create-from/delete,ensure_no_pending_recovery, and the three read-view captures — so a writer on one branch no longer waits for a writer on another, and reads stop serializing behind writers' publish holds.merge_exclusiveis already gone frommain, so a merge now takes the shared permit and only its own source and target branch gates: merges into independent targets no longer queue on a graph-wide guard. performance: merges into independent targets serialize behind graph-wide guards #643 also asks for concurrent-merge measurements and boundary tests that this PR does not carry, so it stays open.One wait remains on the read path. A read on the writer's own handle still waits for that handle's coordinator while a publish on its bound branch is in flight, because
commit_updates_on_branch_with_expectedholdscoordinator.write()across the publish. The RFC lists this under Unresolved questions.Why the old rationale no longer binds: the exclusive gate guarded a mutation advancing a Lance HEAD before discovering an in-flight schema lock. Under RFC 0067 no pre-publication effect moves a HEAD, so a misclassified site can at worst attempt a stale publish, which the CAS refuses.
Commits
engine: shared/exclusive schema-gate slot on the write-queue manager— the mechanism with zero callers:SchemaGateSlot(tokioRwLock+ release epoch),SchemaSharedPermit/SchemaExclusivePermitfollowing theQueueGuarddrop protocol, scheduled acquisition keepingscheduled_lock's three DST invariants,HeldWriteGates. Unit tests incl. the plain-mode fairness pin.engine: flip the schema gate to shared/exclusive at every site— all 22 sites in one commit (no mixed-mechanism window);schema_apply_serial_queue_key()deleted; the staged write returns a typedHeldWriteGatesenvelope instead of an anonymous guardVec.test: pin the shared schema gate's behavior in both directions—cross_branch_writers_overlap_inside_schema_gate(headline),read_capture_proceeds_while_writer_parked,parked_writer_blocks_schema_apply(the mis-classification tripwire).docs: gate order, read-path claim, and the RFC's implementation record—writes.mdfinalization order,architecture.mdread sentence corrected, RFC → accepted/complete with its decision log.docs(release): writers no longer serialize on the schema gate.test(dst): concurrent-universe arm racing writers against a schema apply— the DST coverage the RFC commissions: a schema actor applies monotone additive schemas against main while data writers run; asserts neither side starves, one era commit per apply, and non-vacuous interleaving on plain-mode seeds. Strict replay under the seam scheduler is the ignored hunt's claim for this arm (an apply's table rewrite runs on the singlelance-cpupool thread the arbiter deliberately cannot see, so its stall budget trips under load); the permits' turn/epoch protocol stays pinned bydst_seam_scheduler_bite_and_replay.engine: the write capture parks behind an in-flight schema apply— what the DST arm found on its first run:open_write_txnchecked the durable sentinel before any permit, so a writer arriving on another handle mid-apply was refused with a typed conflict and retried hot (256 immediate retries on its first op), while a writer that entered before the apply parked. Now, when the sentinel probe is positive, the capture parks on the shared side until the apply releases and recaptures under the promoted contract; the common path takes no permit (a capture-scoped permit was tried first and cost the seam-scheduler pin 0.8 s → 5.2 s).test(dst): the schema arm lands applies between writer commits— review found the arm's non-vacuity check could pass with no apply ever landing between two writer commits, and none did: the schema actor applied back to back as the start barrier opened, and the write-preferring lock ran every apply first. The actor now pauses a seeded 2–31 ms before each apply, and the pin requires an apply between writer commits (ConcurrentReport.era_commits_between_data). The hunt runs without reader actors, whose read-only opens take the exclusive side outside the turnssched_escapes == 0certifies.docs: narrow the read-path claim; record the step 2 measurement— the same-handle read wait above. Also: committing detached effects before the gates, the other half of RFC 0067's step 2, was measured and not adopted.docs(rfc): record the fail-fast prototype and the same-branch ceiling.docs(rfc): attribute the shared gate's gain; retract the one-branch claim— see Throughput, one effect at a time below.What review and the DST arm caught
The DST arm, before it went green:
dst_seam_scheduler_bite_and_replayescaped on every op, bisected to that commit alone. Reverted — and since superseded, as refactor: remove table commit promotion #776 removed promotion.std::sync::Barrieris cyclic, so the first N actors passed and the last parked forever — what held the DST suite for four hours), and its thread was not attributed to the seam scheduler.An independent review (Codex,
gpt-6-astra) found three gaps, each verified and fixed in commits 8 and 9:Review of the benchmark asked for one effect at a time. The split, below, retracted the one-branch throughput claim.
Evidence (the RFC's acceptance bar)
cross_handle_branch_gate_serializes_post_effect_publishandoptimize_holds_main_gate_through_disjoint_table_effects(failpoints.rs), and the 16-handlen_concurrent_strict_loads_have_one_same_id_winner_and_keep_disjoint_ids(consistency.rs).mutation_waits_for_mid_apply_schema_gate_then_reprepares(schema_apply.rs) keeps its assertion, with the mechanism now reader-behind-writer.write_cost.rsper-operation counts unchanged (manifest ceiling, schema-fence read/exists counts).concurrent-writesthroughput, below.Throughput, one effect at a time
This change has two effects:
To separate them, a diagnostic build added one process-wide lock around every mutation and load publication to this branch. Captures overlap publications, but publications still serialize. It is never merged.
main, the diagnostic and this PR ran back to back in each cell, 3 runs per cell, medians:--writers 8, 30 s windows;--rows 256;--rows 64.main's side isb14c22c5, which is the merge of #781 (copy-on-write catalog publication), so both sides include copy-on-write. The records' build attestation shows clean trees atb14c22c5and at this branch.main¹
mainand this PR from the earlier session, the diagnostic from this one; this cell reads 0.40 for every build in every session.What the table shows:
mainlocally and 4.3× at +30 ms. The diagnostic, which overlaps captures but not publications, matchesmain.authority_conflictsis 0 in every 8-branch cell.read_capture_proceeds_while_writer_parkedpins, not a throughput claim.Compare within one session only. The local one-branch cell drifts by about 20% between sessions, because re-prepares dominate it: every build exhausts 116–178 re-prepare budgets per cell. The eight-branch cells drift too; this PR's local cell read 177 in an earlier session and 154 here. That earlier session compared
mainand this PR in separate runs: 2.5× local and 5.9× at +30 ms for 8 branches, and +36% locally for 1 branch. The one-branch gain was that drift, and it is retracted.Same-branch starvation on
main(#784)At +30 ms on one branch, with 12 commits per run, one writer on
mainmakes all 12 commits while the other seven make none for the whole window. The writer that just published holds the freshest view and wins every revalidation; the others re-prepare and lose again. This PR spreads the same 12 commits over two or three writers but does not make admission fair.The same-branch ceiling was measured in the RFC's decision log. Because any publication invalidates every other prepared write on the branch, eight writers on one branch cannot exceed one writer (31.7 commits/s locally, 0.87 at +30 ms). Two changes were measured and not adopted:
Raising that ceiling is the group commit of RFC 0067's step 3, proposed in #785.
Validation
Local, on the tree at commit 7:
cargo fmt --check; write-queue unit tests (16); engine ownersschema_apply(30),writes(46),consistency(23),branching(48),maintenance(47),lifecycle(27),composite_flow(3),write_cost(17, per-operation counts unchanged);failpoint_names_guard(8);failpoints(118); the full DST suite fromcrates/omnigraph-dst(54); both workspace clippy graphs with-D warnings;check-docs,typos,check-agents-md,check-workflow-action-pins.Later commits: the DST suite (54) and lib tests after commit 8, and DST clippy. Commits 9–11 are docs only (
check-docs,typos).