Skip to content

engine: shared schema gate and the write critical section (RFC 2026-09-18) - #783

Merged
azimafroozeh merged 14 commits into
mainfrom
claude/impl-shared-schema-gate
Sep 28, 2026
Merged

azimafroozeh merged 14 commits into
mainfrom
claude/impl-shared-schema-gate

Conversation

@ragnorc

@ragnorc ragnorc commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

What & why

Implements RFC 2026-09-18: shared schema gate and the write critical section (#754) — step 2 of RFC 0067's throughput path. Based on main (the RFC merged as #754). Ported onto main after #776 removed table promotion and #779 extracted omnigraph-core.

The process-local schema gate becomes a shared/exclusive lock under one rule — readers of the accepted view share; publishers of the accepted view exclude:

  • Exclusive (unchanged mutual exclusion): schema apply, the system-column upgrade, open, sync_branch, refresh, settle_pending_schema_install, reload_schema_if_source_changed.
  • Shared: mutation/load commit_all, the no-op conditional-mutation CAS, branch merge, ensure_indices, optimize, cleanup, repair, orphaned-branch reconciliation, branch create/create-from/delete, ensure_no_pending_recovery, and the three read-view captures — so a writer on one branch no longer waits for a writer on another, and reads stop serializing behind writers' publish holds.
  • The branch gate stays for same-branch writers; per-table gates are unchanged.
  • The RFC's promotion sub-decision is retired. It was implemented, measured harmful (below), reverted — and then detached-only tables removed promotion altogether, so there is nothing to move.
  • Nothing cross-process changes: the durable sentinel, the mono-branch refusal and the manifest CAS remain the authority.
  • Refs performance: merges into independent targets serialize behind graph-wide guards #643. merge_exclusive is already gone from main, so a merge now takes the shared permit and only its own source and target branch gates: merges into independent targets no longer queue on a graph-wide guard. performance: merges into independent targets serialize behind graph-wide guards #643 also asks for concurrent-merge measurements and boundary tests that this PR does not carry, so it stays open.

One wait remains on the read path. A read on the writer's own handle still waits for that handle's coordinator while a publish on its bound branch is in flight, because commit_updates_on_branch_with_expected holds coordinator.write() across the publish. The RFC lists this under Unresolved questions.

Why the old rationale no longer binds: the exclusive gate guarded a mutation advancing a Lance HEAD before discovering an in-flight schema lock. Under RFC 0067 no pre-publication effect moves a HEAD, so a misclassified site can at worst attempt a stale publish, which the CAS refuses.

Commits

  1. engine: shared/exclusive schema-gate slot on the write-queue manager — the mechanism with zero callers: SchemaGateSlot (tokio RwLock + release epoch), SchemaSharedPermit/SchemaExclusivePermit following the QueueGuard drop protocol, scheduled acquisition keeping scheduled_lock's three DST invariants, HeldWriteGates. Unit tests incl. the plain-mode fairness pin.
  2. engine: flip the schema gate to shared/exclusive at every site — all 22 sites in one commit (no mixed-mechanism window); schema_apply_serial_queue_key() deleted; the staged write returns a typed HeldWriteGates envelope instead of an anonymous guard Vec.
  3. test: pin the shared schema gate's behavior in both directions — cross_branch_writers_overlap_inside_schema_gate (headline), read_capture_proceeds_while_writer_parked, parked_writer_blocks_schema_apply (the mis-classification tripwire).
  4. docs: gate order, read-path claim, and the RFC's implementation record — writes.md finalization order, architecture.md read sentence corrected, RFC → accepted/complete with its decision log.
  5. docs(release): writers no longer serialize on the schema gate.
  6. test(dst): concurrent-universe arm racing writers against a schema apply — the DST coverage the RFC commissions: a schema actor applies monotone additive schemas against main while data writers run; asserts neither side starves, one era commit per apply, and non-vacuous interleaving on plain-mode seeds. Strict replay under the seam scheduler is the ignored hunt's claim for this arm (an apply's table rewrite runs on the single lance-cpu pool thread the arbiter deliberately cannot see, so its stall budget trips under load); the permits' turn/epoch protocol stays pinned by dst_seam_scheduler_bite_and_replay.
  7. engine: the write capture parks behind an in-flight schema apply — what the DST arm found on its first run: open_write_txn checked the durable sentinel before any permit, so a writer arriving on another handle mid-apply was refused with a typed conflict and retried hot (256 immediate retries on its first op), while a writer that entered before the apply parked. Now, when the sentinel probe is positive, the capture parks on the shared side until the apply releases and recaptures under the promoted contract; the common path takes no permit (a capture-scoped permit was tried first and cost the seam-scheduler pin 0.8 s → 5.2 s).
  8. test(dst): the schema arm lands applies between writer commits — review found the arm's non-vacuity check could pass with no apply ever landing between two writer commits, and none did: the schema actor applied back to back as the start barrier opened, and the write-preferring lock ran every apply first. The actor now pauses a seeded 2–31 ms before each apply, and the pin requires an apply between writer commits (ConcurrentReport.era_commits_between_data). The hunt runs without reader actors, whose read-only opens take the exclusive side outside the turns sched_escapes == 0 certifies.
  9. docs: narrow the read-path claim; record the step 2 measurement — the same-handle read wait above. Also: committing detached effects before the gates, the other half of RFC 0067's step 2, was measured and not adopted.
  10. docs(rfc): record the fail-fast prototype and the same-branch ceiling.
  11. docs(rfc): attribute the shared gate's gain; retract the one-branch claim — see Throughput, one effect at a time below.

What review and the DST arm caught

The DST arm, before it went green:

  • Engine: the entry-order inconsistency above.
  • Engine: promotion after guard release makes the next same-branch writer promote the same pin concurrently, so two replays serialize on Lance's commit path — an engine-internal wait the seam arbiter cannot see; dst_seam_scheduler_bite_and_replay escaped on every op, bisected to that commit alone. Reverted — and since superseded, as refactor: remove table commit promotion #776 removed promotion.
  • Harness: the schema actor was spawned but not counted among the start-barrier participants (std::sync::Barrier is cyclic, so the first N actors passed and the last parked forever — what held the DST suite for four hours), and its thread was not attributed to the seam scheduler.

An independent review (Codex, gpt-6-astra) found three gaps, each verified and fixed in commits 8 and 9:

  • the arm's vacuous interleaving check;
  • the hunt's reader actors, which sat outside the seam scheduler's turns;
  • an overstated read-path claim.

Review of the benchmark asked for one effect at a time. The split, below, retracted the one-branch throughput claim.

Evidence (the RFC's acceptance bar)

  • Existing pins survive as-is: cross_handle_branch_gate_serializes_post_effect_publish and optimize_holds_main_gate_through_disjoint_table_effects (failpoints.rs), and the 16-handle n_concurrent_strict_loads_have_one_same_id_winner_and_keep_disjoint_ids (consistency.rs).
  • Re-derived: mutation_waits_for_mid_apply_schema_gate_then_reprepares (schema_apply.rs) keeps its assertion, with the mechanism now reader-behind-writer.
  • write_cost.rs per-operation counts unchanged (manifest ceiling, schema-fence read/exists counts).
  • concurrent-writes throughput, below.

Throughput, one effect at a time

This change has two effects:

  • read-view captures stop waiting behind another writer's publication;
  • publications on different branches overlap.

To separate them, a diagnostic build added one process-wide lock around every mutation and load publication to this branch. Captures overlap publications, but publications still serialize. It is never merged.

main, the diagnostic and this PR ran back to back in each cell, 3 runs per cell, medians:

  • --writers 8, 30 s windows;
  • local FS with --rows 256;
  • RustFS through toxiproxy at +30 ms per round trip with --rows 64.

main's side is b14c22c5, which is the merge of #781 (copy-on-write catalog publication), so both sides include copy-on-write. The records' build attestation shows clean trees at b14c22c5 and at this branch.

Cell, commits/s main Diagnostic This PR
local, 8 branches 69.4 (p95 160 ms) 68.6 (p95 161 ms) 154.0 (p95 70 ms)
+30 ms, 8 branches 0.50 (p50 15.4 s) 0.43 (p50 17.4 s) 2.13 (p50 3.6 s)
local, 1 branch, round 1 24.4 21.7 21.7
local, 1 branch, round 2 20.3 22.2 22.8
+30 ms, 1 branch¹ 0.40 0.40 0.40

¹ main and this PR from the earlier session, the diagnostic from this one; this cell reads 0.40 for every build in every session.

What the table shows:

  • The whole gain is publications on different branches overlapping. At 8 branches this PR is 2.2× main locally and 4.3× at +30 ms. The diagnostic, which overlaps captures but not publications, matches main.
  • Against the RFC's bar, this PR's 8-branch rate is about 7× (local) and 5.3× (+30 ms) its own single-branch rate. authority_conflicts is 0 in every 8-branch cell.
  • One branch is unchanged. Same-branch publications still serialize on the branch gate, as intended. The capture overlap is a mechanism read_capture_proceeds_while_writer_parked pins, not a throughput claim.

Compare within one session only. The local one-branch cell drifts by about 20% between sessions, because re-prepares dominate it: every build exhausts 116–178 re-prepare budgets per cell. The eight-branch cells drift too; this PR's local cell read 177 in an earlier session and 154 here. That earlier session compared main and this PR in separate runs: 2.5× local and 5.9× at +30 ms for 8 branches, and +36% locally for 1 branch. The one-branch gain was that drift, and it is retracted.

Same-branch starvation on main (#784)

At +30 ms on one branch, with 12 commits per run, one writer on main makes all 12 commits while the other seven make none for the whole window. The writer that just published holds the freshest view and wins every revalidation; the others re-prepare and lose again. This PR spreads the same 12 commits over two or three writers but does not make admission fair.

The same-branch ceiling was measured in the RFC's decision log. Because any publication invalidates every other prepared write on the branch, eight writers on one branch cannot exceed one writer (31.7 commits/s locally, 0.87 at +30 ms). Two changes were measured and not adopted:

  • committing detached effects before the gates;
  • failing stale attempts before revalidation's round trips.

Raising that ceiling is the group commit of RFC 0067's step 3, proposed in #785.

Validation

Local, on the tree at commit 7: cargo fmt --check; write-queue unit tests (16); engine owners schema_apply (30), writes (46), consistency (23), branching (48), maintenance (47), lifecycle (27), composite_flow (3), write_cost (17, per-operation counts unchanged); failpoint_names_guard (8); failpoints (118); the full DST suite from crates/omnigraph-dst (54); both workspace clippy graphs with -D warnings; check-docs, typos, check-agents-md, check-workflow-action-pins.
Later commits: the DST suite (54) and lib tests after commit 8, and DST clippy. Commits 9–11 are docs only (check-docs, typos).

@ragnorc

ragnorc commented Sep 28, 2026

Copy link
Copy Markdown
Contributor Author

Following up on @azimafroozeh's review, which came in chat.

1. "The benchmark mixes copy-on-write and the schema gate." It doesn't mix in copy-on-write. The main side is b14c22c5, which is #781's own merge, so both binaries include copy-on-write, and the build attestation shows clean trees on both sides.

The underlying point was right, though: this PR bundles two effects of its own.

  • Read-view captures stop waiting behind a publication.
  • Publications on different branches overlap.

A diagnostic build (this branch plus one process-wide publication lock, so captures overlap but publications still serialize) separates them. main, the diagnostic and this PR ran back to back, in commits/s:

Cell main Diagnostic This PR
local, 8 branches 69.4 68.6 154.0
+30 ms, 8 branches 0.50 0.43 2.13
local, 1 branch, two rounds 24.4, 20.3 21.7, 22.2 21.7, 22.8
  • The whole gain comes from publications on different branches overlapping.
  • Overlapping captures adds nothing measurable.
  • The +36% one-branch figure I reported earlier compared two sessions, and that cell drifts by about 20% between sessions. It's retracted in the description and in the RFC's decision log (35926d3).

2. "With one writer we're bottlenecked by the writer: rejections, more IO, and IO latency dominates." IO latency dominates, but a single writer never rejects: 0 failed revalidations with one writer.

  • At +30 ms, one commit is about 38 serial round trips.
  • About 670 ms of that is under the branch gate: revalidation ~300 ms, detached commit ~85 ms, publication ~285 ms.
  • The remaining ~480 ms is preparation.

Rejections come from several writers on one branch: 6.6 per commit locally and 2.9 at +30 ms, each burning about 300 ms of revalidation IO under the gate. Rejecting stale attempts from memory before any IO was prototyped:

It also can't beat one writer, because every publication invalidates every other prepared write on the branch.

So there are two levers:

  • One writer: fewer serial round trips. Revalidation alone is about 45% of the gate hold at +30 ms, and the manifest fold is RFC 0068's subject.
  • Many writers on one branch: writes that don't invalidate each other. That's the group commit RFC in docs(rfc): group commit #785.

@azimafroozeh azimafroozeh self-assigned this Sep 28, 2026
ragnorc and others added 12 commits September 28, 2026 23:14
The mechanism for RFC 2026-09-18-shared-schema-gate, with zero callers:
SchemaGateSlot (tokio RwLock + release epoch), SchemaSharedPermit /
SchemaExclusivePermit whose drops follow the QueueGuard protocol (turn
first, release inside it, epoch bump last), scheduled shared/exclusive
acquisition that keeps scheduled_lock's three DST invariants (the
contender never leaves the turn loop, installed mode attempts with try_*
only, yield every iteration), HeldWriteGates (field order = drop order,
schema first, matching the former guards[0] release sequence), and the
acquire_schema_shared/exclusive manager methods.

Unit tests: shared permits concurrent; exclusive excludes shared in both
directions; the plain-mode fairness pin (a shared permit requested after
a queued exclusive parks behind it — reds if a tokio upgrade changes the
RwLock policy); one gate per canonical root across handles.
The atomic half of RFC 2026-09-18-shared-schema-gate: all 22 acquisition
sites move to the permit API in one commit so no mixed-mechanism window
exists, and schema_apply_serial_queue_key() is deleted with its synthetic
("__schema_apply__", None) entry.

Exclusive (contract-lifecycle passes, unchanged mutual exclusion):
schema apply, the system-column upgrade, open, sync_branch, refresh,
settle_pending_schema_install, reload_schema_if_source_changed.

Shared (readers of the accepted view — they stop excluding each other
and keep excluding every contract-lifecycle pass): mutation/load
commit_all, the no-op conditional-mutation CAS, branch merge,
ensure_indices, optimize, cleanup, repair, orphaned-branch
reconciliation, branch create/create-from/delete,
ensure_no_pending_recovery, and the three read-view captures (reads no
longer serialize behind writers' publish holds).

The staged write returns HeldWriteGates (shared permit + branch + sorted
table guards; field order preserves the former release order, so seeded
DST grant sequences keep their event order) in place of the anonymous
guard Vec. The apply-side comments are re-derived to name the permit and
the shared-re-entry parking hazard.
Three regression owners for RFC 2026-09-18-shared-schema-gate:

- cross_branch_writers_overlap_inside_schema_gate — the headline: a
  writer parked inside its full envelope on b1 no longer blocks a b2
  writer, which runs to completion while A is parked (impossible under
  the exclusive mutex).
- read_capture_proceeds_while_writer_parked — reads capture their
  catalog under a shared permit and stop waiting for a writer's publish
  hold.
- parked_writer_blocks_schema_apply — the reverse direction and the
  mis-classification tripwire: a held shared permit blocks the apply's
  exclusive acquisition outright; a writer wrongly left off the gate
  would let the apply proceed mid-write and red this test.

The existing mid-apply pin keeps its assertion; its comment now names
the shared-behind-exclusive mechanism.
writes.md names the shared permit in the finalization order and the
promotion-after-release step; architecture.md's read sentence now states
what is true (a capture takes a shared schema permit, no branch or table
gate, and does not wait for writers — before this change it took the
exclusive gate). The RFC flips to accepted/complete and its decision log
records the implementation resolutions (optimize/cleanup shared;
RO-open/reload exclusive as recorded conservatism; the tripwire and
fairness pins; the DST arm with strict replay).
The coverage RFC 2026-09-18-shared-schema-gate commissions: the concurrent
universe gains a schema actor that applies monotone additive schemas
against main while the data writers run. dst_schema_apply_racing_writers_first_contact
asserts every data write and every apply commits (neither side starves),
each apply lands exactly one era commit, at least one seed interleaves
writer commits (a non-vacuous concurrency claim), and under the seam
scheduler the permits' turn/epoch protocol keeps the schedule seed-ordered
(sched_escapes == 0, strict replay). The ignored hunt sweeps the
interleaving space.
The DST schema-apply-racing-writers arm found the write capture outside
the gate: open_write_txn checked the durable sentinel before any permit,
so a writer arriving on another handle while an apply was in flight was
refused with a typed conflict and retried hot until the apply ended (256
immediate retries exhausted the arm's OCC budget on its first op), while
a writer that had entered before the apply parked behind the exclusive
side.

The capture now probes the sentinel (schema_apply_sentinel_present,
factored out of ensure_schema_apply_not_locked) and, when it stands,
parks on the shared side until the apply releases, then recaptures under
the promoted contract; a cross-process apply grants the permit at once
and the bounded loop ends in the same typed refusal as before. The
common path takes no permit: a capture-scoped shared permit was tried
first and cost every reprepare attempt two arbiter turns under the DST
seam scheduler (dst_seam_scheduler_bite_and_replay went from 0.8 s to
5.2 s, near its escape budget); with the park-on-refusal shape it runs
in 0.6 s.

Evidence: the arm's seeds commit every write and every apply; the
mid-apply pin keeps its assertion (its mutation captures before the
apply starts).
Review follow-up. The first-contact pin's non-vacuity check counted only
writer alternations, so it could not show that an apply ever contended
with the writers — and none had: the schema actor applied back to back
the moment the start barrier opened, and the write-preferring lock ran
every apply before the first writer commit. The concurrent report now
counts era commits that land between two writer commits
(era_commits_between_data), the pin requires one somewhere across its
seeds, and the schema actor pauses a seeded 2-31 ms before each apply so
the applies spread across the writers' lifetime. Every apply of every
seed now lands between writer commits.

The hunt runs without readers: a reader's read-only open takes the
schema gate's exclusive side with no arbiter hook, so its transitions
fell outside the turns sched_escapes == 0 certifies. The pin's dead
scheduler branch, left from when its scheduler seed moved to the hunt,
is removed.
Review follow-up. A read on the writer's own handle still waits while
that handle publishes on its bound branch: the publish holds the
coordinator lock across the manifest compare-and-swap and the read
capture takes it. The architecture guide, the release note, the RFC's
behavior section and a test comment now say that, and the remaining
wait is listed under the RFC's unresolved questions.

The RFC's decision log records the independent review and the step 2
measurement: committing detached effects before the gates is 13-15% of
the gate hold and, because revalidation fails on any head move, would
turn most same-branch attempts into dead detached commits, so it is not
adopted. RFC 0067's plan text gains a dated amendment pointing there,
and its step 3 no longer describes a promotion pass that detached-only
tables removed.
The in-process stale-attempt check caught every failed revalidation but
gained only at object-store latency (0.40 to 0.53 commits/s with eight
writers on one branch at +30 ms) and lost locally (25.5 to 23.2), and it
left the starvation of #784 unchanged. The measurement bounds the whole
family: with branch-wide revalidation, eight writers on one branch cannot
exceed one writer alone, so raising same-branch throughput needs group
commit's footprint admission rather than more work on the gate.
…laim

A diagnostic build with one process-wide publication lock shows the whole
eight-branch gain comes from overlapping publications across branches;
interleaved reruns show no one-branch difference, so the earlier +36% was
drift between sessions.
@azimafroozeh
azimafroozeh force-pushed the claude/impl-shared-schema-gate branch from 35926d3 to bc09eec Compare September 28, 2026 21:27

@azimafroozeh azimafroozeh left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, Thanks!

@azimafroozeh
azimafroozeh added this pull request to the merge queue Sep 28, 2026
Merged via the queue into main with commit 8bc9aa5 Sep 28, 2026
30 checks passed
@azimafroozeh
azimafroozeh deleted the claude/impl-shared-schema-gate branch September 29, 2026 15:09
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants