From 900d0f2b6c1a293eb5e78ea1e34bc7534fefb90e Mon Sep 17 00:00:00 2001 From: Ragnor Comerford Date: Sat, 19 Sep 2026 14:39:13 +0200 Subject: [PATCH 01/13] engine: shared/exclusive schema-gate slot on the write-queue manager MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The mechanism for RFC 2026-09-18-shared-schema-gate, with zero callers: SchemaGateSlot (tokio RwLock + release epoch), SchemaSharedPermit / SchemaExclusivePermit whose drops follow the QueueGuard protocol (turn first, release inside it, epoch bump last), scheduled shared/exclusive acquisition that keeps scheduled_lock's three DST invariants (the contender never leaves the turn loop, installed mode attempts with try_* only, yield every iteration), HeldWriteGates (field order = drop order, schema first, matching the former guards[0] release sequence), and the acquire_schema_shared/exclusive manager methods. Unit tests: shared permits concurrent; exclusive excludes shared in both directions; the plain-mode fairness pin (a shared permit requested after a queued exclusive parks behind it — reds if a tokio upgrade changes the RwLock policy); one gate per canonical root across handles. --- crates/omnigraph/src/db/write_queue.rs | 318 ++++++++++++++++++++++++- 1 file changed, 314 insertions(+), 4 deletions(-) diff --git a/crates/omnigraph/src/db/write_queue.rs b/crates/omnigraph/src/db/write_queue.rs index 2f2f5f66b..900652b1f 100644 --- a/crates/omnigraph/src/db/write_queue.rs +++ b/crates/omnigraph/src/db/write_queue.rs @@ -14,7 +14,8 @@ //! Serialization remains in-process only; cross-process writers on one graph //! remain one-winner-CAS at publish. //! -//! ## Why exclusive `tokio::sync::Mutex<()>` per key +//! ## Lock shapes: exclusive `tokio::sync::Mutex<()>` per key, one +//! shared/exclusive schema gate //! //! Every writer stages detached from its table's pin and becomes visible //! only through the manifest CAS (RFC 0067), so the queue is not a @@ -22,9 +23,17 @@ //! keep same-process writers on one `(table_key, branch_ref)` from //! interleaving between revalidation and publication (the loser's detached //! versions would be wasted garbage) and to serialize destructive ref -//! deletion and fork creation against live writers. Every writer takes the -//! same exclusive lock; a shared/exclusive split would add a -//! writer-classification surface that's easy to get wrong. +//! deletion and fork creation against live writers. Table and branch keys +//! stay exclusive. The graph-global schema gate is the one +//! shared/exclusive slot ([`SchemaGateSlot`]): a pass that only READS the +//! accepted contract/catalog view takes a shared permit, and only a pass +//! that can CHANGE which view is accepted (schema apply, the system-column +//! upgrade, and the contract install/discard/reload passes) takes the +//! exclusive side — the classification rule and its safety argument are +//! RFC 2026-09-18-shared-schema-gate. Under the old protocol every writer +//! took this gate exclusively because a mutation could advance a Lance +//! HEAD before discovering an in-flight apply; detached staging abolished +//! that failure, leaving the gate as pure serialization. //! //! ## Sorted-order acquisition //! @@ -145,6 +154,167 @@ pub(crate) fn schema_apply_serial_queue_key() -> TableQueueKey { ("__schema_apply__".to_string(), None) } +/// The graph-global schema gate: the one shared/exclusive slot. +/// +/// It serializes every graph-global schema writer (schema apply and the +/// system-column upgrade) and every pass that installs, discards, or +/// republishes the accepted schema-contract view against each other — +/// exclusively — while readers of the accepted view (ordinary writers, +/// merges, maintenance, branch control, read-view captures) share. +/// +/// The gate is non-reentrant in BOTH modes on one task: an exclusive +/// holder re-acquiring either side self-deadlocks exactly like the +/// per-key mutexes, and a shared holder re-acquiring the shared side can +/// park forever behind a queued writer (tokio's `RwLock` is +/// write-preferring in plain mode). Callers therefore never take the gate +/// twice on one call path; `refresh_coordinator_only` exists for exactly +/// this reason. +/// +/// Fairness is asymmetric by mode: PLAIN mode inherits tokio's +/// write-preferring FIFO (once an exclusive acquisition is queued, later +/// shared permits park behind it, so writers cannot starve a schema +/// apply); INSTALLED (DST) mode never enters the native waiter queue — +/// both sides use `try_*` inside the turn loop, so grant order and +/// starvation-freedom are properties of the seed, asserted only by the +/// plain-mode unit test. +/// +/// The release epoch plays the same role as [`QueueSlot::releases`]: any +/// permit drop (shared or exclusive) takes a turn, releases inside it, +/// and bumps the epoch last, so contenders re-attempt only after a +/// turn-ordered release and same-seed runs cannot diverge on the unlock. +#[derive(Default)] +pub(crate) struct SchemaGateSlot { + lock: Arc>, + releases: std::sync::atomic::AtomicU64, +} + +/// A shared schema permit: proof that no contract-lifecycle pass is +/// concurrently swapping the accepted schema/catalog view. Held by +/// readers of that view for the duration of their gate-ordered work +/// (writers: through manifest publish). +#[must_use = "dropping the permit releases the shared schema gate"] +pub(crate) struct SchemaSharedPermit { + inner: Option>, + slot: Arc, +} + +/// An exclusive schema permit: sole ownership of the accepted-view +/// transition. Held by schema apply, the system-column upgrade, and the +/// contract install/discard/reload passes. +#[must_use = "dropping the permit releases the exclusive schema gate"] +pub(crate) struct SchemaExclusivePermit { + inner: Option>, + slot: Arc, +} + +impl Drop for SchemaSharedPermit { + fn drop(&mut self) { + // Same protocol as `QueueGuard`: turn first, release inside it, + // bump the epoch last so an observed bump implies a genuinely + // released reader slot. + let _turn = crate::dst_gate::turn(); + drop(self.inner.take()); + self.slot + .releases + .fetch_add(1, std::sync::atomic::Ordering::SeqCst); + } +} + +impl Drop for SchemaExclusivePermit { + fn drop(&mut self) { + let _turn = crate::dst_gate::turn(); + drop(self.inner.take()); + self.slot + .releases + .fetch_add(1, std::sync::atomic::Ordering::SeqCst); + } +} + +/// Scheduled shared acquisition of the schema gate; the exact +/// [`scheduled_lock`] protocol on the read side. Uninstalled: plain +/// blocking `read_owned` (fair, write-preferring). Installed: stay in the +/// turn loop, `try_read_owned` only when the release epoch moved, yield +/// every iteration. +async fn scheduled_schema_shared(slot: Arc) -> SchemaSharedPermit { + let mut wait_epoch: Option = None; + loop { + match crate::dst_gate::turn() { + None => { + let guard = Arc::clone(&slot.lock).read_owned().await; + return SchemaSharedPermit { + inner: Some(guard), + slot, + }; + } + Some(_turn) => { + let epoch_now = slot.releases.load(std::sync::atomic::Ordering::SeqCst); + if wait_epoch.is_none_or(|e| epoch_now != e) { + match Arc::clone(&slot.lock).try_read_owned() { + Ok(guard) => { + return SchemaSharedPermit { + inner: Some(guard), + slot: Arc::clone(&slot), + }; + } + Err(_) => wait_epoch = Some(epoch_now), + } + } + // else: no release since the failed attempt — a no-op turn. + } + } + tokio::task::yield_now().await; + } +} + +/// Scheduled exclusive acquisition of the schema gate; the write-side +/// twin of [`scheduled_schema_shared`]. +async fn scheduled_schema_exclusive(slot: Arc) -> SchemaExclusivePermit { + let mut wait_epoch: Option = None; + loop { + match crate::dst_gate::turn() { + None => { + let guard = Arc::clone(&slot.lock).write_owned().await; + return SchemaExclusivePermit { + inner: Some(guard), + slot, + }; + } + Some(_turn) => { + let epoch_now = slot.releases.load(std::sync::atomic::Ordering::SeqCst); + if wait_epoch.is_none_or(|e| epoch_now != e) { + match Arc::clone(&slot.lock).try_write_owned() { + Ok(guard) => { + return SchemaExclusivePermit { + inner: Some(guard), + slot: Arc::clone(&slot), + }; + } + Err(_) => wait_epoch = Some(epoch_now), + } + } + // else: no release since the failed attempt — a no-op turn. + } + } + tokio::task::yield_now().await; + } +} + +/// Ordered write-gate envelope for an RFC-022 effect writer: shared +/// schema permit, then the branch gate, then the lex-sorted table gates, +/// held from revalidation through manifest publish. +/// +/// Field order is drop order — schema releases first, then branch, then +/// tables — preserving the exact release sequence the former homogeneous +/// guard Vec produced, so seeded DST grant sequences keep their event +/// order at release. +#[must_use = "dropping the gates releases the write envelope"] +pub(crate) struct HeldWriteGates { + pub(crate) schema: SchemaSharedPermit, + /// `queue[0]` is the branch gate, followed by the lex-sorted table + /// gates. + pub(crate) queue: Vec, +} + /// Non-cloneable ownership of the sole immutable export cut for one graph. /// /// The root registry stores weak references, so retaining the manager here is @@ -190,6 +360,9 @@ pub(crate) struct WriteQueueManager { /// destructive controls share the read side, so they remain mutually /// concurrent but cannot remove a path/version beneath a live cut. export_gate: Arc>, + /// The graph-global shared/exclusive schema gate; see + /// [`SchemaGateSlot`]. + schema_gate: Arc, } impl WriteQueueManager { @@ -249,6 +422,23 @@ impl WriteQueueManager { scheduled_lock(self.branch_slot(&key)).await } + /// Take the schema gate's shared side: proof that no + /// contract-lifecycle pass is concurrently swapping the accepted + /// schema/catalog view. Acquire BEFORE the branch gate and any table + /// queue; never re-acquire either side while holding a permit (the + /// gate is non-reentrant — see [`SchemaGateSlot`]). + pub(crate) async fn acquire_schema_shared(&self) -> SchemaSharedPermit { + scheduled_schema_shared(Arc::clone(&self.schema_gate)).await + } + + /// Take the schema gate's exclusive side: sole ownership of the + /// accepted-view transition. Shared holders drain first; new shared + /// permits queue behind this acquisition in plain mode. Same + /// non-reentrancy contract as [`Self::acquire_schema_shared`]. + pub(crate) async fn acquire_schema_exclusive(&self) -> SchemaExclusivePermit { + scheduled_schema_exclusive(Arc::clone(&self.schema_gate)).await + } + /// Reserve the sole immutable export cut without waiting. pub(crate) fn try_acquire_export_cut(self: &Arc) -> Option { let permit = Arc::clone(&self.export_gate).try_write_owned().ok()?; @@ -399,6 +589,126 @@ mod tests { assert!(second.try_acquire_export_cut().is_some()); } + #[tokio::test] + async fn schema_shared_permits_are_concurrent() { + let qm = Arc::new(WriteQueueManager::new()); + let first = qm.acquire_schema_shared().await; + let qm2 = Arc::clone(&qm); + let second = timeout(Duration::from_secs(2), async move { + qm2.acquire_schema_shared().await + }) + .await + .expect("a second shared permit must not wait behind the first"); + drop(first); + drop(second); + } + + #[tokio::test] + async fn schema_exclusive_excludes_shared_both_directions() { + let qm = Arc::new(WriteQueueManager::new()); + + // Held exclusive blocks a shared acquire. + let exclusive = qm.acquire_schema_exclusive().await; + let qm2 = Arc::clone(&qm); + let blocked = timeout(Duration::from_millis(200), async move { + qm2.acquire_schema_shared().await + }) + .await; + assert!( + blocked.is_err(), + "a shared permit must wait behind a held exclusive permit" + ); + drop(exclusive); + let qm2 = Arc::clone(&qm); + let shared = timeout(Duration::from_secs(2), async move { + qm2.acquire_schema_shared().await + }) + .await + .expect("shared must acquire once the exclusive permit releases"); + + // Held shared blocks an exclusive acquire. + let qm2 = Arc::clone(&qm); + let blocked = timeout(Duration::from_millis(200), async move { + qm2.acquire_schema_exclusive().await + }) + .await; + assert!( + blocked.is_err(), + "an exclusive permit must wait behind a held shared permit" + ); + drop(shared); + let qm2 = Arc::clone(&qm); + timeout(Duration::from_secs(2), async move { + qm2.acquire_schema_exclusive().await + }) + .await + .expect("exclusive must acquire once the shared permit releases"); + } + + /// The plain-mode no-starvation pin: tokio's `RwLock` is + /// write-preferring, so once an exclusive acquisition is queued, a + /// LATER shared acquisition parks behind it instead of overtaking. If + /// a tokio upgrade ever changes that policy, this test reds and the + /// schema gate's fairness claim (RFC 2026-09-18-shared-schema-gate) + /// must be re-derived. + #[tokio::test] + async fn queued_schema_exclusive_blocks_later_shared() { + let qm = Arc::new(WriteQueueManager::new()); + let held_shared = qm.acquire_schema_shared().await; + + let qm_writer = Arc::clone(&qm); + let writer = tokio::spawn(async move { qm_writer.acquire_schema_exclusive().await }); + // Give the exclusive acquisition time to enter the waiter queue. + tokio::time::sleep(Duration::from_millis(100)).await; + + let qm2 = Arc::clone(&qm); + let overtaking = timeout(Duration::from_millis(200), async move { + qm2.acquire_schema_shared().await + }) + .await; + assert!( + overtaking.is_err(), + "a shared permit requested after a queued exclusive must park behind it" + ); + + drop(held_shared); + let exclusive = timeout(Duration::from_secs(2), writer) + .await + .expect("queued exclusive must acquire once the shared permit releases") + .expect("writer task must not panic"); + drop(exclusive); + let qm2 = Arc::clone(&qm); + timeout(Duration::from_secs(2), async move { + qm2.acquire_schema_shared().await + }) + .await + .expect("the parked shared permit must acquire after the exclusive releases"); + } + + #[tokio::test] + async fn schema_gate_shared_across_for_root_handles() { + let root = format!("memory://schema-gate/{}", ulid::Ulid::new()); + let first = WriteQueueManager::for_root(&root); + let second = WriteQueueManager::for_root(&root); + + let exclusive = first.acquire_schema_exclusive().await; + let second2 = Arc::clone(&second); + let blocked = timeout(Duration::from_millis(200), async move { + second2.acquire_schema_shared().await + }) + .await; + assert!( + blocked.is_err(), + "handles for one root must exclude on one schema gate" + ); + drop(exclusive); + timeout(Duration::from_secs(2), async move { + second.acquire_schema_shared().await + }) + .await + .expect("the second handle's shared permit must acquire after release"); + } + #[tokio::test] async fn acquire_many_sorts_keys_deterministically() { // Two callers passing keys in different orders must acquire in From b9f064064e5d3c6444cc3270b355e7efac6852c2 Mon Sep 17 00:00:00 2001 From: Ragnor Comerford Date: Sat, 19 Sep 2026 14:48:09 +0200 Subject: [PATCH 02/13] engine: flip the schema gate to shared/exclusive at every site MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The atomic half of RFC 2026-09-18-shared-schema-gate: all 22 acquisition sites move to the permit API in one commit so no mixed-mechanism window exists, and schema_apply_serial_queue_key() is deleted with its synthetic ("__schema_apply__", None) entry. Exclusive (contract-lifecycle passes, unchanged mutual exclusion): schema apply, the system-column upgrade, open, sync_branch, refresh, settle_pending_schema_install, reload_schema_if_source_changed. Shared (readers of the accepted view — they stop excluding each other and keep excluding every contract-lifecycle pass): mutation/load commit_all, the no-op conditional-mutation CAS, branch merge, ensure_indices, optimize, cleanup, repair, orphaned-branch reconciliation, branch create/create-from/delete, ensure_no_pending_recovery, and the three read-view captures (reads no longer serialize behind writers' publish holds). The staged write returns HeldWriteGates (shared permit + branch + sorted table guards; field order preserves the former release order, so seeded DST grant sequences keep their event order) in place of the anonymous guard Vec. The apply-side comments are re-derived to name the permit and the shared-re-entry parking hazard. --- crates/omnigraph/src/db/omnigraph.rs | 62 +++++-------------- crates/omnigraph/src/db/omnigraph/optimize.rs | 15 ++--- crates/omnigraph/src/db/omnigraph/repair.rs | 5 +- .../src/db/omnigraph/schema_apply.rs | 29 +++++---- .../src/db/omnigraph/system_column_upgrade.rs | 3 +- .../omnigraph/src/db/omnigraph/table_ops.rs | 5 +- crates/omnigraph/src/db/write_queue.rs | 33 +++++----- crates/omnigraph/src/exec/merge.rs | 5 +- crates/omnigraph/src/exec/mutation.rs | 12 ++-- crates/omnigraph/src/exec/staging.rs | 32 +++++----- crates/omnigraph/src/loader/mod.rs | 13 ++-- 11 files changed, 84 insertions(+), 130 deletions(-) diff --git a/crates/omnigraph/src/db/omnigraph.rs b/crates/omnigraph/src/db/omnigraph.rs index cb3211e82..0f79ebb27 100644 --- a/crates/omnigraph/src/db/omnigraph.rs +++ b/crates/omnigraph/src/db/omnigraph.rs @@ -731,9 +731,7 @@ impl Omnigraph { let storage = storage_for_uri(&root)?; let identity = write_queue_root_identity(&root)?; let queues = crate::db::write_queue::WriteQueueManager::for_root(&identity); - let _schema_gate = queues - .acquire(&crate::db::write_queue::schema_apply_serial_queue_key()) - .await; + let _schema_gate = queues.acquire_schema_shared().await; crate::db::upgrade::legacy_sidecars::refuse_pending_recovery(&root, storage.as_ref()).await } @@ -787,9 +785,7 @@ impl Omnigraph { // Hold the same schema gate through format preflight and contract // capture. A v3 live or staged IR must refuse before the local write // probe, coordinator open, or the staged-contract pass can change files. - let schema_contract_guard = write_queue - .acquire(&crate::db::write_queue::schema_apply_serial_queue_key()) - .await; + let schema_contract_guard = write_queue.acquire_schema_exclusive().await; crate::db::schema_state::refuse_unsupported_schema_versions(&root, storage.as_ref()) .await?; // Read-write opens write before the first user mutation (the @@ -1083,8 +1079,8 @@ impl Omnigraph { /// validating the files does not make `self.catalog()` current after another /// handle applies a schema. Control/legacy-adapter bridges use this capture /// for planning and conservative table-gate enumeration. The caller MUST - /// already hold `schema_apply_serial_queue_key`; this helper does not acquire - /// it because the gate is a non-reentrant mutex. + /// already hold a schema permit (either side); this helper does not acquire + /// one because the gate is non-reentrant on one task. pub(crate) async fn load_accepted_catalog_with_schema_gate_held(&self) -> Result> { let catalog = self.build_accepted_catalog_with_schema_gate_held().await?; let snapshot = self.coordinator.read().await.snapshot(); @@ -1987,10 +1983,7 @@ impl Omnigraph { // authority window. This also // captures the schema contract and target coordinator coherently across // a concurrent schema apply. Lock order remains schema -> coordinator. - let _schema_guard = self - .write_queue() - .acquire(&crate::db::write_queue::schema_apply_serial_queue_key()) - .await; + let _schema_permit = self.write_queue().acquire_schema_exclusive().await; let (schema_ir, _) = load_validated_schema_contract(self.uri(), Arc::clone(&self.storage)).await?; let branch = normalize_branch_name(branch)?; @@ -2036,10 +2029,7 @@ impl Omnigraph { // `reload_schema_if_source_changed` takes the coordinator read // lock, and Tokio's RwLock is not reentrant. Pinned by // `composite_flow_schema_apply_then_branch_ops_no_deadlock_in_refresh`. - let _serial = self - .write_queue - .acquire(&crate::db::write_queue::schema_apply_serial_queue_key()) - .await; + let _serial = self.write_queue.acquire_schema_exclusive().await; let mut coord = self.coordinator.write().await; coord.refresh().await?; let outcome = recover_schema_state_files( @@ -2107,10 +2097,7 @@ impl Omnigraph { return Ok(()); } let result = { - let _serial = self - .write_queue - .acquire(&crate::db::write_queue::schema_apply_serial_queue_key()) - .await; + let _serial = self.write_queue.acquire_schema_exclusive().await; let snapshot = self.coordinator.read().await.snapshot(); recover_schema_state_files( &self.root_uri, @@ -2141,10 +2128,7 @@ impl Omnigraph { // across the complete source/IR/state read and ArcSwap publication so a // concurrent apply cannot interleave its sequential file promotions // with this reload. - let _schema_guard = self - .write_queue - .acquire(&crate::db::write_queue::schema_apply_serial_queue_key()) - .await; + let _schema_permit = self.write_queue.acquire_schema_exclusive().await; fail(&SCHEMA_RELOAD_BEFORE_CONTRACT_READ)?; let schema_path = schema_source_uri(&self.root_uri); let schema_source = self.storage.read_text(&schema_path).await?; @@ -2252,10 +2236,7 @@ impl Omnigraph { let target = target.into(); let validate_live_snapshot = matches!(&target, ReadTarget::Branch(_)); let bind_historical_aliases = matches!(&target, ReadTarget::Snapshot(_)); - let _schema_guard = self - .write_queue() - .acquire(&crate::db::write_queue::schema_apply_serial_queue_key()) - .await; + let _schema_permit = self.write_queue().acquire_schema_shared().await; let catalog = self.build_accepted_catalog_with_schema_gate_held().await?; let mut resolved = self.resolve_target_after_schema_validation(target).await?; if validate_live_snapshot { @@ -2273,10 +2254,7 @@ impl Omnigraph { } pub(crate) async fn capture_current_read_view(&self) -> Result<(ResolvedTarget, Arc)> { - let _schema_guard = self - .write_queue() - .acquire(&crate::db::write_queue::schema_apply_serial_queue_key()) - .await; + let _schema_permit = self.write_queue().acquire_schema_shared().await; let current_branch = self .coordinator .read() @@ -2296,10 +2274,7 @@ impl Omnigraph { &self, version: u64, ) -> Result<(Snapshot, Arc)> { - let _schema_guard = self - .write_queue() - .acquire(&crate::db::write_queue::schema_apply_serial_queue_key()) - .await; + let _schema_permit = self.write_queue().acquire_schema_shared().await; let catalog = self.load_accepted_catalog_with_schema_gate_held().await?; let branch = self .coordinator @@ -3026,10 +3001,7 @@ impl Omnigraph { let source = self.active_branch().await; self.settle_pending_schema_install().await?; fail(&BRANCH_CONTROL_PRE_GATES)?; - let _schema_guard = self - .write_queue() - .acquire(&crate::db::write_queue::schema_apply_serial_queue_key()) - .await; + let _schema_permit = self.write_queue().acquire_schema_shared().await; let _branch_guards = self .write_queue() .acquire_branches(&[source.clone(), Some(target.clone())]) @@ -3112,10 +3084,7 @@ impl Omnigraph { self.ensure_schema_state_valid().await?; self.settle_pending_schema_install().await?; fail(&BRANCH_CONTROL_PRE_GATES)?; - let _schema_guard = self - .write_queue() - .acquire(&crate::db::write_queue::schema_apply_serial_queue_key()) - .await; + let _schema_permit = self.write_queue().acquire_schema_shared().await; let _branch_guards = self .write_queue() .acquire_branches(&[branch.clone(), Some(target_branch.clone())]) @@ -3175,10 +3144,7 @@ impl Omnigraph { self.ensure_schema_state_valid().await?; self.settle_pending_schema_install().await?; fail(&BRANCH_CONTROL_PRE_GATES)?; - let _schema_guard = self - .write_queue() - .acquire(&crate::db::write_queue::schema_apply_serial_queue_key()) - .await; + let _schema_permit = self.write_queue().acquire_schema_shared().await; let _branch_guard = self.write_queue().acquire_branch(Some(&branch)).await; // Purge only after taking the branch gate. Merge capture takes the // same branch-gate -> cache-lock order, so no later insert for this diff --git a/crates/omnigraph/src/db/omnigraph/optimize.rs b/crates/omnigraph/src/db/omnigraph/optimize.rs index 9fc1094fb..5a5e2c2bf 100644 --- a/crates/omnigraph/src/db/omnigraph/optimize.rs +++ b/crates/omnigraph/src/db/omnigraph/optimize.rs @@ -236,8 +236,7 @@ pub async fn optimize_all_datasets(db: &Omnigraph) -> Result branch -> sorted tables. Planning reads // catalog index intent, so it must use an operation-local accepted catalog // under the same schema gate as schema apply and the exact RFC-022 writers. - let schema_gate_key = crate::db::write_queue::schema_apply_serial_queue_key(); - let schema_guard = db.write_queue().acquire(&schema_gate_key).await; + let schema_permit = db.write_queue().acquire_schema_shared().await; db.refresh_coordinator_only().await?; db.ensure_schema_apply_not_locked("optimize").await?; let catalog = db.load_accepted_catalog_with_schema_gate_held().await?; @@ -358,7 +357,7 @@ pub async fn optimize_all_datasets(db: &Omnigraph) -> Result Result { - let _schema = db - .write_queue() - .acquire(&crate::db::write_queue::schema_apply_serial_queue_key()) - .await; + let _schema = db.write_queue().acquire_schema_shared().await; let catalog = db.catalog(); let graph_branches = cleanup_graph_branches(db).await?; let _branches = db.write_queue().acquire_branches(&graph_branches).await; diff --git a/crates/omnigraph/src/db/omnigraph/repair.rs b/crates/omnigraph/src/db/omnigraph/repair.rs index 4258ff5bc..b54a86110 100644 --- a/crates/omnigraph/src/db/omnigraph/repair.rs +++ b/crates/omnigraph/src/db/omnigraph/repair.rs @@ -158,10 +158,7 @@ pub async fn repair_all_datasets(db: &Omnigraph, options: RepairOptions) -> Resu // final publish all remain under schema -> main -> sorted-table gates. This // prevents a concurrent drop/re-add from pairing the old dataset path with // the replacement's same public alias and new identity. - let _schema_guard = db - .write_queue() - .acquire(&crate::db::write_queue::schema_apply_serial_queue_key()) - .await; + let _schema_permit = db.write_queue().acquire_schema_shared().await; db.refresh_coordinator_only().await?; db.ensure_schema_apply_not_locked("repair").await?; let catalog = db.load_accepted_catalog_with_schema_gate_held().await?; diff --git a/crates/omnigraph/src/db/omnigraph/schema_apply.rs b/crates/omnigraph/src/db/omnigraph/schema_apply.rs index d55ee94a0..612cf7fbe 100644 --- a/crates/omnigraph/src/db/omnigraph/schema_apply.rs +++ b/crates/omnigraph/src/db/omnigraph/schema_apply.rs @@ -254,15 +254,16 @@ where // before planning against the accepted contract. db.settle_pending_schema_install().await?; - // Process-local schema-control gate. RFC-022 mutation/load commit paths - // acquire this before their branch/table gates and retain it through - // publication. Taking it before the durable sentinel closes the old race in - // which schema apply could create the sentinel while a mutation already held - // a table queue, causing that mutation to advance Lance HEAD and only then - // discover the schema lock. The native sentinel remains the cross-handle / - // crash-visible authority; this queue removes the avoidable same-handle race. - let schema_gate_key = crate::db::write_queue::schema_apply_serial_queue_key(); - let _schema_gate = db.write_queue().acquire(&schema_gate_key).await; + // Process-local schema gate, EXCLUSIVE side: schema apply is a + // contract-lifecycle pass, so it excludes every shared holder (writers, + // maintenance, branch control, read captures) and they exclude it. The + // permit is taken before the branch/table gates and before the durable + // sentinel, and retained through sentinel release, so no shared holder + // can revalidate against a contract this apply is about to replace. + // The native sentinel remains the cross-handle / crash-visible + // authority; this permit removes the avoidable same-handle race (RFC + // 2026-09-18-shared-schema-gate). + let _schema_gate = db.write_queue().acquire_schema_exclusive().await; acquire_schema_apply_lock(db).await?; let result = apply_schema_with_lock(db, desired_schema_source, options, actor, validate_catalog).await; @@ -615,10 +616,12 @@ where .datasets() .map(|entry| (entry.type_key.clone(), entry.native_dataset_branch.clone())) .collect(); - // The outer `apply_schema` holds the schema-control serialization key from - // before sentinel creation through sentinel release. Per-table guards here - // therefore cover only the concrete table effects; acquiring the schema key - // again would deadlock because these queues are intentionally non-reentrant. + // The outer `apply_schema` holds the exclusive schema permit from before + // sentinel creation through sentinel release. Per-table guards here + // therefore cover only the concrete table effects; re-acquiring either + // side of the schema gate on this task deadlocks — the gate is + // non-reentrant, and even a shared re-entry parks behind any queued + // writer under the plain-mode write-preferring lock. let _main_branch_guard = db.write_queue().acquire_branch(None).await; let _schema_apply_queue_guards = db .write_queue() diff --git a/crates/omnigraph/src/db/omnigraph/system_column_upgrade.rs b/crates/omnigraph/src/db/omnigraph/system_column_upgrade.rs index 95e36dd2e..24af25133 100644 --- a/crates/omnigraph/src/db/omnigraph/system_column_upgrade.rs +++ b/crates/omnigraph/src/db/omnigraph/system_column_upgrade.rs @@ -175,8 +175,7 @@ pub(super) async fn upgrade_system_columns( if !options.check { db.settle_pending_schema_install().await?; } - let schema_gate_key = crate::db::write_queue::schema_apply_serial_queue_key(); - let _schema_gate = db.write_queue().acquire(&schema_gate_key).await; + let _schema_gate = db.write_queue().acquire_schema_exclusive().await; db.refresh_coordinator_only().await?; let stamp = crate::db::manifest::internal_schema_stamp_at(db.uri(), None) .await? diff --git a/crates/omnigraph/src/db/omnigraph/table_ops.rs b/crates/omnigraph/src/db/omnigraph/table_ops.rs index 2ae1c0daf..799337c11 100644 --- a/crates/omnigraph/src/db/omnigraph/table_ops.rs +++ b/crates/omnigraph/src/db/omnigraph/table_ops.rs @@ -358,10 +358,7 @@ async fn maintain_indices_for_branch( .iter() .map(|target| (target.table_key.clone(), active_branch.clone())) .collect(); - let _schema_guard = db - .write_queue() - .acquire(&crate::db::write_queue::schema_apply_serial_queue_key()) - .await; + let _schema_permit = db.write_queue().acquire_schema_shared().await; let _branch_guard = db .write_queue() .acquire_branch(active_branch.as_deref()) diff --git a/crates/omnigraph/src/db/write_queue.rs b/crates/omnigraph/src/db/write_queue.rs index 900652b1f..e62263293 100644 --- a/crates/omnigraph/src/db/write_queue.rs +++ b/crates/omnigraph/src/db/write_queue.rs @@ -145,15 +145,6 @@ async fn scheduled_lock(slot: Arc) -> QueueGuard { /// serialize at the queue. pub(crate) type TableQueueKey = (String, Option); -/// The write-queue key that serializes every graph-global schema writer -/// (schema apply and the system-column upgrade) against each other and -/// against the passes that install or discard a staged schema contract. The -/// name cannot collide with real table keys (those are `node:`/`edge:` -/// prefixed). -pub(crate) fn schema_apply_serial_queue_key() -> TableQueueKey { - ("__schema_apply__".to_string(), None) -} - /// The graph-global schema gate: the one shared/exclusive slot. /// /// It serializes every graph-global schema writer (schema apply and the @@ -309,10 +300,20 @@ async fn scheduled_schema_exclusive(slot: Arc) -> SchemaExclusiv /// order at release. #[must_use = "dropping the gates releases the write envelope"] pub(crate) struct HeldWriteGates { - pub(crate) schema: SchemaSharedPermit, - /// `queue[0]` is the branch gate, followed by the lex-sorted table - /// gates. - pub(crate) queue: Vec, + _schema: SchemaSharedPermit, + /// `[0]` is the branch gate, followed by the lex-sorted table gates. + _queue: Vec, +} + +impl HeldWriteGates { + /// Assemble the envelope from the gates in acquisition order: the + /// shared schema permit, then the branch gate and sorted table gates. + pub(crate) fn new(schema: SchemaSharedPermit, queue: Vec) -> Self { + Self { + _schema: schema, + _queue: queue, + } + } } /// Non-cloneable ownership of the sole immutable export cut for one graph. @@ -638,7 +639,7 @@ mod tests { ); drop(shared); let qm2 = Arc::clone(&qm); - timeout(Duration::from_secs(2), async move { + let _exclusive = timeout(Duration::from_secs(2), async move { qm2.acquire_schema_exclusive().await }) .await @@ -678,7 +679,7 @@ mod tests { .expect("writer task must not panic"); drop(exclusive); let qm2 = Arc::clone(&qm); - timeout(Duration::from_secs(2), async move { + let _shared = timeout(Duration::from_secs(2), async move { qm2.acquire_schema_shared().await }) .await @@ -702,7 +703,7 @@ mod tests { "handles for one root must exclude on one schema gate" ); drop(exclusive); - timeout(Duration::from_secs(2), async move { + let _shared = timeout(Duration::from_secs(2), async move { second.acquire_schema_shared().await }) .await diff --git a/crates/omnigraph/src/exec/merge.rs b/crates/omnigraph/src/exec/merge.rs index e5efff274..fe08849f4 100644 --- a/crates/omnigraph/src/exec/merge.rs +++ b/crates/omnigraph/src/exec/merge.rs @@ -5333,10 +5333,7 @@ impl Omnigraph { // Holding both branch gates through publication prevents a target // delete/recreate from reusing the branch name underneath a plan (ABA). self.settle_pending_schema_install().await?; - let _schema_guard = self - .write_queue() - .acquire(&crate::db::write_queue::schema_apply_serial_queue_key()) - .await; + let _schema_permit = self.write_queue().acquire_schema_shared().await; let _branch_guards = self .write_queue() .acquire_branches(&[source_branch.clone(), target_branch.clone()]) diff --git a/crates/omnigraph/src/exec/mutation.rs b/crates/omnigraph/src/exec/mutation.rs index b2cf97bcf..15d4f5987 100644 --- a/crates/omnigraph/src/exec/mutation.rs +++ b/crates/omnigraph/src/exec/mutation.rs @@ -976,12 +976,10 @@ impl Omnigraph { // A no-op has no table transaction, so it never reaches // `commit_all`. It still needs a linearization point for // the caller's CAS promise: under the same schema -> branch - // ordering as effectful writes, re-read the complete + // ordering as effectful writes (shared permit — this pass + // only reads the accepted view), re-read the complete // authority and map a moved caller head to terminal 412. - let _schema_guard = self - .write_queue() - .acquire(&crate::db::write_queue::schema_apply_serial_queue_key()) - .await; + let _schema_permit = self.write_queue().acquire_schema_shared().await; let _branch_guard = self .write_queue() .acquire_branch(requested.as_deref()) @@ -1002,7 +1000,7 @@ impl Omnigraph { let lineage_intent = self .new_lineage_intent_for_branch(requested.as_deref(), actor_id) .await?; - // `_queue_guards` holds the root-shared schema gate, branch + // `_held_gates` holds the shared schema permit, branch // effect gate, and sorted table gates acquired by `commit_all`. // They remain held through manifest publication, covering the // complete same-process effect lifetime. They are a local @@ -1011,7 +1009,7 @@ impl Omnigraph { let super::staging::CommittedMutation { updates, expected_versions, - guards: _queue_guards, + gates: _held_gates, } = staged.commit_all(self, requested.as_deref(), &txn).await?; // Failpoint for the detached-effects → publisher boundary: // every table effect is committed detached but nothing is diff --git a/crates/omnigraph/src/exec/staging.rs b/crates/omnigraph/src/exec/staging.rs index 42e1987f9..290546bf8 100644 --- a/crates/omnigraph/src/exec/staging.rs +++ b/crates/omnigraph/src/exec/staging.rs @@ -825,10 +825,11 @@ pub(crate) struct CommittedMutation { /// publisher checks them together with native branch identity, exact graph /// head, and schema identity as one authority precondition. pub(crate) expected_versions: crate::db::manifest::ExpectedTableVersions, - /// Root schema, coarse branch, and sorted `(table, branch)` guards. The - /// caller MUST hold the complete set across manifest publish (see - /// `commit_all`) so no same-process writer interleaves after revalidation. - pub(crate) guards: Vec, + /// The write envelope: shared schema permit, coarse branch gate, and + /// sorted `(table, branch)` guards. The caller MUST hold the complete + /// set across manifest publish (see `commit_all`) so no same-process + /// writer interleaves after revalidation. + pub(crate) gates: crate::db::write_queue::HeldWriteGates, } decide_seam! { @@ -844,7 +845,7 @@ impl StagedMutation { /// base (RFC 0067), and return the publisher input plus guards. No Lance /// HEAD moves here, and the published pins stay detached. /// - /// **Caller must hold the returned `_guards` Vec across the + /// **Caller must hold the returned gates across the /// subsequent manifest publish.** Releasing guards before publish /// would let another same-process writer publish between our detached /// commits and our publish. The exact publisher precondition would still @@ -881,15 +882,16 @@ impl StagedMutation { for entry in &staged { queue_keys.push((entry.table_key.clone(), branch.map(str::to_string))); } - // Total order shared with schema apply: schema gate, branch gate, then - // sorted per-table gates. Hold the full set through manifest publish. - let schema_guard = db - .write_queue() - .acquire(&crate::db::write_queue::schema_apply_serial_queue_key()) - .await; + // Total order shared with schema apply: schema permit (SHARED — a + // writer only reads the accepted contract view; schema apply's + // exclusive side excludes every writer, RFC + // 2026-09-18-shared-schema-gate), branch gate, then sorted + // per-table gates. Hold the full set through manifest publish. + let schema = db.write_queue().acquire_schema_shared().await; let branch_guard = db.write_queue().acquire_branch(branch).await; - let mut guards = vec![schema_guard, branch_guard]; - guards.extend(db.write_queue().acquire_many(&queue_keys).await); + let mut queue = vec![branch_guard]; + queue.extend(db.write_queue().acquire_many(&queue_keys).await); + let gates = crate::db::write_queue::HeldWriteGates::new(schema, queue); // Re-capture manifest pins under the queue (PR 2 / MR-686). // @@ -937,7 +939,7 @@ impl StagedMutation { return Ok(CommittedMutation { updates: Vec::new(), expected_versions, - guards, + gates, }); } @@ -996,7 +998,7 @@ impl StagedMutation { Ok(CommittedMutation { updates, expected_versions, - guards, + gates, }) } } diff --git a/crates/omnigraph/src/loader/mod.rs b/crates/omnigraph/src/loader/mod.rs index d1392c6d0..6b12c31f9 100644 --- a/crates/omnigraph/src/loader/mod.rs +++ b/crates/omnigraph/src/loader/mod.rs @@ -893,15 +893,16 @@ async fn load_jsonl_reader_once( .await?; fail(&catalog::MUTATION_POST_STAGE_PRE_EFFECT_GATE)?; let lineage_intent = db.new_lineage_intent_for_branch(branch, actor_id).await?; - // `_queue_guards` holds the root-shared schema → branch → sorted-table - // gates across manifest publication. This closes same-process - // interleaving across the effect lifetime. The exact publisher token - // remains the persistent correctness authority; these local gates do not - // expand the documented single-writer-process boundary. + // `held_gates` holds the root-shared schema permit → branch → + // sorted-table gates across manifest publication. This closes + // same-process interleaving across the effect lifetime. The exact + // publisher token remains the persistent correctness authority; these + // local gates do not expand the documented single-writer-process + // boundary. let crate::exec::staging::CommittedMutation { updates, expected_versions, - guards: _queue_guards, + gates: _held_gates, } = staged.commit_all(db, branch, &txn).await?; // Same detached-effects → publisher boundary as mutations: every table // effect is committed detached, but the graph manifest has not published From d6ddc2129d92c1a6c39868d405a6004e114d0766 Mon Sep 17 00:00:00 2001 From: Ragnor Comerford Date: Sat, 19 Sep 2026 14:53:20 +0200 Subject: [PATCH 03/13] test: pin the shared schema gate's behavior in both directions MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Three regression owners for RFC 2026-09-18-shared-schema-gate: - cross_branch_writers_overlap_inside_schema_gate — the headline: a writer parked inside its full envelope on b1 no longer blocks a b2 writer, which runs to completion while A is parked (impossible under the exclusive mutex). - read_capture_proceeds_while_writer_parked — reads capture their catalog under a shared permit and stop waiting for a writer's publish hold. - parked_writer_blocks_schema_apply — the reverse direction and the mis-classification tripwire: a held shared permit blocks the apply's exclusive acquisition outright; a writer wrongly left off the gate would let the apply proceed mid-write and red this test. The existing mid-apply pin keeps its assertion; its comment now names the shared-behind-exclusive mechanism. --- crates/omnigraph/tests/failpoints.rs | 119 +++++++++++++++++++++++++ crates/omnigraph/tests/schema_apply.rs | 64 ++++++++++++- 2 files changed, 181 insertions(+), 2 deletions(-) diff --git a/crates/omnigraph/tests/failpoints.rs b/crates/omnigraph/tests/failpoints.rs index 9f59cf58e..acea15fcc 100644 --- a/crates/omnigraph/tests/failpoints.rs +++ b/crates/omnigraph/tests/failpoints.rs @@ -1825,6 +1825,125 @@ async fn cross_handle_branch_gate_serializes_post_effect_publish() { ); } +/// The shared schema gate's headline behavior (RFC +/// 2026-09-18-shared-schema-gate): writers on DISJOINT branches overlap +/// inside the schema gate. Writer A parks inside its full write envelope +/// (shared schema permit + b1 branch gate + table gates, detached commits +/// done, publish pending); writer B on b2 must run to completion while A is +/// parked — impossible under the former exclusive schema mutex, where B +/// would queue behind A's publish hold. +#[tokio::test(flavor = "multi_thread", worker_threads = 4)] +#[serial] +async fn cross_branch_writers_overlap_inside_schema_gate() { + let _scenario = FailScenario::setup(); + let dir = tempfile::tempdir().unwrap(); + let db = helpers::init_and_load(&dir).await; + db.branch_create("b1").await.unwrap(); + db.branch_create("b2").await.unwrap(); + let db = std::sync::Arc::new(db); + + let in_envelope = + helpers::failpoint::Rendezvous::park_first(&catalog::MUTATION_POST_FINALIZE_PRE_PUBLISHER); + + let writer_a_db = std::sync::Arc::clone(&db); + let writer_a = tokio::spawn(async move { + writer_a_db + .mutate( + "b1", + MUTATION_QUERIES, + "insert_person", + &mixed_params(&[("$name", "OnBranchOne")], &[("$age", 41)]), + ) + .await + }); + in_envelope.wait_until_reached().await; + + let writer_b_db = std::sync::Arc::clone(&db); + let writer_b = tokio::spawn(async move { + writer_b_db + .mutate( + "b2", + MUTATION_QUERIES, + "insert_person", + &mixed_params(&[("$name", "OnBranchTwo")], &[("$age", 42)]), + ) + .await + }); + let b_result = tokio::time::timeout(std::time::Duration::from_secs(10), writer_b) + .await + .expect("B on a disjoint branch must complete while A holds its envelope parked") + .unwrap() + .expect("B's insert must publish"); + assert_eq!(b_result.affected_nodes, 1); + assert!( + !writer_a.is_finished(), + "A must still be parked inside its envelope while B published" + ); + + in_envelope.release(); + writer_a + .await + .unwrap() + .expect("A must publish normally after release"); + assert_eq!( + helpers::count_rows_branch(&db, "b1", "node:Person").await, + 5 + ); + assert_eq!( + helpers::count_rows_branch(&db, "b2", "node:Person").await, + 5 + ); +} + +/// Reads capture their catalog under a SHARED schema permit, so a read no +/// longer waits for a writer's publish hold (RFC +/// 2026-09-18-shared-schema-gate). Park a writer inside its envelope; a +/// read on a second handle must complete while the writer is parked — +/// under the former exclusive mutex this read would block until the +/// writer's guards dropped. +#[tokio::test(flavor = "multi_thread", worker_threads = 4)] +#[serial] +async fn read_capture_proceeds_while_writer_parked() { + let _scenario = FailScenario::setup(); + let dir = tempfile::tempdir().unwrap(); + let uri = dir.path().to_str().unwrap().to_string(); + let db = helpers::init_and_load(&dir).await; + drop(db); + + let db_a = std::sync::Arc::new(helpers::session(Omnigraph::open(&uri).await.unwrap())); + let db_b = std::sync::Arc::new(helpers::session(Omnigraph::open(&uri).await.unwrap())); + + let in_envelope = + helpers::failpoint::Rendezvous::park_first(&catalog::MUTATION_POST_FINALIZE_PRE_PUBLISHER); + let writer_db = std::sync::Arc::clone(&db_a); + let writer = tokio::spawn(async move { + writer_db + .mutate( + "main", + MUTATION_QUERIES, + "insert_person", + &mixed_params(&[("$name", "HoldsEnvelope")], &[("$age", 61)]), + ) + .await + }); + in_envelope.wait_until_reached().await; + + let rows = tokio::time::timeout( + std::time::Duration::from_secs(5), + count_rows(&db_b, "node:Person"), + ) + .await + .expect("a read must capture its catalog while the writer's envelope is held"); + assert_eq!(rows, 4, "the parked write must not be visible yet"); + + in_envelope.release(); + writer + .await + .unwrap() + .expect("the parked writer must publish after release"); + assert_eq!(count_rows(&db_b, "node:Person").await, 5); +} + // Atomic schema apply: schema apply writes staging files first, then commits // the manifest, then renames staging → final. Tests below inject crashes at // the two boundaries and assert that reopening the graph yields a consistent diff --git a/crates/omnigraph/tests/schema_apply.rs b/crates/omnigraph/tests/schema_apply.rs index 98d32301b..656b674b8 100644 --- a/crates/omnigraph/tests/schema_apply.rs +++ b/crates/omnigraph/tests/schema_apply.rs @@ -370,8 +370,10 @@ async fn mutation_waits_for_mid_apply_schema_gate_then_reprepares() { mutation_rv.release(); // Give the already-runnable mutation repeated scheduler turns. It must stay - // pending on the schema gate; completing here means it either advanced under - // an in-flight migration or returned a spurious post-prepare failure. + // pending on the schema gate — its SHARED permit parks behind the apply's + // held EXCLUSIVE permit (RFC 2026-09-18-shared-schema-gate); completing + // here means it either advanced under an in-flight migration or returned + // a spurious post-prepare failure. for _ in 0..128 { tokio::task::yield_now().await; if mutation_task.is_finished() { @@ -393,6 +395,64 @@ async fn mutation_waits_for_mid_apply_schema_gate_then_reprepares() { assert_eq!(count_rows(&db, "node:Person").await, 5); } +/// The reverse-direction pin for the shared/exclusive schema gate — the test +/// that catches a mis-classified writer: a writer holding its SHARED permit +/// (parked inside its envelope after detached commits, before publish) must +/// block a schema apply's EXCLUSIVE acquisition entirely, before the apply +/// creates its sentinel or touches any file. If a writer site were wrongly +/// left off the gate, the apply would proceed mid-write and this test reds. +#[cfg(feature = "failpoints")] +#[tokio::test(flavor = "multi_thread", worker_threads = 4)] +#[serial_test::serial] +async fn parked_writer_blocks_schema_apply() { + use omnigraph::seams::catalog; + + let dir = tempfile::tempdir().unwrap(); + let db = Arc::new(init_and_load(&dir).await); + let desired = TEST_SCHEMA.replace(" age: I32?\n}", " age: I32?\n motto: String?\n}"); + + let in_envelope = + helpers::failpoint::Rendezvous::park_first(&catalog::MUTATION_POST_FINALIZE_PRE_PUBLISHER); + let writer_db = Arc::clone(&db); + let writer = tokio::spawn(async move { + writer_db + .mutate( + "main", + MUTATION_QUERIES, + "insert_person", + &mixed_params(&[("$name", "gate-holder")], &[("$age", 27)]), + ) + .await + }); + in_envelope.wait_until_reached().await; + + let schema_db = Arc::clone(&db); + let schema_task = tokio::spawn(async move { schema_db.apply_schema(&desired).await }); + // The apply must park on the exclusive side behind the writer's shared + // permit: repeated scheduler turns, never finished. + for _ in 0..128 { + tokio::task::yield_now().await; + if schema_task.is_finished() { + break; + } + } + assert!( + !schema_task.is_finished(), + "schema apply must wait behind a writer's held shared schema permit", + ); + + in_envelope.release(); + writer + .await + .unwrap() + .expect("the parked writer must publish after release"); + schema_task + .await + .unwrap() + .expect("schema apply must complete once the writer's envelope releases"); + assert_eq!(count_rows(&db, "node:Person").await, 5); +} + /// ReadOnly opens participate in the process-local schema publication gate even /// though they perform no recovery writes. Park an open immediately before its /// source/IR/state read: schema apply must not reach staging until that coherent From 64bc2569187b6b59183ad9d33cfdaefc28da3148 Mon Sep 17 00:00:00 2001 From: Ragnor Comerford Date: Sat, 19 Sep 2026 16:05:09 +0200 Subject: [PATCH 04/13] docs: gate order, read-path claim, and the RFC's implementation record MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit writes.md names the shared permit in the finalization order and the promotion-after-release step; architecture.md's read sentence now states what is true (a capture takes a shared schema permit, no branch or table gate, and does not wait for writers — before this change it took the exclusive gate). The RFC flips to accepted/complete and its decision log records the implementation resolutions (optimize/cleanup shared; RO-open/reload exclusive as recorded conservatism; the tripwire and fairness pins; the DST arm with strict replay). --- docs/dev/architecture.md | 4 +++- docs/dev/writes.md | 12 +++++++++--- docs/rfcs/2026-09-18-shared-schema-gate.md | 18 +++++++++++------- docs/rfcs/README.md | 2 +- 4 files changed, 24 insertions(+), 12 deletions(-) diff --git a/docs/dev/architecture.md b/docs/dev/architecture.md index e8a8035da..252ce1f7c 100644 --- a/docs/dev/architecture.md +++ b/docs/dev/architecture.md @@ -127,7 +127,9 @@ gates for single-writer ownership. ## Concurrency and support boundary -- Reads are snapshot-isolated and do not take write gates. +- Reads are snapshot-isolated. A read-view capture takes a shared schema + permit (so it cannot observe a contract mid-swap) and no branch or table + gate; it does not wait for writers. - Write preparation may overlap. Durable effects are ordered by the shared schema, branch, and sorted-table gates, then fenced again by persisted authority and Lance transaction identity. diff --git a/docs/dev/writes.md b/docs/dev/writes.md index 0853151ba..8ed34286a 100644 --- a/docs/dev/writes.md +++ b/docs/dev/writes.md @@ -21,7 +21,7 @@ prepare logical change and validate it ↓ stage exact Lance transactions (no HEAD movement) ↓ -acquire schema → branch → sorted-table gates, recheck the complete authority +acquire shared schema permit → branch → sorted-table gates, recheck the complete authority ↓ commit each participant as a detached version of its pin ↓ @@ -99,12 +99,18 @@ history for collision, expected-version, and lineage validation. Finalization acquires the root-shared gate order: -1. schema; +1. the schema gate — a shared permit for ordinary writers (only a + contract-lifecycle pass such as schema apply or the system-column + upgrade takes it exclusively, so cross-branch writers do not serialize + on it; see + [RFC 2026-09-18-shared-schema-gate](../rfcs/2026-09-18-shared-schema-gate.md)); 2. target branch; 3. touched `(table identity, physical branch)` entries in deterministic order; 4. coordinator publication. -These gates order work inside one process. Correctness still depends on the +Promotion runs after the mutation, load, and ensure-indices writers release +this envelope; it needs no gate for correctness. These gates order work +inside one process. Correctness still depends on the persisted manifest precondition and the exact Lance transaction identity each pin records. A retryable pre-effect attempt discards all staged work, captures a new `WriteTxn`, and repeats boundedly; it never reuses batches against a new base. diff --git a/docs/rfcs/2026-09-18-shared-schema-gate.md b/docs/rfcs/2026-09-18-shared-schema-gate.md index 18ddcc167..4df0ce8d8 100644 --- a/docs/rfcs/2026-09-18-shared-schema-gate.md +++ b/docs/rfcs/2026-09-18-shared-schema-gate.md @@ -2,12 +2,12 @@ rfc: "2026-09-18-shared-schema-gate" title: "Shared schema gate and the write critical section" track: maintainer -status: draft -implementation: not-started +status: accepted +implementation: complete authors: - ragnorc created: 2026-09-18 -updated: 2026-09-18 +updated: 2026-09-19 discussion: null supersedes: [] superseded_by: [] @@ -259,13 +259,17 @@ one-commit revert. - Group commit (RFC 0067 path step 3) restructures the same critical section at the publisher; it composes with this change (it needs concurrent arrivals, which this change creates) and is a separate proposal. -- Whether optimize and cleanup should take the exclusive side out of caution - rather than the shared side. They hold every table gate they touch, so - they already serialize with writers at table grain; this RFC proposes - shared and records the question for review. ## Decision log - 2026-09-18 — Drafted against the detached-commit engine, with the concurrent-writes instrument's sequential cross-engine baselines as the motivating evidence. Blocked on RFC 0067 merging. +- 2026-09-19 — Implemented. Optimize and cleanup take the shared side (they + hold every table gate they touch, so they already serialize with writers + at table grain); read-only open and reload stay exclusive as recorded + conservatism. The mis-classification tripwire is + `parked_writer_blocks_schema_apply`; the plain-mode fairness pin is + `queued_schema_exclusive_blocks_later_shared`; the DST concurrent + universe gained the schema-apply-racing-writers arm with strict replay + asserted (`sched_escapes == 0`). diff --git a/docs/rfcs/README.md b/docs/rfcs/README.md index f406881ab..ce5df7e1a 100644 --- a/docs/rfcs/README.md +++ b/docs/rfcs/README.md @@ -218,7 +218,7 @@ then dated RFCs by date. | [2026-09-10](2026-09-10-server-lifecycle-and-online-deployment.md) | Server lifecycle and online deployment | maintainer | draft | not-started | | [2026-09-14](2026-09-14-compatibility-surfaces.md) | Compatibility surfaces | maintainer | draft | not-started | | [2026-09-16](2026-09-16-session-settings.md) | Session settings | maintainer | draft | in-progress | -| [2026-09-18](2026-09-18-shared-schema-gate.md) | Shared schema gate and the write critical section | maintainer | draft | not-started | +| [2026-09-18](2026-09-18-shared-schema-gate.md) | Shared schema gate and the write critical section | maintainer | accepted | complete | | [2026-09-21](2026-09-21-detached-only-tables.md) | Detached-only tables | maintainer | accepted | in-progress | | [2026-09-24](2026-09-24-shared-expression-model.md) | Shared expression model | maintainer | draft | in-progress | | [2026-09-26](2026-09-26-self-contained-server-testing.md) | Self-contained server testing with GQT and DST | maintainer | draft | not-started | From d74c0d47b505634f53c0e6fc76e91c694f24aa10 Mon Sep 17 00:00:00 2001 From: Ragnor Comerford Date: Fri, 25 Sep 2026 10:27:18 +0200 Subject: [PATCH 05/13] docs(release): writers no longer serialize on the schema gate (RFC 2026-09-18) --- docs/releases/v0.12.0.md | 10 ++++++++++ 1 file changed, 10 insertions(+) diff --git a/docs/releases/v0.12.0.md b/docs/releases/v0.12.0.md index 1c01c507a..9a1600e16 100644 --- a/docs/releases/v0.12.0.md +++ b/docs/releases/v0.12.0.md @@ -277,6 +277,16 @@ Unreleased. ## Compatibility and behavior changes +- **Writers stop serializing on the process-local schema gate (RFC + 2026-09-18, shared schema gate).** Mutations, loads, merges, index + maintenance, optimize, cleanup, repair, branch control and read-view + captures take a shared permit on the schema gate; schema apply, the + system-column upgrade, open, refresh and reload take it exclusive. A + writer on one branch no longer waits for a writer on another, and a read + no longer waits behind a writer's publish hold; same-branch writers still + serialize on the branch gate, and a schema apply still excludes every + writer for its whole pass. No flag; per-operation object-store request + counts are unchanged. - **Every table write is a detached Lance commit, published once and never promoted (RFC 0067 and RFC "Detached-only tables").** Each table effect is a detached commit of the pinned base, published as the pin From 2906ac77cc3426d20cc7034f58e781a78c877077 Mon Sep 17 00:00:00 2001 From: Ragnor Comerford Date: Fri, 25 Sep 2026 10:21:45 +0200 Subject: [PATCH 06/13] test(dst): concurrent-universe arm racing writers against a schema apply The coverage RFC 2026-09-18-shared-schema-gate commissions: the concurrent universe gains a schema actor that applies monotone additive schemas against main while the data writers run. dst_schema_apply_racing_writers_first_contact asserts every data write and every apply commits (neither side starves), each apply lands exactly one era commit, at least one seed interleaves writer commits (a non-vacuous concurrency claim), and under the seam scheduler the permits' turn/epoch protocol keeps the schedule seed-ordered (sched_escapes == 0, strict replay). The ignored hunt sweeps the interleaving space. --- crates/omnigraph-dst/src/concurrent.rs | 161 +++++++++++++++++++++-- crates/omnigraph-dst/src/lance_faults.rs | 4 +- crates/omnigraph-dst/tests/scenarios.rs | 133 +++++++++++++++++++ 3 files changed, 289 insertions(+), 9 deletions(-) diff --git a/crates/omnigraph-dst/src/concurrent.rs b/crates/omnigraph-dst/src/concurrent.rs index 3a37e86e3..975ffc8a5 100644 --- a/crates/omnigraph-dst/src/concurrent.rs +++ b/crates/omnigraph-dst/src/concurrent.rs @@ -52,7 +52,7 @@ use omnigraph::storage::{ObjectStorageAdapter, StorageAdapter}; use crate::fixtures::{ MUTATION_QUERIES, TEST_DATA, TEST_SCHEMA, mixed_params, mutate_on, person_rows_target, - query_main, + query_main, schema_with_extras, }; use crate::rand::SplitMix64; @@ -133,6 +133,17 @@ pub struct ConcurrentScenario { /// op's Lance-realm writes may land (declared, not a hole: that IS a /// torn state for recovery to judge). pub kill_writer: Option<(usize, usize)>, + /// ARM — schema apply as the EXCLUSIVE gate's writer role (RFC + /// 2026-09-18-shared-schema-gate): a dedicated schema actor (scheduler + /// id = `writers + 2`) performs this many monotone additive applies + /// (`schema_with_extras(1..=n)`) against main while the data writers + /// race under their SHARED permits. Seed triple drawn only when + /// nonzero, after every existing draw, so all existing scenarios keep + /// their exact draw sequences. Refused together with `branch_cycles` + /// (apply's mono-branch refusal would make every cycle a vacuous + /// conflict). This is the deterministic coverage of the + /// shared/exclusive boundary: writers-racing-apply had none before. + pub schema_ops: usize, /// ARM 3 — branch verbs under concurrency: a dedicated BRANCH ACTOR /// (writer id = `writers`, one past the data writers) runs this many /// fork→write→merge→delete cycles against main while the writers race — @@ -201,6 +212,9 @@ pub struct ConcurrentReport { pub branch_committed: usize, pub branch_merges: usize, pub branch_retries: usize, + /// Schema actor totals: applies committed / legal-conflict retries. + pub schema_committed: usize, + pub schema_retries: usize, /// Reader-actor rounds completed (their oracles red by panic en route). pub reader_rounds: usize, /// Marked storage faults actually delivered to writers (bite evidence @@ -1521,6 +1535,88 @@ fn writer_life( /// error surface: only `kind: Conflict` is legal; anything else panics /// naming the op — this arm's whole point is learning what a live peer's /// maintenance actually surfaces. +/// The schema actor (RFC 2026-09-18-shared-schema-gate): monotone +/// additive applies racing the data writers. Each apply takes the schema +/// gate's EXCLUSIVE side inside the engine, so it drains every writer's +/// shared permit and blocks new ones — the boundary this actor exists to +/// exercise deterministically. Legal outcomes en route: success, or a +/// typed `kind: Conflict` retry (a moved graph head between capture and +/// publish). Anything else reds naming the apply. +fn schema_life( + root: &str, + storage: Arc, + ops: usize, + seeds3: (u64, u64, u64), + start: Arc, + sched_ctx: Option<(Arc, usize)>, +) -> (usize, usize) { + let (runtime_seed, ulid_seed, _workload_seed) = seeds3; + let _ = rand::rng().reseed(); + let runtime = tokio::runtime::Builder::new_current_thread() + .enable_time() + .rng_seed(tokio::runtime::RngSeed::from_bytes( + &runtime_seed.to_le_bytes(), + )) + .build_local(Default::default()) + .expect("schema runtime"); + runtime.block_on(Box::pin(async move { + let _ids = omnigraph::dst_ids::IDS + .install(Arc::new(omnigraph::dst_ids::SeededUlids::new(ulid_seed))); + let _clock = omnigraph::dst_clock::CLOCK + .install(Arc::new(omnigraph::dst_clock::LogicalClock::default())); + let _gate = sched_ctx.as_ref().map(|(s, actor)| { + let (hook_sched, hook_actor) = (s.clone(), *actor); + omnigraph::dst_gate::GATE.install(Arc::new(omnigraph::dst_gate::TurnFn(move || { + hook_sched + .enter(hook_actor) + .map(|g| Box::new(g) as Box) + }))) + }); + let db = Session::from_defaults( + Arc::new( + Omnigraph::open_with_storage(root, storage) + .await + .expect("schema-actor handle on shared root"), + ), + SessionSettings::default(), + ); + let _finish = sched_ctx + .as_ref() + .map(|(s, actor)| s.finish_on_drop(*actor)); + start.wait(); + if let Some((s, _)) = &sched_ctx { + s.arm(); + } + let mut committed = 0usize; + let mut retries = 0usize; + for count in 1..=ops { + let desired = schema_with_extras(count); + let mut occ_retries = 0usize; + loop { + match Box::pin(db.apply_schema(&desired)).await { + Ok(_) => break, + Err(err) => { + let rendered = format!("{err:?}"); + assert!( + rendered.contains("kind: Conflict"), + "schema apply {count}: illegal rejection while racing \ + live writers: {rendered}" + ); + occ_retries += 1; + assert!( + occ_retries < 256, + "schema apply {count}: livelocked on legal conflicts" + ); + retries += 1; + } + } + } + committed += 1; + } + (committed, retries) + })) +} + fn maintenance_life( root: &str, storage: Arc, @@ -1871,6 +1967,14 @@ pub fn run_concurrent_universe(root: &str, sc: &ConcurrentScenario) -> Concurren } else { Vec::new() }; + assert!( + !(sc.schema_ops > 0 && sc.branch_cycles > 0), + "schema_ops races the mono-branch refusal: a branch actor makes every \ + apply a vacuous legal conflict — arm one or the other" + ); + // Drawn last so every existing scenario keeps its exact draw sequence. + let schema_seeds = + (sc.schema_ops > 0).then(|| (seeds.next_u64(), seeds.next_u64(), seeds.next_u64())); crate::harness::clear_process_slots(); crate::env_knobs::require_pool_env(); @@ -1929,9 +2033,13 @@ pub fn run_concurrent_universe(root: &str, sc: &ConcurrentScenario) -> Concurren })); // ---- the race: one OS thread per participant ---- + // Every actor that calls `start.wait()` must be counted: the + // barrier is cyclic, so an undercount lets the first N pass and + // parks the last arrival forever. let participants = sc.writers + usize::from(maintenance_seeds.is_some()) + usize::from(branch_seeds.is_some()) + + usize::from(schema_seeds.is_some()) + sc.readers; let start = Arc::new(std::sync::Barrier::new(participants)); // ARM 2: the dying writer gets its own kill wrapper over the @@ -1959,8 +2067,9 @@ pub fn run_concurrent_universe(root: &str, sc: &ConcurrentScenario) -> Concurren }) .collect(); // the universe's arbiter. Scheduler ids: writers 0..N, - // branch actor N (its claim id), maintenance N+1. Readers - // stay ungated (read-only; initial scope). + // branch actor N (its claim id), maintenance N+1, schema + // actor N+2. Readers stay ungated (read-only; initial + // scope). let scheduler: Option> = sc .seam_schedule .then(|| SeamScheduler::new(sc.seed ^ SEAM_SCHED_SALT)); @@ -1968,6 +2077,9 @@ pub fn run_concurrent_universe(root: &str, sc: &ConcurrentScenario) -> Concurren for w in 0..sc.writers { s.register(w); } + if schema_seeds.is_some() { + s.register(sc.writers + 2); + } if branch_seeds.is_some() { s.register(sc.writers); } @@ -1989,11 +2101,12 @@ pub fn run_concurrent_universe(root: &str, sc: &ConcurrentScenario) -> Concurren crate::lance_faults::set_seam_scheduler(scheduler.clone().map(|s| (s, sc.writers))); type BranchStats = (Vec, usize, usize); #[allow(clippy::type_complexity)] // scoped result tuple of the race - let (results, maintenance_stats, branch_stats, reader_rounds): ( + let (results, maintenance_stats, branch_stats, reader_rounds, schema_stats): ( Vec<(Vec, usize, usize)>, (usize, usize, usize), BranchStats, usize, + (usize, usize), ) = std::thread::scope(|writers| { let handles: Vec<_> = writer_seeds .iter() @@ -2064,6 +2177,26 @@ pub fn run_concurrent_universe(root: &str, sc: &ConcurrentScenario) -> Concurren }) .expect("spawn maintenance thread") }); + let schema_handle = schema_seeds.map(|seeds3| { + let schema_actor = sc.writers + 2; + let storage: Arc = match &scheduler { + Some(s) => Arc::new(ScheduledStorage::new( + storage.clone(), + s.clone(), + schema_actor, + )), + None => storage.clone(), + }; + let sched_ctx = scheduler.clone().map(|s| (s, schema_actor)); + let start = start.clone(); + std::thread::Builder::new() + .name("dst-schema-actor".into()) + .stack_size(crate::harness::UNIVERSE_STACK_BYTES) + .spawn_scoped(writers, move || { + schema_life(root, storage, sc.schema_ops, seeds3, start, sched_ctx) + }) + .expect("spawn schema actor thread") + }); let branch_handle = branch_seeds.map(|seeds3| { let storage: Arc = match &scheduler { Some(s) => Arc::new(ScheduledStorage::new( @@ -2114,6 +2247,13 @@ pub fn run_concurrent_universe(root: &str, sc: &ConcurrentScenario) -> Concurren Err(panic) => std::panic::resume_unwind(panic), } } + let schema_stats = match schema_handle { + None => (0usize, 0usize), + Some(h) => match h.join() { + Ok(s) => s, + Err(panic) => std::panic::resume_unwind(panic), + }, + }; let m = match maintenance_handle { None => (0, 0, 0), Some(h) => match h.join() { @@ -2135,7 +2275,7 @@ pub fn run_concurrent_universe(root: &str, sc: &ConcurrentScenario) -> Concurren Err(panic) => std::panic::resume_unwind(panic), } } - (all, m, b, r_total) + (all, m, b, r_total, schema_stats) }); // The race is over: clear the Lance-realm arbiter slot so // the final audit's reads run ungated (and never count as @@ -2143,6 +2283,7 @@ pub fn run_concurrent_universe(root: &str, sc: &ConcurrentScenario) -> Concurren crate::lance_faults::set_seam_scheduler(None); let (maintenance_committed, maintenance_retries, maintenance_cleanups) = maintenance_stats; + let (schema_committed, schema_retries) = schema_stats; let (branch_claims, branch_merges, branch_retries) = branch_stats; let branch_committed = branch_claims.len(); let recovery_reopens: usize = results.iter().map(|(_, r, _)| *r).sum(); @@ -2204,6 +2345,7 @@ pub fn run_concurrent_universe(root: &str, sc: &ConcurrentScenario) -> Concurren // always red. let maintenance_active = sc.maintenance_ops > 0; let branch_active = sc.branch_cycles > 0; + let schema_active = sc.schema_ops > 0; let mut below_horizon = 0usize; let mut islands = 0usize; let mut maintenance_commit_count = 0usize; @@ -2276,11 +2418,12 @@ pub fn run_concurrent_universe(root: &str, sc: &ConcurrentScenario) -> Concurren assert!( maintenance_active || recovery_legal - || branch_active, + || branch_active + || schema_active, "commit {id}: empty person-diff with \ no maintenance actor, no dying \ - writer, and no branch actor — \ - unattributable commit" + writer, no branch actor, and no \ + schema actor — unattributable commit" ); maintenance_commit_count += 1; } else { @@ -2441,6 +2584,8 @@ pub fn run_concurrent_universe(root: &str, sc: &ConcurrentScenario) -> Concurren maintenance_retries, maintenance_cleanups, maintenance_commits: maintenance_commit_count, + schema_committed, + schema_retries, below_horizon, branch_committed, branch_merges, diff --git a/crates/omnigraph-dst/src/lance_faults.rs b/crates/omnigraph-dst/src/lance_faults.rs index af9b22c6e..2eb2983d2 100644 --- a/crates/omnigraph-dst/src/lance_faults.rs +++ b/crates/omnigraph-dst/src/lance_faults.rs @@ -256,7 +256,8 @@ fn seam_scheduler() -> Option<(Arc, usize)> { /// Thread-name attribution: the actors' OS threads carry their /// identities — `dst-writer-N` (scheduler id N), `dst-branch-actor` (id = -/// writers), `dst-maintenance` (id = writers+1). Lance-realm calls executed +/// writers), `dst-maintenance` (id = writers+1), `dst-schema-actor` (id = +/// writers+2). Lance-realm calls executed /// INLINE on an actor's thread inherit its name and take turns; calls from /// Lance's own pool threads (lance-cpu, lance-io) carry other names and run /// UNGATED — the measured coverage gap (`note_unattributed`), never a @@ -270,6 +271,7 @@ fn actor_from_thread(writers: usize) -> Option { match name { "dst-branch-actor" => Some(writers), "dst-maintenance" => Some(writers + 1), + "dst-schema-actor" => Some(writers + 2), _ => None, } } diff --git a/crates/omnigraph-dst/tests/scenarios.rs b/crates/omnigraph-dst/tests/scenarios.rs index 963c33135..16ae0e64b 100644 --- a/crates/omnigraph-dst/tests/scenarios.rs +++ b/crates/omnigraph-dst/tests/scenarios.rs @@ -4581,6 +4581,7 @@ fn dst_concurrent_two_writers_first_contact() { writers: 2, ops_per_writer: 12, maintenance_ops: 0, + schema_ops: 0, kill_writer: None, branch_cycles: 0, readers: 0, @@ -4630,6 +4631,7 @@ fn dst_concurrent_contention_hunt() { writers: 4, ops_per_writer: 20, maintenance_ops: 0, + schema_ops: 0, kill_writer: None, branch_cycles: 0, readers: 0, @@ -4669,6 +4671,7 @@ fn dst_maintenance_actor_first_contact() { writers: 2, ops_per_writer: 12, maintenance_ops: 8, + schema_ops: 0, kill_writer: None, branch_cycles: 0, readers: 0, @@ -4698,6 +4701,128 @@ fn dst_maintenance_actor_first_contact() { } } +/// Schema apply racing live writers — the deterministic coverage the +/// shared/exclusive schema gate ships with (RFC 2026-09-18-shared-schema-gate; +/// before it, writers-racing-apply had NO DST coverage). A dedicated schema +/// actor performs monotone additive applies (each takes the gate's EXCLUSIVE +/// side, draining every writer's shared permit) while two data writers race +/// under shared permits. Oracles: every writer claim commits (no wedge under +/// schema contention), every apply commits (writers cannot starve the +/// exclusive side), each apply lands exactly one empty-person-diff era +/// commit, and — across the seed budget — the writers genuinely interleaved +/// (alternations ≥ 1 somewhere, or the green is vacuous). Plain mode only: +/// an apply's table rewrite runs on the single `lance-cpu` pool thread, +/// which the seam arbiter deliberately cannot see, so under the seam +/// scheduler its stall budget trips on a loaded machine (measured: 0, 4, +/// 12, 18 or 27 escapes across runs of one seed). The strict-replay claim +/// for this arm (`sched_escapes == 0`) is therefore the hunt's, run +/// explicitly on an idle machine; the permits' turn/epoch protocol itself +/// is pinned by `dst_seam_scheduler_bite_and_replay`. +#[test] +#[serial] +fn dst_schema_apply_racing_writers_first_contact() { + use omnigraph_dst::concurrent::{ConcurrentScenario, run_concurrent_universe}; + let mut interleaved_somewhere = false; + for seed in dst_seeds(&[24_301, 24_302, 24_303]) { + let sched = false; + let root = format!("shared-memory://dst-s24-schema-{seed}"); + let sc = ConcurrentScenario { + seed, + writers: 2, + ops_per_writer: 12, + maintenance_ops: 0, + schema_ops: 3, + kill_writer: None, + branch_cycles: 0, + readers: 1, + writer_fault_pct: 0, + seam_schedule: sched, + park_deleter_hold: false, + }; + let report = run_concurrent_universe(&root, &sc); + assert_eq!( + report.committed, 24, + "every data write must commit despite schema contention" + ); + assert_eq!( + report.schema_committed, 3, + "every schema apply must commit; writers cannot starve the exclusive side" + ); + assert_eq!( + report.maintenance_commits, 3, + "each apply lands exactly one empty-person-diff era commit" + ); + if sched { + assert_eq!( + report.sched_escapes, 0, + "the schema permits' turn/epoch protocol must keep the \ + interleaving seed-ordered (strict replay)" + ); + } + interleaved_somewhere |= report.alternations >= 1; + println!( + "dst s24 schema [seed={seed} sched={sched}]: committed={} occ_retries={} \ + schema(committed={} retries={}) alternations={} sched_escapes={}", + report.committed, + report.occ_retries, + report.schema_committed, + report.schema_retries, + report.alternations, + report.sched_escapes + ); + } + assert!( + interleaved_somewhere, + "no seed produced interleaved writer commits — a vacuous green for the \ + concurrency claim; widen the seed budget" + ); +} + +/// The schema-arm hunt instrument: wider seeds, scheduler on, faults on — +/// run explicitly when hunting interleavings around the shared/exclusive +/// boundary (`OMNIGRAPH_DST_SEEDS` widens the search). +#[test] +#[serial] +#[ignore = "hunt: schema-apply-vs-writers interleaving search — run explicitly"] +fn dst_schema_apply_racing_writers_hunt() { + use omnigraph_dst::concurrent::{ConcurrentScenario, run_concurrent_universe}; + // 24_304 is the first-contact scenario's own shape under the scheduler + // (see the pin for why strict replay is a hunt claim for this arm). + for seed in dst_seeds(&[ + 24_304, 24_310, 24_311, 24_312, 24_313, 24_314, 24_315, 24_316, 24_317, + ]) { + let root = format!("shared-memory://dst-s24-schema-hunt-{seed}"); + let sc = ConcurrentScenario { + seed, + writers: 3, + ops_per_writer: 10, + maintenance_ops: 0, + schema_ops: 4, + kill_writer: None, + branch_cycles: 0, + readers: 2, + writer_fault_pct: 10, + seam_schedule: true, + park_deleter_hold: false, + }; + let report = run_concurrent_universe(&root, &sc); + assert_eq!(report.committed, 30); + assert_eq!(report.schema_committed, 4); + assert_eq!(report.sched_escapes, 0, "strict replay must hold"); + println!( + "dst s24 schema-hunt [seed={seed}]: committed={} schema(committed={} \ + retries={}) faults={} alternations={} sched(turns={} escapes={})", + report.committed, + report.schema_committed, + report.schema_retries, + report.writer_faults_injected, + report.alternations, + report.sched_turns, + report.sched_escapes + ); + } +} + /// ARM 2 — crash one writer mid-op while the other keeps racing: /// writer 0's adapter-realm storage dies at its k-th write-class call /// (post-mortem refusal, no revive — the one-participant process-death @@ -4719,6 +4844,7 @@ fn dst_crash_one_writer_first_contact() { writers: 2, ops_per_writer: 12, maintenance_ops: 0, + schema_ops: 0, kill_writer: Some((0, kill_at)), branch_cycles: 0, readers: 0, @@ -4772,6 +4898,7 @@ fn dst_branch_actor_first_contact() { writers: 2, ops_per_writer: 12, maintenance_ops: 0, + schema_ops: 0, kill_writer: None, branch_cycles: 4, readers: 0, @@ -4821,6 +4948,7 @@ fn dst_concurrent_fleet() { writers: 3, ops_per_writer: 10, maintenance_ops: 0, + schema_ops: 0, kill_writer: None, branch_cycles: 0, // Readers in EVERY fleet arm — live differential reads during @@ -4842,6 +4970,7 @@ fn dst_concurrent_fleet() { ( "crash", ConcurrentScenario { + schema_ops: 0, kill_writer: Some((0, 7 + (seed as usize % 17))), ..base.clone() }, @@ -4915,6 +5044,7 @@ fn dst_seam_scheduler_bite_and_replay() { writers: 2, ops_per_writer: 8, maintenance_ops: 0, + schema_ops: 0, kill_writer: None, branch_cycles: 0, readers: 0, @@ -5016,6 +5146,7 @@ fn dst_optimize_races_branch_delete() { writers: 3, ops_per_writer: 10, maintenance_ops: 4, + schema_ops: 0, kill_writer: None, branch_cycles: 3, readers: 0, @@ -5084,6 +5215,7 @@ fn dst_optimize_races_branch_delete_seed_search() { writers: 2, ops_per_writer: 6, maintenance_ops: 0, + schema_ops: 0, kill_writer: None, branch_cycles: 3, readers: 0, @@ -5180,6 +5312,7 @@ fn dst_optimize_races_branch_delete_directed_hold() { writers: 2, ops_per_writer: 6, maintenance_ops: 0, + schema_ops: 0, kill_writer: None, branch_cycles: 3, readers: 0, From 90860e6434b30f08fe895127bdded0d147c1e09d Mon Sep 17 00:00:00 2001 From: Ragnor Comerford Date: Fri, 25 Sep 2026 15:01:33 +0200 Subject: [PATCH 07/13] engine: the write capture parks behind an in-flight schema apply The DST schema-apply-racing-writers arm found the write capture outside the gate: open_write_txn checked the durable sentinel before any permit, so a writer arriving on another handle while an apply was in flight was refused with a typed conflict and retried hot until the apply ended (256 immediate retries exhausted the arm's OCC budget on its first op), while a writer that had entered before the apply parked behind the exclusive side. The capture now probes the sentinel (schema_apply_sentinel_present, factored out of ensure_schema_apply_not_locked) and, when it stands, parks on the shared side until the apply releases, then recaptures under the promoted contract; a cross-process apply grants the permit at once and the bounded loop ends in the same typed refusal as before. The common path takes no permit: a capture-scoped shared permit was tried first and cost every reprepare attempt two arbiter turns under the DST seam scheduler (dst_seam_scheduler_bite_and_replay went from 0.8 s to 5.2 s, near its escape budget); with the park-on-refusal shape it runs in 0.6 s. Evidence: the arm's seeds commit every write and every apply; the mid-apply pin keeps its assertion (its mutation captures before the apply starts). --- crates/omnigraph/src/db/omnigraph.rs | 29 ++++++++++- .../src/db/omnigraph/schema_apply.rs | 9 +++- docs/dev/writes.md | 12 +++-- docs/rfcs/2026-09-18-shared-schema-gate.md | 52 ++++++++++++++++--- 4 files changed, 87 insertions(+), 15 deletions(-) diff --git a/crates/omnigraph/src/db/omnigraph.rs b/crates/omnigraph/src/db/omnigraph.rs index 0f79ebb27..e085f731c 100644 --- a/crates/omnigraph/src/db/omnigraph.rs +++ b/crates/omnigraph/src/db/omnigraph.rs @@ -1216,6 +1216,10 @@ impl Omnigraph { schema_apply::ensure_schema_apply_not_locked(self, operation).await } + pub(crate) async fn schema_apply_sentinel_present(&self) -> Result { + schema_apply::schema_apply_sentinel_present(self).await + } + /// Engine-facing trait surface around `TableStore`. /// /// This is the **only** accessor for engine code reaching into the @@ -1355,7 +1359,24 @@ impl Omnigraph { const MAX_CAPTURE_RETRIES: usize = 8; let branch = normalize_branch_name(branch.unwrap_or("main"))?; + let mut parked_behind_apply = false; for _ in 0..MAX_CAPTURE_RETRIES { + // A standing sentinel means a contract-lifecycle pass is in + // flight. An apply on this root holds the exclusive schema + // permit for its whole pass, so parking on the shared side + // waits it out; the writer then recaptures under the promoted + // contract instead of being refused and retrying hot (RFC + // 2026-09-18-shared-schema-gate). A cross-process apply grants + // the permit at once, so the bounded loop still ends in the + // sentinel refusal below. The common path takes no permit here: + // `commit_all` takes the writer's, and the gate is never held + // twice on one call path. + if self.schema_apply_sentinel_present().await? { + drop(self.write_queue().acquire_schema_shared().await); + parked_behind_apply = true; + tokio::task::yield_now().await; + continue; + } // A schema apply publishes graph_head before promoting its staged // contract. Read one fully validated IR/catalog, capture coherent // manifest authority, then re-read the durable schema marker (the @@ -1363,8 +1384,6 @@ impl Omnigraph { // only (old head, old schema) or (new head, new schema), never the // intermediate (new head, old schema) state, without paying for a // second full schema parse during capture. - self.ensure_schema_apply_not_locked("write preparation") - .await?; let (schema_ir, schema_state) = load_validated_schema_contract(self.uri(), Arc::clone(&self.storage)).await?; let (branch_identifier, graph_head, effective_graph_head, snapshot, manifest_probe) = @@ -1401,6 +1420,12 @@ impl Omnigraph { }); } + if parked_behind_apply { + // The sentinel outlived every park: a cross-process apply (or one + // that died holding it) — the same typed refusal as before. + self.ensure_schema_apply_not_locked("write preparation") + .await?; + } Err(OmniError::manifest_read_set_changed( format!("write_authority:{}", branch.as_deref().unwrap_or("main")), None, diff --git a/crates/omnigraph/src/db/omnigraph/schema_apply.rs b/crates/omnigraph/src/db/omnigraph/schema_apply.rs index 612cf7fbe..8a693ea08 100644 --- a/crates/omnigraph/src/db/omnigraph/schema_apply.rs +++ b/crates/omnigraph/src/db/omnigraph/schema_apply.rs @@ -1200,8 +1200,15 @@ pub(super) async fn release_schema_apply_lock(db: &Omnigraph) -> Result<()> { db.refresh_coordinator_only().await } +/// Whether a schema apply's durable sentinel stands on this graph: the +/// cross-handle and cross-process signal that a contract-lifecycle pass is +/// in flight. +pub(super) async fn schema_apply_sentinel_present(db: &Omnigraph) -> Result { + db.coordinator.read().await.schema_apply_locked().await +} + pub(super) async fn ensure_schema_apply_not_locked(db: &Omnigraph, operation: &str) -> Result<()> { - if db.coordinator.read().await.schema_apply_locked().await? { + if schema_apply_sentinel_present(db).await? { return Err(OmniError::manifest_conflict(format!( "{} is unavailable while schema apply is in progress", operation diff --git a/docs/dev/writes.md b/docs/dev/writes.md index 8ed34286a..6a2733707 100644 --- a/docs/dev/writes.md +++ b/docs/dev/writes.md @@ -97,6 +97,10 @@ handoff use the existing complete fold. This is disposable process memory; `__manifest` remains the only durable graph authority. Publication still scans history for collision, expected-version, and lineage validation. +The capture itself (`open_write_txn`) takes no permit on the common path: +when a schema-apply sentinel stands it parks on the shared side until the +apply releases, then recaptures under the promoted contract. + Finalization acquires the root-shared gate order: 1. the schema gate — a shared permit for ordinary writers (only a @@ -108,11 +112,9 @@ Finalization acquires the root-shared gate order: 3. touched `(table identity, physical branch)` entries in deterministic order; 4. coordinator publication. -Promotion runs after the mutation, load, and ensure-indices writers release -this envelope; it needs no gate for correctness. These gates order work -inside one process. Correctness still depends on the -persisted manifest precondition and the exact Lance transaction identity each -pin records. A retryable pre-effect attempt discards all staged work, captures a new +These gates order work inside one process. Correctness still depends on the persisted manifest +precondition and the exact Lance transaction identity each pin records. A +retryable pre-effect attempt discards all staged work, captures a new `WriteTxn`, and repeats boundedly; it never reuses batches against a new base. ## Writer adapters diff --git a/docs/rfcs/2026-09-18-shared-schema-gate.md b/docs/rfcs/2026-09-18-shared-schema-gate.md index 4df0ce8d8..6b981b4d0 100644 --- a/docs/rfcs/2026-09-18-shared-schema-gate.md +++ b/docs/rfcs/2026-09-18-shared-schema-gate.md @@ -37,7 +37,8 @@ The branch gate stays: same-branch writers still serialize from revalidation through publication, per RFC 0067's own recommendation. The per-table gates stay: they keep two same-table stagers from wasting one staging. Promotion — already correct with no gate held — moves after guard release as a separate, -independently revertible sub-decision. +independently revertible sub-decision (superseded: see Promotion outside the +guards). The net effect is the second step of RFC 0067's throughput path: cross-branch writers, independent merges, and reads stop serializing process-wide on one @@ -113,7 +114,10 @@ in-repo precedent for a shared/exclusive permit pair (`ExportCutPermit` / - `SchemaSharedPermit` — read side; taken by commit_all, the no-op conditional-mutation CAS, merge, ensure_indices and the full-text rebuild, optimize, cleanup, repair, branch create/create-from/delete, and the - read-view captures. + read-view captures. The write capture (`open_write_txn`) takes it only + while a schema-apply sentinel stands: it parks on the shared side until + the apply releases, then recaptures; the common path takes no permit + there. - `SchemaExclusivePermit` — write side; taken by schema apply, the system-column upgrade, and the contract-lifecycle passes on the handle: open, refresh, `settle_pending_schema_install`, `reload_schema_if_source_changed`, @@ -142,6 +146,13 @@ interleavings silently stop replaying; with it, the DST arbiter sees the same event vocabulary it sees today. ### Promotion outside the guards +> Superseded. The implementation measured this move as harmful (decision +> log, 2026-09-25: the successor writer promotes the same pin concurrently +> and the two replays serialize on Lance's commit path), and +> [Detached-only tables](2026-09-21-detached-only-tables.md) then removed +> promotion altogether, so there is nothing left to move. The +> `HeldWriteGates` envelope helper this section introduced stays. + `promote_held_all` runs today inside all three gates although its own contract states the write is already durable and graph-visible and a failure @@ -182,8 +193,8 @@ implicit in a mutex. - One coherent accepted view (invariant 3) is what the exclusive side protects: a contract-lifecycle pass still swaps the accepted view with no reader or writer in flight, because shared holders drain first. -- Recovery/pending-pin semantics (RFC 0067) are unchanged; promotion's - gate-free correctness is already the engine's documented contract. +- Recovery and pin semantics (RFC 0067, as amended by detached-only tables) + are unchanged. - The deny-list line "process-local locks presented as distributed writer fencing" is reaffirmed, not weakened: the gate remains an in-process contention structure; the durable `__schema_apply_lock__` sentinel, the @@ -250,7 +261,7 @@ The acceptance bar for the implementation PR, mapped to existing owners: ## Rollout One implementation PR after acceptance: lock + permits + classification + -promotion move + test re-derivations + the DST arm + doc updates, with the +test re-derivations + the DST arm + doc updates, with the before/after instrument runs in the PR body. No flag; reversibility is a one-commit revert. @@ -271,5 +282,32 @@ one-commit revert. conservatism. The mis-classification tripwire is `parked_writer_blocks_schema_apply`; the plain-mode fairness pin is `queued_schema_exclusive_blocks_later_shared`; the DST concurrent - universe gained the schema-apply-racing-writers arm with strict replay - asserted (`sched_escapes == 0`). + universe gained the schema-apply-racing-writers arm (strict replay: see + 2026-09-25). +- 2026-09-25 — Two findings from the DST arm's first runs, both now in the + implementation. (1) The write capture sat outside the gate: + `open_write_txn` checked the durable sentinel before any permit, so a + writer arriving on another handle while an apply was in flight was refused + with a typed conflict and retried hot until the apply ended, while one that + had entered before the apply parked. A capture-scoped shared permit fixed + that but cost every attempt two arbiter turns under the seam scheduler + (the scheduler pin went from 0.8 s to 5.2 s, near its escape budget); the + shipped shape parks on the shared side only when the sentinel probe is + positive, then recaptures — zero cost on the common path, and the arm's + seeds commit every write and every apply. (2) Promotion after guard + release was measured harmful: with the envelope released first, the next + same-branch writer finds the pin still pending and replays the same twin, + so two promoters serialize on Lance's commit path — an engine-internal + wait the seam arbiter cannot see (the scheduler pin escaped on every op, + bisected to that commit alone). It was reverted; the `HeldWriteGates` + helper stays. Strict replay for the schema arm is a + hunt claim, not a CI pin: an apply's table rewrite runs on the single + `lance-cpu` pool thread, invisible to the arbiter, so its stall budget + trips under load; the plain-mode seeds are the pin and + `dst_seam_scheduler_bite_and_replay` pins the permits' turn/epoch protocol. +- 2026-09-28 — Ported onto main after + [detached-only tables](2026-09-21-detached-only-tables.md) removed + promotion, which retires the promotion sub-decision outright, and after + the `omnigraph-core` extraction. The 22 acquisition sites, the classification + and the DST arm carry over unchanged; the write capture's sentinel probe + now uses the coordinator's `schema_apply_locked`. From e2685ff1158547c581d72c609bd44d35d4bac37a Mon Sep 17 00:00:00 2001 From: Ragnor Comerford Date: Mon, 28 Sep 2026 14:12:47 +0200 Subject: [PATCH 08/13] test(dst): the schema arm lands applies between writer commits MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Review follow-up. The first-contact pin's non-vacuity check counted only writer alternations, so it could not show that an apply ever contended with the writers — and none had: the schema actor applied back to back the moment the start barrier opened, and the write-preferring lock ran every apply before the first writer commit. The concurrent report now counts era commits that land between two writer commits (era_commits_between_data), the pin requires one somewhere across its seeds, and the schema actor pauses a seeded 2-31 ms before each apply so the applies spread across the writers' lifetime. Every apply of every seed now lands between writer commits. The hunt runs without readers: a reader's read-only open takes the schema gate's exclusive side with no arbiter hook, so its transitions fell outside the turns sched_escapes == 0 certifies. The pin's dead scheduler branch, left from when its scheduler seed moved to the hunt, is removed. --- crates/omnigraph-dst/src/concurrent.rs | 29 ++++++++++++++++++++- crates/omnigraph-dst/tests/scenarios.rs | 34 ++++++++++++++----------- 2 files changed, 47 insertions(+), 16 deletions(-) diff --git a/crates/omnigraph-dst/src/concurrent.rs b/crates/omnigraph-dst/src/concurrent.rs index 975ffc8a5..db16c2669 100644 --- a/crates/omnigraph-dst/src/concurrent.rs +++ b/crates/omnigraph-dst/src/concurrent.rs @@ -203,6 +203,11 @@ pub struct ConcurrentReport { /// maintenance actor, a dying/faulted writer's recovery pass, or a /// branch actor (none of these writes a writer-encoded person value). pub maintenance_commits: usize, + /// Empty person-diff commits (maintenance, schema applies) that landed + /// with a writer's data commit both before and after them in main's + /// lineage: evidence that such a commit interleaved with the writers + /// rather than running entirely before or after them. + pub era_commits_between_data: usize, /// Era commits legally unreadable at final audit because a concurrent /// Cleanup retired their versions (the retention horizon, live). The /// prefix-membership judge covers the claims that landed there. @@ -1550,7 +1555,7 @@ fn schema_life( start: Arc, sched_ctx: Option<(Arc, usize)>, ) -> (usize, usize) { - let (runtime_seed, ulid_seed, _workload_seed) = seeds3; + let (runtime_seed, ulid_seed, workload_seed) = seeds3; let _ = rand::rng().reseed(); let runtime = tokio::runtime::Builder::new_current_thread() .enable_time() @@ -1589,7 +1594,15 @@ fn schema_life( } let mut committed = 0usize; let mut retries = 0usize; + // Seeded pauses spread the applies across the writers' lifetime. A + // back-to-back burst queues the exclusive side the moment the barrier + // opens and, the lock being write-preferring, runs every apply before + // the first writer commit — no contention at all. REAL time, like the + // writers' think time: this universe holds no virtual clock. + let mut stream = SplitMix64(workload_seed); for count in 1..=ops { + let pause_ms = 2 + stream.next_u64() % 30; + tokio::time::sleep(std::time::Duration::from_millis(pause_ms)).await; let desired = schema_with_extras(count); let mut occ_retries = 0usize; loop { @@ -2349,6 +2362,9 @@ pub fn run_concurrent_universe(root: &str, sc: &ConcurrentScenario) -> Concurren let mut below_horizon = 0usize; let mut islands = 0usize; let mut maintenance_commit_count = 0usize; + // `true` for a data commit, `false` for an empty-diff + // commit, in lineage order. + let mut era_sequence: Vec = Vec::new(); let mut attributed: Vec = Vec::new(); let mut prev: Option> = None; let mut s0: Option> = None; @@ -2426,7 +2442,9 @@ pub fn run_concurrent_universe(root: &str, sc: &ConcurrentScenario) -> Concurren schema actor — unattributable commit" ); maintenance_commit_count += 1; + era_sequence.push(false); } else { + era_sequence.push(true); // Decode every change; ARM 3: a commit // whose changes ALL belong to the branch // actor is a MERGE commit folding a whole @@ -2584,6 +2602,15 @@ pub fn run_concurrent_universe(root: &str, sc: &ConcurrentScenario) -> Concurren maintenance_retries, maintenance_cleanups, maintenance_commits: maintenance_commit_count, + era_commits_between_data: era_sequence + .iter() + .enumerate() + .filter(|(i, data)| { + !**data + && era_sequence[..*i].iter().any(|d| *d) + && era_sequence[i + 1..].iter().any(|d| *d) + }) + .count(), schema_committed, schema_retries, below_horizon, diff --git a/crates/omnigraph-dst/tests/scenarios.rs b/crates/omnigraph-dst/tests/scenarios.rs index 16ae0e64b..093745aba 100644 --- a/crates/omnigraph-dst/tests/scenarios.rs +++ b/crates/omnigraph-dst/tests/scenarios.rs @@ -4710,7 +4710,9 @@ fn dst_maintenance_actor_first_contact() { /// schema contention), every apply commits (writers cannot starve the /// exclusive side), each apply lands exactly one empty-person-diff era /// commit, and — across the seed budget — the writers genuinely interleaved -/// (alternations ≥ 1 somewhere, or the green is vacuous). Plain mode only: +/// (`alternations` ≥ 1 somewhere) and at least one apply landed between two +/// writer commits (`era_commits_between_data` ≥ 1 somewhere); otherwise the +/// green is vacuous. Plain mode only: /// an apply's table rewrite runs on the single `lance-cpu` pool thread, /// which the seam arbiter deliberately cannot see, so under the seam /// scheduler its stall budget trips on a loaded machine (measured: 0, 4, @@ -4723,8 +4725,8 @@ fn dst_maintenance_actor_first_contact() { fn dst_schema_apply_racing_writers_first_contact() { use omnigraph_dst::concurrent::{ConcurrentScenario, run_concurrent_universe}; let mut interleaved_somewhere = false; + let mut apply_between_writes = false; for seed in dst_seeds(&[24_301, 24_302, 24_303]) { - let sched = false; let root = format!("shared-memory://dst-s24-schema-{seed}"); let sc = ConcurrentScenario { seed, @@ -4736,7 +4738,7 @@ fn dst_schema_apply_racing_writers_first_contact() { branch_cycles: 0, readers: 1, writer_fault_pct: 0, - seam_schedule: sched, + seam_schedule: false, park_deleter_hold: false, }; let report = run_concurrent_universe(&root, &sc); @@ -4752,23 +4754,17 @@ fn dst_schema_apply_racing_writers_first_contact() { report.maintenance_commits, 3, "each apply lands exactly one empty-person-diff era commit" ); - if sched { - assert_eq!( - report.sched_escapes, 0, - "the schema permits' turn/epoch protocol must keep the \ - interleaving seed-ordered (strict replay)" - ); - } interleaved_somewhere |= report.alternations >= 1; + apply_between_writes |= report.era_commits_between_data >= 1; println!( - "dst s24 schema [seed={seed} sched={sched}]: committed={} occ_retries={} \ - schema(committed={} retries={}) alternations={} sched_escapes={}", + "dst s24 schema [seed={seed}]: committed={} occ_retries={} \ + schema(committed={} retries={}) alternations={} applies_between_writes={}", report.committed, report.occ_retries, report.schema_committed, report.schema_retries, report.alternations, - report.sched_escapes + report.era_commits_between_data ); } assert!( @@ -4776,11 +4772,19 @@ fn dst_schema_apply_racing_writers_first_contact() { "no seed produced interleaved writer commits — a vacuous green for the \ concurrency claim; widen the seed budget" ); + assert!( + apply_between_writes, + "no seed landed a schema apply between two writer commits — the applies \ + never contended with the writers; widen the seed budget" + ); } /// The schema-arm hunt instrument: wider seeds, scheduler on, faults on — /// run explicitly when hunting interleavings around the shared/exclusive -/// boundary (`OMNIGRAPH_DST_SEEDS` widens the search). +/// boundary (`OMNIGRAPH_DST_SEEDS` widens the search). No readers: a reader +/// opens read-only handles, which take the schema gate's exclusive side with +/// no arbiter hook, so their gate transitions would fall outside the turns +/// that `sched_escapes == 0` certifies. #[test] #[serial] #[ignore = "hunt: schema-apply-vs-writers interleaving search — run explicitly"] @@ -4800,7 +4804,7 @@ fn dst_schema_apply_racing_writers_hunt() { schema_ops: 4, kill_writer: None, branch_cycles: 0, - readers: 2, + readers: 0, writer_fault_pct: 10, seam_schedule: true, park_deleter_hold: false, From 327cb69832f920590c43b1dc058aa9d6998c6510 Mon Sep 17 00:00:00 2001 From: Ragnor Comerford Date: Mon, 28 Sep 2026 14:12:47 +0200 Subject: [PATCH 09/13] docs: narrow the read-path claim; record the step 2 measurement Review follow-up. A read on the writer's own handle still waits while that handle publishes on its bound branch: the publish holds the coordinator lock across the manifest compare-and-swap and the read capture takes it. The architecture guide, the release note, the RFC's behavior section and a test comment now say that, and the remaining wait is listed under the RFC's unresolved questions. The RFC's decision log records the independent review and the step 2 measurement: committing detached effects before the gates is 13-15% of the gate hold and, because revalidation fails on any head move, would turn most same-branch attempts into dead detached commits, so it is not adopted. RFC 0067's plan text gains a dated amendment pointing there, and its step 3 no longer describes a promotion pass that detached-only tables removed. --- crates/omnigraph/tests/failpoints.rs | 4 +- docs/dev/architecture.md | 6 ++- docs/releases/v0.12.0.md | 7 +-- docs/rfcs/0067-detached-table-commits.md | 10 ++++- docs/rfcs/2026-09-18-shared-schema-gate.md | 51 ++++++++++++++++++++-- 5 files changed, 68 insertions(+), 10 deletions(-) diff --git a/crates/omnigraph/tests/failpoints.rs b/crates/omnigraph/tests/failpoints.rs index acea15fcc..2ad8c810e 100644 --- a/crates/omnigraph/tests/failpoints.rs +++ b/crates/omnigraph/tests/failpoints.rs @@ -1895,8 +1895,8 @@ async fn cross_branch_writers_overlap_inside_schema_gate() { ); } -/// Reads capture their catalog under a SHARED schema permit, so a read no -/// longer waits for a writer's publish hold (RFC +/// Reads capture their catalog under a SHARED schema permit, so a read on +/// another handle no longer waits for a writer's publish hold (RFC /// 2026-09-18-shared-schema-gate). Park a writer inside its envelope; a /// read on a second handle must complete while the writer is parked — /// under the former exclusive mutex this read would block until the diff --git a/docs/dev/architecture.md b/docs/dev/architecture.md index 252ce1f7c..e2b26d2b0 100644 --- a/docs/dev/architecture.md +++ b/docs/dev/architecture.md @@ -129,7 +129,11 @@ gates for single-writer ownership. - Reads are snapshot-isolated. A read-view capture takes a shared schema permit (so it cannot observe a contract mid-swap) and no branch or table - gate; it does not wait for writers. + gate, so it does not wait for writers on other handles. It does take its + handle's coordinator lock, which a publish on that handle's bound branch + holds for the manifest compare-and-swap: a read and a write sharing one + handle and branch (for example, the server's requests on `main`) still + wait for each other there. - Write preparation may overlap. Durable effects are ordered by the shared schema, branch, and sorted-table gates, then fenced again by persisted authority and Lance transaction identity. diff --git a/docs/releases/v0.12.0.md b/docs/releases/v0.12.0.md index 9a1600e16..642b28849 100644 --- a/docs/releases/v0.12.0.md +++ b/docs/releases/v0.12.0.md @@ -283,9 +283,10 @@ Unreleased. captures take a shared permit on the schema gate; schema apply, the system-column upgrade, open, refresh and reload take it exclusive. A writer on one branch no longer waits for a writer on another, and a read - no longer waits behind a writer's publish hold; same-branch writers still - serialize on the branch gate, and a schema apply still excludes every - writer for its whole pass. No flag; per-operation object-store request + no longer waits behind another handle's publish hold (a read and a write + on the same handle and branch still wait for each other while it + publishes); same-branch writers still serialize on the branch gate, and + a schema apply still excludes every writer for its whole pass. No flag; per-operation object-store request counts are unchanged. - **Every table write is a detached Lance commit, published once and never promoted (RFC 0067 and RFC "Detached-only tables").** Each table effect diff --git a/docs/rfcs/0067-detached-table-commits.md b/docs/rfcs/0067-detached-table-commits.md index 0438c0789..508a9b099 100644 --- a/docs/rfcs/0067-detached-table-commits.md +++ b/docs/rfcs/0067-detached-table-commits.md @@ -1202,6 +1202,13 @@ per write, which compaction bounds. process-wide. Evidence needed: the `writes.rs` concurrency matrix rerun with staging outside the gates, and the whole-run cost instrument showing the critical section shrank to revalidate-and-publish. + *Amended 2026-09-28:* the shared schema gate shipped + ([Shared schema gate](2026-09-18-shared-schema-gate.md)); committing the + detached effects before the gates was measured and not adopted. The + detached commit is 13–15% of the gate hold, and because revalidation + fails on any move of the branch head, under same-branch contention each + losing attempt would write a commit that is dead on arrival. That RFC's + decision log has the numbers and the levers they point to. 3. **Group commit at the branch publisher.** One publisher task per branch per process owns the CAS. Writers stage, then submit an entry (captured authority, staged pins, lineage row) and wait. While one CAS is in @@ -1217,7 +1224,8 @@ per write, which compaction bounds. table, and the head row moving from H to the last commit under the same row-level check as today; every entry gets the same durable outcome and is acknowledged only after the CAS returns; promotion runs afterwards in - one pass. Consequences a separate RFC must decide: time travel by + one pass (since removed: [detached-only tables](2026-09-21-detached-only-tables.md) + leaves no promotion to run). Consequences a separate RFC must decide: time travel by manifest version becomes time travel by batch, each commit row must carry the table pins it changed so the change feed stays per commit, and a snapshot at an interior commit of a batch is derivable but not diff --git a/docs/rfcs/2026-09-18-shared-schema-gate.md b/docs/rfcs/2026-09-18-shared-schema-gate.md index 6b981b4d0..76c79f374 100644 --- a/docs/rfcs/2026-09-18-shared-schema-gate.md +++ b/docs/rfcs/2026-09-18-shared-schema-gate.md @@ -91,9 +91,12 @@ No API, format, wire, or configuration change. Observable differences: - Cross-branch concurrent writers scale instead of serializing process-wide; same-branch writers keep today's ordering and conflict behavior. -- A read no longer waits for an unrelated writer's publish window to capture - its catalog view; it waits only for an in-flight contract-lifecycle pass, - as it must. +- A read no longer waits for another handle's publish window to capture its + catalog view; on the schema gate it waits only for an in-flight + contract-lifecycle pass, as it must. A read through the writer's own + handle still waits while that handle publishes on its bound branch: the + publish holds the handle's coordinator lock across the manifest + compare-and-swap, and the capture needs it (see Unresolved questions). - A schema apply or system-column upgrade still waits for every in-flight shared holder to drain, then excludes all of them — the same fairness as today's mutex, made explicit by a write-preferring lock: once the exclusive @@ -270,6 +273,13 @@ one-commit revert. - Group commit (RFC 0067 path step 3) restructures the same critical section at the publisher; it composes with this change (it needs concurrent arrivals, which this change creates) and is a separate proposal. +- Same-handle reads still wait on publication. A publish on a handle's bound + branch holds that handle's coordinator lock across the manifest + compare-and-swap (`commit_updates_on_branch_with_expected`), and a read + capture takes the same lock. The server shares one handle per graph, so its + reads on `main` still wait for its writes to `main`. Removing that wait means + publishing without holding the coordinator lock and installing the new view + afterwards; it is a separate change. ## Decision log @@ -311,3 +321,38 @@ one-commit revert. the `omnigraph-core` extraction. The 22 acquisition sites, the classification and the DST arm carry over unchanged; the write capture's sentinel probe now uses the coordinator's `schema_apply_locked`. +- 2026-09-28 — An independent review (Codex, `gpt-6-astra`) found three gaps, + each verified in the code. (1) The documentation overstated the read path: + a read on the writer's own handle still waits for that handle's coordinator + during a publish on its bound branch; the text now says so and the + remaining wait is listed under Unresolved questions. (2) The DST arm's + non-vacuity check counted only writer alternations; the harness now counts + applies that land between two writer commits, and requiring one exposed that + none ever had: the schema actor applied back to back the moment the start + barrier opened, and the write-preferring lock ran every apply before the + first writer commit. The actor now pauses a seeded 2–31 ms before each apply, + and every apply of every seed lands between writer commits. (3) The hunt ran + reader actors whose read-only opens take the exclusive side with no arbiter + hook, outside the turns `sched_escapes == 0` certifies; it runs without + readers now. +- 2026-09-28 — The second half of RFC 0067's step 2, committing the detached + table effects before taking the gates, was measured and not adopted. With + timing probes in the write path, a single writer on one branch holds the + gates for 7.3 ms locally (revalidation 0.6 ms, detached commit 1.1 ms, + publication 5.6 ms) and 670 ms at +30 ms per round trip (298, 85 and 284 ms): + the detached commit is 13–15% of the hold. Revalidation fails on any move of + the branch head, so under same-branch contention nearly every attempt that + waited for the gate loses: 6.6 failed revalidations per publication locally + with eight writers, and 2.9 at +30 ms, where each failure holds the gate for + about 300 ms before giving up (about 19 s of a 30 s window, more than the + successful writes used). Committing detached before the gates would make + each of those losers write a commit that is dead on arrival — roughly seven + times the table writes, most of them garbage for the collector — to save at + most 15% of the hold. The per-table gate plan in RFC 0067 would also invert + the lock order: every production path takes table gates inside the branch + gate, and the key includes the branch, so they add no exclusion today. The + levers the measurement points to are a cheap in-process check that fails a + stale attempt before revalidation's round trips, fewer round trips in + revalidation itself, and footprint-granular admission, which is RFC 0067's + group-commit rule. + From 8d6b2cdcb97c16867ad5716ab18387ffd1671ad6 Mon Sep 17 00:00:00 2001 From: Ragnor Comerford Date: Mon, 28 Sep 2026 14:40:20 +0200 Subject: [PATCH 10/13] docs(rfc): record the fail-fast prototype and the same-branch ceiling The in-process stale-attempt check caught every failed revalidation but gained only at object-store latency (0.40 to 0.53 commits/s with eight writers on one branch at +30 ms) and lost locally (25.5 to 23.2), and it left the starvation of #784 unchanged. The measurement bounds the whole family: with branch-wide revalidation, eight writers on one branch cannot exceed one writer alone, so raising same-branch throughput needs group commit's footprint admission rather than more work on the gate. --- docs/rfcs/2026-09-18-shared-schema-gate.md | 16 ++++++++++++++++ 1 file changed, 16 insertions(+) diff --git a/docs/rfcs/2026-09-18-shared-schema-gate.md b/docs/rfcs/2026-09-18-shared-schema-gate.md index 76c79f374..f6ed2a9af 100644 --- a/docs/rfcs/2026-09-18-shared-schema-gate.md +++ b/docs/rfcs/2026-09-18-shared-schema-gate.md @@ -355,4 +355,20 @@ one-commit revert. stale attempt before revalidation's round trips, fewer round trips in revalidation itself, and footprint-granular admission, which is RFC 0067's group-commit rule. +- 2026-09-28 — The first lever was prototyped and measured, not adopted. A + stale attempt was failed from the handle's in-memory view (same branch + incarnation, a strictly newer manifest version than the capture, a + different head) before revalidation's round trips; it caught every failed + revalidation in both settings. With eight writers on one branch it raised + throughput from 0.40 to 0.53 commits/s at +30 ms and cut manifest reads per + commit from 60 to 43, but lowered it from 25.5 to 23.2 locally and widened + the local p95 from 232 ms to 1.9 s: where revalidation is already cheap, + failing faster only sends losers back to re-prepare and re-queue sooner. + It also left the starvation of #784 unchanged. The measurement bounds this + whole family of fixes. Because any publication invalidates every other + attempt on the branch, eight writers on one branch cannot exceed one writer + alone (31.7 commits/s locally, 0.87 at +30 ms); this change already reaches + 80% and 46% of that. Raising the ceiling requires same-branch writes that + do not invalidate each other, which is the footprint admission of RFC + 0067's group commit (step 3), not further work on the gate. From 3d0532e49c9b6bbf4efea4c40edc91c435476dd6 Mon Sep 17 00:00:00 2001 From: Ragnor Comerford Date: Mon, 28 Sep 2026 16:29:20 +0200 Subject: [PATCH 11/13] docs(rfc): attribute the shared gate's gain; retract the one-branch claim A diagnostic build with one process-wide publication lock shows the whole eight-branch gain comes from overlapping publications across branches; interleaved reruns show no one-branch difference, so the earlier +36% was drift between sessions. --- docs/rfcs/2026-09-18-shared-schema-gate.md | 52 +++++++++++++++++----- 1 file changed, 42 insertions(+), 10 deletions(-) diff --git a/docs/rfcs/2026-09-18-shared-schema-gate.md b/docs/rfcs/2026-09-18-shared-schema-gate.md index f6ed2a9af..9b6d77297 100644 --- a/docs/rfcs/2026-09-18-shared-schema-gate.md +++ b/docs/rfcs/2026-09-18-shared-schema-gate.md @@ -361,14 +361,46 @@ one-commit revert. different head) before revalidation's round trips; it caught every failed revalidation in both settings. With eight writers on one branch it raised throughput from 0.40 to 0.53 commits/s at +30 ms and cut manifest reads per - commit from 60 to 43, but lowered it from 25.5 to 23.2 locally and widened - the local p95 from 232 ms to 1.9 s: where revalidation is already cheap, - failing faster only sends losers back to re-prepare and re-queue sooner. - It also left the starvation of #784 unchanged. The measurement bounds this - whole family of fixes. Because any publication invalidates every other - attempt on the branch, eight writers on one branch cannot exceed one writer - alone (31.7 commits/s locally, 0.87 at +30 ms); this change already reaches - 80% and 46% of that. Raising the ceiling requires same-branch writes that - do not invalidate each other, which is the footprint admission of RFC - 0067's group commit (step 3), not further work on the gate. + commit from 60 to 43. Locally it gave nothing: 23.2 commits/s sits inside + the 20–25 band that interleaved reruns later showed for every build (see + the attribution entry below). It also widened the local p95 from 232 ms to + 1.9 s, because where revalidation is already cheap, failing faster only + sends losers back to re-prepare and re-queue sooner. It left the + starvation of #784 unchanged. The measurement bounds this whole family of + fixes. Because any publication invalidates every other attempt on the + branch, eight writers on one branch cannot exceed one writer alone (31.7 + commits/s locally, 0.87 at +30 ms); eight writers reach about 65–80% and + 46% of that. Raising the ceiling requires same-branch writes that do not + invalidate each other, which is the footprint admission of RFC 0067's group + commit (step 3), not further work on the gate. +- 2026-09-28 — Attribution, after review asked for one effect at a time. + This change has two effects: + - read-view captures stop waiting behind another writer's publication; + - publications on different branches overlap. + + A diagnostic build separated them: this change plus one process-wide lock + around every mutation and load publication, so captures overlap + publications but publications still serialize. It was never merged. + + The three builds ran back to back in each cell, three runs per cell. The + comparison does not mix in the copy-on-write catalog publication (#781): + `main`'s side is `b14c22c5`, which is #781's own merge. + + | Cell | `main` | Diagnostic | This change | + |---|---|---|---| + | local, 8 branches | 69.4 | 68.6 | 154.0 | + | +30 ms, 8 branches | 0.50 | 0.43 | 2.13 | + + The whole eight-branch gain comes from overlapping publications; overlapping + captures adds nothing measurable. On one branch, locally, two interleaved + rounds gave, in commits/s: + - `main`: 24.4 and 20.3; + - diagnostic: 21.7 and 22.2; + - this change: 21.7 and 22.8. + + That is no difference. The earlier +36% (18.8 to 25.5) compared runs from + two sessions, and that cell drifts by about 20% between sessions: it is + dominated by re-prepares, with 116 to 178 exhausted re-prepare budgets per + cell in every build. The capture overlap stays a mechanism that + `read_capture_proceeds_while_writer_parked` pins, not a throughput claim. From bc09eec262dfd0e919858f09beacdcbc9c74ef50 Mon Sep 17 00:00:00 2001 From: azim afroozeh <13484327+azimafroozeh@users.noreply.github.com> Date: Mon, 28 Sep 2026 23:24:44 +0200 Subject: [PATCH 12/13] test --- ...ncurrent_reprepare_refresh_blocks_read.gqt | 36 +++++++++++++++++++ 1 file changed, 36 insertions(+) create mode 100644 crates/omnigraph-gqt/cases/concurrent_reprepare_refresh_blocks_read.gqt diff --git a/crates/omnigraph-gqt/cases/concurrent_reprepare_refresh_blocks_read.gqt b/crates/omnigraph-gqt/cases/concurrent_reprepare_refresh_blocks_read.gqt new file mode 100644 index 000000000..eb87c3fbb --- /dev/null +++ b/crates/omnigraph-gqt/cases/concurrent_reprepare_refresh_blocks_read.gqt @@ -0,0 +1,36 @@ +# issue: none +--- runner +timeout_ms: 10000 +environments: + - target: omnigraph-engine-dst + storage: in-memory-object-store + seeds: [0, 42] + +--- schema +node Person { + name: String @key +} +--- seed +{"type":"Person","data":{"name":"alice"}} +--- mutate +branch create feature +--- expect ok +--- concurrent +w1: query add_bob() { insert Person { name: "bob" } } +w2: query add_carol() { insert Person { name: "carol" } } +w3 on feature: query add_dave() { insert Person { name: "dave" } } +r1: query all() { match { $p: Person } return { $p.name } } +order: w2 park put nodes/, w3 park put feature/_versions/, w1, w2 put nodes/, r1 start, r1, w3 put feature/_versions/, w3, w2 +--- expect +w1: ok +w2: ok +w3: ok +r1: ok +--- query +query all() { match { $p: Person } return { $p.name } } +--- expect unordered +{"p.name":"alice"} +{"p.name":"bob"} +{"p.name":"carol"} +--- expect shape +p.name: String From ec28c2222c7be91314a935afbdd1dfddc1f8e360 Mon Sep 17 00:00:00 2001 From: azim afroozeh <13484327+azimafroozeh@users.noreply.github.com> Date: Tue, 29 Sep 2026 00:29:28 +0200 Subject: [PATCH 13/13] review --- crates/omnigraph-core/src/branch_control.rs | 9 + crates/omnigraph-dst/src/concurrent.rs | 15 +- crates/omnigraph-dst/tests/scenarios.rs | 52 ++- ...ncurrent_reprepare_refresh_blocks_read.gqt | 2 +- .../benches/scenarios/concurrent_writes.rs | 5 +- crates/omnigraph/src/db/omnigraph.rs | 78 ++-- .../src/db/omnigraph/schema_apply.rs | 14 +- .../src/db/omnigraph/system_column_upgrade.rs | 12 +- crates/omnigraph/src/db/write_queue.rs | 262 ++++++----- crates/omnigraph/src/exec/mutation.rs | 2 +- crates/omnigraph/src/exec/staging.rs | 5 +- crates/omnigraph/src/loader/mod.rs | 2 +- crates/omnigraph/src/seams/catalog.rs | 2 + crates/omnigraph/tests/failpoints.rs | 56 +++ crates/omnigraph/tests/schema_apply.rs | 32 +- docs/dev/architecture.md | 6 +- docs/dev/writes.md | 13 +- docs/releases/v0.12.0.md | 12 +- docs/rfcs/0067-detached-table-commits.md | 10 +- docs/rfcs/2026-09-18-shared-schema-gate.md | 425 +++++++++++++----- 20 files changed, 680 insertions(+), 334 deletions(-) diff --git a/crates/omnigraph-core/src/branch_control.rs b/crates/omnigraph-core/src/branch_control.rs index 6cb9ff97d..73841a93b 100644 --- a/crates/omnigraph-core/src/branch_control.rs +++ b/crates/omnigraph-core/src/branch_control.rs @@ -776,6 +776,14 @@ decide_seam! { pub static BRANCH_CREATE_POST_NATIVE = ("branch_create.post_native", BranchCreate, [Fail]); } +decide_seam! { + /// After the namespace inventory (refs listed, path collision checked, + /// a ref-less tree reclaimed), before the native create. No CAS covers + /// this window: the caller's exclusive schema permit is what keeps a + /// sibling create from passing its own inventory meanwhile. + pub static BRANCH_CREATE_POST_INVENTORY_PRE_NATIVE = ("branch_create.post_inventory_pre_native", BranchCreate, [Fail]); +} + /// Archived legacy ancestors still own their native path, even without refs. /// Flat generated names need no archive probe; slash names cost one per ancestor. async fn refuse_archived_path_ancestor(dataset: &Dataset, branch: &str) -> Result<()> { @@ -848,6 +856,7 @@ pub async fn create_branch_recoverably( { return Err(authority_appeared_after_absence(source, branch).await?); } + fail(&BRANCH_CREATE_POST_INVENTORY_PRE_NATIVE)?; for attempt in 0..2 { let native_error = match crate::lance_clone::create_branch(source, branch, source_version) diff --git a/crates/omnigraph-dst/src/concurrent.rs b/crates/omnigraph-dst/src/concurrent.rs index db16c2669..c0d5f73d8 100644 --- a/crates/omnigraph-dst/src/concurrent.rs +++ b/crates/omnigraph-dst/src/concurrent.rs @@ -1534,12 +1534,6 @@ fn writer_life( })) } -/// ARM 1 — the maintenance actor's life: races Optimize / Cleanup(keep=1) / -/// ensure_indices against the data writers from its own thread + handle. -/// Returns (committed, legal retries, cleanups run). STRICT first-contact -/// error surface: only `kind: Conflict` is legal; anything else panics -/// naming the op — this arm's whole point is learning what a live peer's -/// maintenance actually surfaces. /// The schema actor (RFC 2026-09-18-shared-schema-gate): monotone /// additive applies racing the data writers. Each apply takes the schema /// gate's EXCLUSIVE side inside the engine, so it drains every writer's @@ -1630,6 +1624,12 @@ fn schema_life( })) } +/// ARM 1 — the maintenance actor's life: races Optimize / Cleanup(keep=1) / +/// ensure_indices against the data writers from its own thread + handle. +/// Returns (committed, legal retries, cleanups run). STRICT first-contact +/// error surface: only `kind: Conflict` is legal; anything else panics +/// naming the op — this arm's whole point is learning what a live peer's +/// maintenance actually surfaces. fn maintenance_life( root: &str, storage: Arc, @@ -1952,12 +1952,13 @@ pub fn run_concurrent_universe(root: &str, sc: &ConcurrentScenario) -> Concurren // Birth certificate. NOTE the envelope: this line does NOT promise replay. println!( "dst concurrent universe [root={root} seed={} writers={} ops_per_writer={} \ - maintenance_ops={} kill_writer={:?} branch_cycles={}] \ + maintenance_ops={} schema_ops={} kill_writer={:?} branch_cycles={}] \ envelope=bite+oracles-hold (no replay claim)", sc.seed, sc.writers, sc.ops_per_writer, sc.maintenance_ops, + sc.schema_ops, sc.kill_writer, sc.branch_cycles ); diff --git a/crates/omnigraph-dst/tests/scenarios.rs b/crates/omnigraph-dst/tests/scenarios.rs index 093745aba..8cce39dde 100644 --- a/crates/omnigraph-dst/tests/scenarios.rs +++ b/crates/omnigraph-dst/tests/scenarios.rs @@ -4709,10 +4709,10 @@ fn dst_maintenance_actor_first_contact() { /// under shared permits. Oracles: every writer claim commits (no wedge under /// schema contention), every apply commits (writers cannot starve the /// exclusive side), each apply lands exactly one empty-person-diff era -/// commit, and — across the seed budget — the writers genuinely interleaved -/// (`alternations` ≥ 1 somewhere) and at least one apply landed between two -/// writer commits (`era_commits_between_data` ≥ 1 somewhere); otherwise the -/// green is vacuous. Plain mode only: +/// commit, and — in every seed — the writers genuinely interleaved +/// (`alternations` ≥ 1) and at least one apply landed between two writer +/// commits (`era_commits_between_data` ≥ 1); otherwise the green is +/// vacuous. Plain mode only: /// an apply's table rewrite runs on the single `lance-cpu` pool thread, /// which the seam arbiter deliberately cannot see, so under the seam /// scheduler its stall budget trips on a loaded machine (measured: 0, 4, @@ -4724,8 +4724,6 @@ fn dst_maintenance_actor_first_contact() { #[serial] fn dst_schema_apply_racing_writers_first_contact() { use omnigraph_dst::concurrent::{ConcurrentScenario, run_concurrent_universe}; - let mut interleaved_somewhere = false; - let mut apply_between_writes = false; for seed in dst_seeds(&[24_301, 24_302, 24_303]) { let root = format!("shared-memory://dst-s24-schema-{seed}"); let sc = ConcurrentScenario { @@ -4754,8 +4752,16 @@ fn dst_schema_apply_racing_writers_first_contact() { report.maintenance_commits, 3, "each apply lands exactly one empty-person-diff era commit" ); - interleaved_somewhere |= report.alternations >= 1; - apply_between_writes |= report.era_commits_between_data >= 1; + assert!( + report.alternations >= 1, + "seed {seed}: the writers never interleaved — a vacuous green for the \ + concurrency claim" + ); + assert!( + report.era_commits_between_data >= 1, + "seed {seed}: no schema apply landed between two writer commits — the \ + applies never contended with the writers" + ); println!( "dst s24 schema [seed={seed}]: committed={} occ_retries={} \ schema(committed={} retries={}) alternations={} applies_between_writes={}", @@ -4767,16 +4773,6 @@ fn dst_schema_apply_racing_writers_first_contact() { report.era_commits_between_data ); } - assert!( - interleaved_somewhere, - "no seed produced interleaved writer commits — a vacuous green for the \ - concurrency claim; widen the seed budget" - ); - assert!( - apply_between_writes, - "no seed landed a schema apply between two writer commits — the applies \ - never contended with the writers; widen the seed budget" - ); } /// The schema-arm hunt instrument: wider seeds, scheduler on, faults on — @@ -4962,8 +4958,26 @@ fn dst_concurrent_fleet() { seam_schedule: seam, park_deleter_hold: false, }; - let arms: [(&str, ConcurrentScenario); 6] = [ + let arms: [(&str, ConcurrentScenario); 8] = [ ("race", base.clone()), + // Schema arms: plain mode only (the first-contact pin says why). + ( + "schema", + ConcurrentScenario { + schema_ops: 3, + seam_schedule: false, + ..base.clone() + }, + ), + ( + "schema+maint", + ConcurrentScenario { + schema_ops: 3, + maintenance_ops: 4, + seam_schedule: false, + ..base.clone() + }, + ), ( "maint", ConcurrentScenario { diff --git a/crates/omnigraph-gqt/cases/concurrent_reprepare_refresh_blocks_read.gqt b/crates/omnigraph-gqt/cases/concurrent_reprepare_refresh_blocks_read.gqt index eb87c3fbb..1130e86f7 100644 --- a/crates/omnigraph-gqt/cases/concurrent_reprepare_refresh_blocks_read.gqt +++ b/crates/omnigraph-gqt/cases/concurrent_reprepare_refresh_blocks_read.gqt @@ -20,7 +20,7 @@ w1: query add_bob() { insert Person { name: "bob" } } w2: query add_carol() { insert Person { name: "carol" } } w3 on feature: query add_dave() { insert Person { name: "dave" } } r1: query all() { match { $p: Person } return { $p.name } } -order: w2 park put nodes/, w3 park put feature/_versions/, w1, w2 put nodes/, r1 start, r1, w3 put feature/_versions/, w3, w2 +order: w2 park put nodes/, w3 park put _versions/, w1, w2 put nodes/, r1 start, r1, w2, w3 put _versions/, w3 --- expect w1: ok w2: ok diff --git a/crates/omnigraph/benches/scenarios/concurrent_writes.rs b/crates/omnigraph/benches/scenarios/concurrent_writes.rs index d890a7773..eb9fd3694 100644 --- a/crates/omnigraph/benches/scenarios/concurrent_writes.rs +++ b/crates/omnigraph/benches/scenarios/concurrent_writes.rs @@ -14,8 +14,9 @@ //! //! Workload: insert-only `Chunk` rows with disjoint keys per worker //! (`cw-w{worker}-{seq}`), the shape that exercises the write path's real -//! serialization — the process-global write queue and the exclusive schema -//! gate every writer crosses in `commit_all` — without manufacturing key +//! serialization — the process-global write queue and the schema gate +//! every writer crosses in `commit_all` (shared since RFC +//! 2026-09-18-shared-schema-gate) — without manufacturing key //! conflicts. `Omnigraph::mutate` replays a typed read-set/authority //! conflict (`ReadSetChanged`: the graph head moved under a concurrent //! writer) itself, up to `MAX_PRE_EFFECT_REPREPARES` times for an diff --git a/crates/omnigraph/src/db/omnigraph.rs b/crates/omnigraph/src/db/omnigraph.rs index e085f731c..56f253466 100644 --- a/crates/omnigraph/src/db/omnigraph.rs +++ b/crates/omnigraph/src/db/omnigraph.rs @@ -62,11 +62,12 @@ use super::commit_graph::GraphCommit; use super::manifest::{GenesisManifestAttempt, ManifestChange, TableRegistration, TableTombstone}; use super::schema_state::{ SCHEMA_SOURCE_FILENAME, SchemaContractText, SchemaStagingPolicy, SchemaStateRecovery, - load_validated_schema_contract, load_validated_schema_contract_for_source, - read_accepted_schema_ir, read_schema_contract_text, read_schema_contract_text_for_source, - read_schema_state_identity, recover_schema_state_files, render_schema_contract, schema_ir_uri, - schema_source_staging_uri, schema_source_uri, schema_state_uri, validate_schema_contract, - validate_schema_contract_text, validate_schema_ir_against_snapshot, write_schema_contract, + StagedContract, inspect_staged_contract, load_validated_schema_contract, + load_validated_schema_contract_for_source, read_accepted_schema_ir, read_schema_contract_text, + read_schema_contract_text_for_source, read_schema_state_identity, recover_schema_state_files, + render_schema_contract, schema_ir_uri, schema_source_staging_uri, schema_source_uri, + schema_state_uri, validate_schema_contract, validate_schema_contract_text, + validate_schema_ir_against_snapshot, write_schema_contract, }; use super::snapshot::Snapshot; use super::{ @@ -831,7 +832,7 @@ impl Omnigraph { .await?; // A staged schema contract names the graph commit that publishes // it: install it when that commit is in lineage, discard it - // otherwise. The caller holds the shared schema gate. + // otherwise. The caller holds the exclusive schema permit. recover_schema_state_files( &root, Arc::clone(&storage), @@ -1359,24 +1360,29 @@ impl Omnigraph { const MAX_CAPTURE_RETRIES: usize = 8; let branch = normalize_branch_name(branch.unwrap_or("main"))?; - let mut parked_behind_apply = false; - for _ in 0..MAX_CAPTURE_RETRIES { - // A standing sentinel means a contract-lifecycle pass is in - // flight. An apply on this root holds the exclusive schema - // permit for its whole pass, so parking on the shared side - // waits it out; the writer then recaptures under the promoted - // contract instead of being refused and retrying hot (RFC - // 2026-09-18-shared-schema-gate). A cross-process apply grants - // the permit at once, so the bounded loop still ends in the - // sentinel refusal below. The common path takes no permit here: - // `commit_all` takes the writer's, and the gate is never held - // twice on one call path. + let mut captures = 0; + loop { + // A standing sentinel under a busy gate is an apply in this + // process: park on the shared side, then recapture under the + // promoted contract. Under a free gate it is another process's + // apply, or a dead one: the typed refusal, once a second listing + // confirms it (RFC 2026-09-18-shared-schema-gate). if self.schema_apply_sentinel_present().await? { - drop(self.write_queue().acquire_schema_shared().await); - parked_behind_apply = true; + match self.write_queue().try_acquire_schema_shared() { + Some(free) => { + drop(free); + self.ensure_schema_apply_not_locked("write preparation") + .await?; + } + None => drop(self.write_queue().acquire_schema_shared().await), + } tokio::task::yield_now().await; continue; } + captures += 1; + if captures > MAX_CAPTURE_RETRIES { + break; + } // A schema apply publishes graph_head before promoting its staged // contract. Read one fully validated IR/catalog, capture coherent // manifest authority, then re-read the durable schema marker (the @@ -1420,12 +1426,6 @@ impl Omnigraph { }); } - if parked_behind_apply { - // The sentinel outlived every park: a cross-process apply (or one - // that died holding it) — the same typed refusal as before. - self.ensure_schema_apply_not_locked("write preparation") - .await?; - } Err(OmniError::manifest_read_set_changed( format!("write_authority:{}", branch.as_deref().unwrap_or("main")), None, @@ -2199,6 +2199,28 @@ impl Omnigraph { Ok(()) } + /// The reprepare refresh: a write whose authority moved recaptures from + /// the coordinator alone, and takes the contract-lifecycle pass of + /// [`refresh`](Self::refresh) only when a published staging is waiting + /// to be installed. `open_write_txn` reads the accepted contract from + /// the store on every capture, so the schema view and the `PromoteOnly` + /// pass add nothing to an ordinary reprepare; taking their exclusive + /// permit there made every same-branch reprepare a process-wide barrier + /// (RFC 2026-09-18-shared-schema-gate, 2026-09-29 entry). + pub(crate) async fn refresh_for_reprepare(&self) -> Result<()> { + self.refresh_coordinator_only().await?; + if matches!( + inspect_staged_contract(&self.root_uri, self.storage.as_ref(), false).await?, + StagedContract::Marked { + published: true, + .. + } + ) { + self.refresh().await?; + } + Ok(()) + } + pub async fn resolve_snapshot(&self, branch: &str) -> Result { self.ensure_schema_state_valid().await?; self.coordinator @@ -3026,7 +3048,7 @@ impl Omnigraph { let source = self.active_branch().await; self.settle_pending_schema_install().await?; fail(&BRANCH_CONTROL_PRE_GATES)?; - let _schema_permit = self.write_queue().acquire_schema_shared().await; + let _schema_permit = self.write_queue().acquire_schema_exclusive().await; let _branch_guards = self .write_queue() .acquire_branches(&[source.clone(), Some(target.clone())]) @@ -3109,7 +3131,7 @@ impl Omnigraph { self.ensure_schema_state_valid().await?; self.settle_pending_schema_install().await?; fail(&BRANCH_CONTROL_PRE_GATES)?; - let _schema_permit = self.write_queue().acquire_schema_shared().await; + let _schema_permit = self.write_queue().acquire_schema_exclusive().await; let _branch_guards = self .write_queue() .acquire_branches(&[branch.clone(), Some(target_branch.clone())]) diff --git a/crates/omnigraph/src/db/omnigraph/schema_apply.rs b/crates/omnigraph/src/db/omnigraph/schema_apply.rs index 8a693ea08..407686e7a 100644 --- a/crates/omnigraph/src/db/omnigraph/schema_apply.rs +++ b/crates/omnigraph/src/db/omnigraph/schema_apply.rs @@ -159,6 +159,13 @@ decide_seam! { pub static SCHEMA_APPLY_POST_LOCK_PRE_EFFECT = ("schema_apply.post_lock_pre_effect", Unreachable, [Fail]); } +decide_seam! { + /// Right after the durable sentinel lands, under the exclusive schema + /// permit and before any planning: the first crossing proves the apply + /// got past every shared holder. + pub static SCHEMA_APPLY_POST_SENTINEL = ("schema_apply.post_sentinel", Unreachable, [Fail]); +} + async fn plan_schema_for_apply_from_accepted( db: &Omnigraph, desired_schema_source: &str, @@ -265,8 +272,11 @@ where // 2026-09-18-shared-schema-gate). let _schema_gate = db.write_queue().acquire_schema_exclusive().await; acquire_schema_apply_lock(db).await?; - let result = - apply_schema_with_lock(db, desired_schema_source, options, actor, validate_catalog).await; + let result = async { + fail(&SCHEMA_APPLY_POST_SENTINEL)?; + apply_schema_with_lock(db, desired_schema_source, options, actor, validate_catalog).await + } + .await; let release_result = release_schema_apply_lock(db).await; if release_result.is_err() { // Liveness: the next write entry on this handle retries the release diff --git a/crates/omnigraph/src/db/omnigraph/system_column_upgrade.rs b/crates/omnigraph/src/db/omnigraph/system_column_upgrade.rs index 24af25133..738cf2255 100644 --- a/crates/omnigraph/src/db/omnigraph/system_column_upgrade.rs +++ b/crates/omnigraph/src/db/omnigraph/system_column_upgrade.rs @@ -175,7 +175,17 @@ pub(super) async fn upgrade_system_columns( if !options.check { db.settle_pending_schema_install().await?; } - let _schema_gate = db.write_queue().acquire_schema_exclusive().await; + let (_shared_gate, _exclusive_gate): ( + Option, + Option, + ) = if options.check { + (Some(db.write_queue().acquire_schema_shared().await), None) + } else { + ( + None, + Some(db.write_queue().acquire_schema_exclusive().await), + ) + }; db.refresh_coordinator_only().await?; let stamp = crate::db::manifest::internal_schema_stamp_at(db.uri(), None) .await? diff --git a/crates/omnigraph/src/db/write_queue.rs b/crates/omnigraph/src/db/write_queue.rs index e62263293..89e0d9e6b 100644 --- a/crates/omnigraph/src/db/write_queue.rs +++ b/crates/omnigraph/src/db/write_queue.rs @@ -148,10 +148,13 @@ pub(crate) type TableQueueKey = (String, Option); /// The graph-global schema gate: the one shared/exclusive slot. /// /// It serializes every graph-global schema writer (schema apply and the -/// system-column upgrade) and every pass that installs, discards, or -/// republishes the accepted schema-contract view against each other — -/// exclusively — while readers of the accepted view (ordinary writers, -/// merges, maintenance, branch control, read-view captures) share. +/// system-column upgrade), every pass that installs, discards, or +/// republishes the accepted schema-contract view, and every pass that +/// changes the live branch-ref set without a CAS over it (branch create +/// and create-from, whose namespace inventory precedes the native create) +/// against each other — exclusively — while readers of the accepted view +/// (ordinary writers, merges, maintenance, branch delete, read-view +/// captures) share. /// /// The gate is non-reentrant in BOTH modes on one task: an exclusive /// holder re-acquiring either side self-deadlocks exactly like the @@ -179,6 +182,18 @@ pub(crate) struct SchemaGateSlot { releases: std::sync::atomic::AtomicU64, } +impl SchemaGateSlot { + /// The release protocol shared by both permits, the same as + /// `QueueGuard`'s: turn first, release inside it, bump the epoch last so + /// an observed bump implies a genuinely released slot. + fn release(&self, guard: Option) { + let _turn = crate::dst_gate::turn(); + drop(guard); + self.releases + .fetch_add(1, std::sync::atomic::Ordering::SeqCst); + } +} + /// A shared schema permit: proof that no contract-lifecycle pass is /// concurrently swapping the accepted schema/catalog view. Held by /// readers of that view for the duration of their gate-ordered work @@ -190,8 +205,8 @@ pub(crate) struct SchemaSharedPermit { } /// An exclusive schema permit: sole ownership of the accepted-view -/// transition. Held by schema apply, the system-column upgrade, and the -/// contract install/discard/reload passes. +/// transition. Held by schema apply, the system-column upgrade, the +/// contract install/discard/reload passes, and branch create. #[must_use = "dropping the permit releases the exclusive schema gate"] pub(crate) struct SchemaExclusivePermit { inner: Option>, @@ -200,54 +215,34 @@ pub(crate) struct SchemaExclusivePermit { impl Drop for SchemaSharedPermit { fn drop(&mut self) { - // Same protocol as `QueueGuard`: turn first, release inside it, - // bump the epoch last so an observed bump implies a genuinely - // released reader slot. - let _turn = crate::dst_gate::turn(); - drop(self.inner.take()); - self.slot - .releases - .fetch_add(1, std::sync::atomic::Ordering::SeqCst); + self.slot.release(self.inner.take()); } } impl Drop for SchemaExclusivePermit { fn drop(&mut self) { - let _turn = crate::dst_gate::turn(); - drop(self.inner.take()); - self.slot - .releases - .fetch_add(1, std::sync::atomic::Ordering::SeqCst); + self.slot.release(self.inner.take()); } } -/// Scheduled shared acquisition of the schema gate; the exact -/// [`scheduled_lock`] protocol on the read side. Uninstalled: plain -/// blocking `read_owned` (fair, write-preferring). Installed: stay in the -/// turn loop, `try_read_owned` only when the release epoch moved, yield -/// every iteration. -async fn scheduled_schema_shared(slot: Arc) -> SchemaSharedPermit { +/// One side of the schema gate on the [`scheduled_lock`] protocol: plain +/// blocking acquire uninstalled; installed, try only when the release +/// epoch moved, yielding every turn. +async fn scheduled_schema_permit( + slot: &SchemaGateSlot, + acquire: impl AsyncFnOnce(Arc>) -> G, + try_acquire: impl Fn(Arc>) -> Option, +) -> G { let mut wait_epoch: Option = None; loop { match crate::dst_gate::turn() { - None => { - let guard = Arc::clone(&slot.lock).read_owned().await; - return SchemaSharedPermit { - inner: Some(guard), - slot, - }; - } + None => return acquire(Arc::clone(&slot.lock)).await, Some(_turn) => { let epoch_now = slot.releases.load(std::sync::atomic::Ordering::SeqCst); if wait_epoch.is_none_or(|e| epoch_now != e) { - match Arc::clone(&slot.lock).try_read_owned() { - Ok(guard) => { - return SchemaSharedPermit { - inner: Some(guard), - slot: Arc::clone(&slot), - }; - } - Err(_) => wait_epoch = Some(epoch_now), + match try_acquire(Arc::clone(&slot.lock)) { + Some(guard) => return guard, + None => wait_epoch = Some(epoch_now), } } // else: no release since the failed attempt — a no-op turn. @@ -257,36 +252,29 @@ async fn scheduled_schema_shared(slot: Arc) -> SchemaSharedPermi } } -/// Scheduled exclusive acquisition of the schema gate; the write-side -/// twin of [`scheduled_schema_shared`]. +async fn scheduled_schema_shared(slot: Arc) -> SchemaSharedPermit { + let guard = scheduled_schema_permit( + &slot, + async |lock| lock.read_owned().await, + |lock| lock.try_read_owned().ok(), + ) + .await; + SchemaSharedPermit { + inner: Some(guard), + slot, + } +} + async fn scheduled_schema_exclusive(slot: Arc) -> SchemaExclusivePermit { - let mut wait_epoch: Option = None; - loop { - match crate::dst_gate::turn() { - None => { - let guard = Arc::clone(&slot.lock).write_owned().await; - return SchemaExclusivePermit { - inner: Some(guard), - slot, - }; - } - Some(_turn) => { - let epoch_now = slot.releases.load(std::sync::atomic::Ordering::SeqCst); - if wait_epoch.is_none_or(|e| epoch_now != e) { - match Arc::clone(&slot.lock).try_write_owned() { - Ok(guard) => { - return SchemaExclusivePermit { - inner: Some(guard), - slot: Arc::clone(&slot), - }; - } - Err(_) => wait_epoch = Some(epoch_now), - } - } - // else: no release since the failed attempt — a no-op turn. - } - } - tokio::task::yield_now().await; + let guard = scheduled_schema_permit( + &slot, + async |lock| lock.write_owned().await, + |lock| lock.try_write_owned().ok(), + ) + .await; + SchemaExclusivePermit { + inner: Some(guard), + slot, } } @@ -301,17 +289,23 @@ async fn scheduled_schema_exclusive(slot: Arc) -> SchemaExclusiv #[must_use = "dropping the gates releases the write envelope"] pub(crate) struct HeldWriteGates { _schema: SchemaSharedPermit, - /// `[0]` is the branch gate, followed by the lex-sorted table gates. - _queue: Vec, + _branch: QueueGuard, + _tables: Vec, } impl HeldWriteGates { /// Assemble the envelope from the gates in acquisition order: the - /// shared schema permit, then the branch gate and sorted table gates. - pub(crate) fn new(schema: SchemaSharedPermit, queue: Vec) -> Self { + /// shared schema permit, then the branch gate, then the sorted table + /// gates. + pub(crate) fn new( + schema: SchemaSharedPermit, + branch: QueueGuard, + tables: Vec, + ) -> Self { Self { _schema: schema, - _queue: queue, + _branch: branch, + _tables: tables, } } } @@ -432,6 +426,18 @@ impl WriteQueueManager { scheduled_schema_shared(Arc::clone(&self.schema_gate)).await } + /// The shared side without waiting: `None` while an exclusive permit is + /// held or queued in this process. A standing schema-apply sentinel with + /// a free gate is therefore another process's apply (or a dead one). + pub(crate) fn try_acquire_schema_shared(&self) -> Option { + let _turn = crate::dst_gate::turn(); + let guard = Arc::clone(&self.schema_gate.lock).try_read_owned().ok()?; + Some(SchemaSharedPermit { + inner: Some(guard), + slot: Arc::clone(&self.schema_gate), + }) + } + /// Take the schema gate's exclusive side: sole ownership of the /// accepted-view transition. Shared holders drain first; new shared /// permits queue behind this acquisition in plain mode. Same @@ -590,30 +596,31 @@ mod tests { assert!(second.try_acquire_export_cut().is_some()); } + /// Shared permits are concurrent; each side excludes the other. The + /// contending acquisitions run on their own tasks, as engine callers do. #[tokio::test] - async fn schema_shared_permits_are_concurrent() { + async fn schema_exclusive_excludes_shared_both_directions() { let qm = Arc::new(WriteQueueManager::new()); + let first = qm.acquire_schema_shared().await; let qm2 = Arc::clone(&qm); - let second = timeout(Duration::from_secs(2), async move { - qm2.acquire_schema_shared().await - }) + let second = timeout( + Duration::from_secs(2), + tokio::spawn(async move { qm2.acquire_schema_shared().await }), + ) .await - .expect("a second shared permit must not wait behind the first"); + .expect("a second shared permit must not wait behind the first") + .unwrap(); drop(first); drop(second); - } - - #[tokio::test] - async fn schema_exclusive_excludes_shared_both_directions() { - let qm = Arc::new(WriteQueueManager::new()); // Held exclusive blocks a shared acquire. let exclusive = qm.acquire_schema_exclusive().await; let qm2 = Arc::clone(&qm); - let blocked = timeout(Duration::from_millis(200), async move { - qm2.acquire_schema_shared().await - }) + let blocked = timeout( + Duration::from_millis(200), + tokio::spawn(async move { qm2.acquire_schema_shared().await }), + ) .await; assert!( blocked.is_err(), @@ -621,17 +628,20 @@ mod tests { ); drop(exclusive); let qm2 = Arc::clone(&qm); - let shared = timeout(Duration::from_secs(2), async move { - qm2.acquire_schema_shared().await - }) + let shared = timeout( + Duration::from_secs(2), + tokio::spawn(async move { qm2.acquire_schema_shared().await }), + ) .await - .expect("shared must acquire once the exclusive permit releases"); + .expect("shared must acquire once the exclusive permit releases") + .unwrap(); // Held shared blocks an exclusive acquire. let qm2 = Arc::clone(&qm); - let blocked = timeout(Duration::from_millis(200), async move { - qm2.acquire_schema_exclusive().await - }) + let blocked = timeout( + Duration::from_millis(200), + tokio::spawn(async move { qm2.acquire_schema_exclusive().await }), + ) .await; assert!( blocked.is_err(), @@ -639,19 +649,18 @@ mod tests { ); drop(shared); let qm2 = Arc::clone(&qm); - let _exclusive = timeout(Duration::from_secs(2), async move { - qm2.acquire_schema_exclusive().await - }) + let _exclusive = timeout( + Duration::from_secs(2), + tokio::spawn(async move { qm2.acquire_schema_exclusive().await }), + ) .await - .expect("exclusive must acquire once the shared permit releases"); + .expect("exclusive must acquire once the shared permit releases") + .unwrap(); } - /// The plain-mode no-starvation pin: tokio's `RwLock` is - /// write-preferring, so once an exclusive acquisition is queued, a - /// LATER shared acquisition parks behind it instead of overtaking. If - /// a tokio upgrade ever changes that policy, this test reds and the - /// schema gate's fairness claim (RFC 2026-09-18-shared-schema-gate) - /// must be re-derived. + /// The plain-mode no-starvation pin: a shared acquisition after a queued + /// exclusive parks behind it (tokio's `RwLock` is write-preferring). A + /// tokio change of policy reds this and the RFC's fairness claim. #[tokio::test] async fn queued_schema_exclusive_blocks_later_shared() { let qm = Arc::new(WriteQueueManager::new()); @@ -659,13 +668,21 @@ mod tests { let qm_writer = Arc::clone(&qm); let writer = tokio::spawn(async move { qm_writer.acquire_schema_exclusive().await }); - // Give the exclusive acquisition time to enter the waiter queue. - tokio::time::sleep(Duration::from_millis(100)).await; + let qm_probe = Arc::clone(&qm); + timeout(Duration::from_secs(2), async move { + while let Some(permit) = qm_probe.try_acquire_schema_shared() { + drop(permit); + tokio::task::yield_now().await; + } + }) + .await + .expect("the exclusive acquisition must enter the waiter queue"); let qm2 = Arc::clone(&qm); - let overtaking = timeout(Duration::from_millis(200), async move { - qm2.acquire_schema_shared().await - }) + let overtaking = timeout( + Duration::from_millis(200), + tokio::spawn(async move { qm2.acquire_schema_shared().await }), + ) .await; assert!( overtaking.is_err(), @@ -679,11 +696,13 @@ mod tests { .expect("writer task must not panic"); drop(exclusive); let qm2 = Arc::clone(&qm); - let _shared = timeout(Duration::from_secs(2), async move { - qm2.acquire_schema_shared().await - }) + let _shared = timeout( + Duration::from_secs(2), + tokio::spawn(async move { qm2.acquire_schema_shared().await }), + ) .await - .expect("the parked shared permit must acquire after the exclusive releases"); + .expect("the parked shared permit must acquire after the exclusive releases") + .unwrap(); } #[tokio::test] @@ -694,20 +713,23 @@ mod tests { let exclusive = first.acquire_schema_exclusive().await; let second2 = Arc::clone(&second); - let blocked = timeout(Duration::from_millis(200), async move { - second2.acquire_schema_shared().await - }) + let blocked = timeout( + Duration::from_millis(200), + tokio::spawn(async move { second2.acquire_schema_shared().await }), + ) .await; assert!( blocked.is_err(), "handles for one root must exclude on one schema gate" ); drop(exclusive); - let _shared = timeout(Duration::from_secs(2), async move { - second.acquire_schema_shared().await - }) + let _shared = timeout( + Duration::from_secs(2), + tokio::spawn(async move { second.acquire_schema_shared().await }), + ) .await - .expect("the second handle's shared permit must acquire after release"); + .expect("the second handle's shared permit must acquire after release") + .unwrap(); } #[tokio::test] diff --git a/crates/omnigraph/src/exec/mutation.rs b/crates/omnigraph/src/exec/mutation.rs index 15d4f5987..bc3ded42c 100644 --- a/crates/omnigraph/src/exec/mutation.rs +++ b/crates/omnigraph/src/exec/mutation.rs @@ -881,7 +881,7 @@ impl Omnigraph { "prepared mutation authority changed before effects; repreparing" ); crate::instrumentation::record_mutation_reprepare(); - self.refresh().await?; + self.refresh_for_reprepare().await?; } result => return result, } diff --git a/crates/omnigraph/src/exec/staging.rs b/crates/omnigraph/src/exec/staging.rs index 290546bf8..1eba08fc4 100644 --- a/crates/omnigraph/src/exec/staging.rs +++ b/crates/omnigraph/src/exec/staging.rs @@ -889,9 +889,8 @@ impl StagedMutation { // per-table gates. Hold the full set through manifest publish. let schema = db.write_queue().acquire_schema_shared().await; let branch_guard = db.write_queue().acquire_branch(branch).await; - let mut queue = vec![branch_guard]; - queue.extend(db.write_queue().acquire_many(&queue_keys).await); - let gates = crate::db::write_queue::HeldWriteGates::new(schema, queue); + let table_guards = db.write_queue().acquire_many(&queue_keys).await; + let gates = crate::db::write_queue::HeldWriteGates::new(schema, branch_guard, table_guards); // Re-capture manifest pins under the queue (PR 2 / MR-686). // diff --git a/crates/omnigraph/src/loader/mod.rs b/crates/omnigraph/src/loader/mod.rs index 6b12c31f9..fd26c7a9b 100644 --- a/crates/omnigraph/src/loader/mod.rs +++ b/crates/omnigraph/src/loader/mod.rs @@ -548,7 +548,7 @@ async fn load_jsonl_data( branch = branch.unwrap_or("main"), "prepared load authority changed before effects; repreparing" ); - db.refresh().await?; + db.refresh_for_reprepare().await?; } result => return result, } diff --git a/crates/omnigraph/src/seams/catalog.rs b/crates/omnigraph/src/seams/catalog.rs index 0de5327b3..28a264039 100644 --- a/crates/omnigraph/src/seams/catalog.rs +++ b/crates/omnigraph/src/seams/catalog.rs @@ -7,6 +7,7 @@ use omnigraph_seams::DecideSeam; omnigraph_seams::catalog! { crate::blob::BLOB_READ_POST_CAPTURE, + crate::branch_control::BRANCH_CREATE_POST_INVENTORY_PRE_NATIVE, crate::branch_control::BRANCH_CREATE_POST_NATIVE, crate::branch_control::BRANCH_DELETE_POST_NATIVE, crate::branch_control::BRANCH_DELETE_POST_ARCHIVE, @@ -52,6 +53,7 @@ omnigraph_seams::catalog! { crate::db::omnigraph::schema_apply::SCHEMA_APPLY_AFTER_STAGING_WRITE, crate::db::omnigraph::schema_apply::SCHEMA_APPLY_BEFORE_STAGING_WRITE, crate::db::omnigraph::schema_apply::SCHEMA_APPLY_POST_LOCK_PRE_EFFECT, + crate::db::omnigraph::schema_apply::SCHEMA_APPLY_POST_SENTINEL, crate::db::omnigraph::schema_apply::SCHEMA_APPLY_POST_TABLE_COMMIT, crate::db::omnigraph::table_ops::ENSURE_INDICES_POST_PHASE_B_PRE_MANIFEST_COMMIT, crate::db::omnigraph::table_ops::ENSURE_INDICES_POST_STAGE_PRE_COMMIT_BTREE, diff --git a/crates/omnigraph/tests/failpoints.rs b/crates/omnigraph/tests/failpoints.rs index 2ad8c810e..0f1b65e5b 100644 --- a/crates/omnigraph/tests/failpoints.rs +++ b/crates/omnigraph/tests/failpoints.rs @@ -1895,6 +1895,62 @@ async fn cross_branch_writers_overlap_inside_schema_gate() { ); } +/// Branch create takes the schema gate's EXCLUSIVE side: no CAS covers its +/// namespace inventory, so a sibling create with disjoint branch gates must +/// wait at the gate, then refuse on the collision the first leaves behind. +#[tokio::test(flavor = "multi_thread", worker_threads = 4)] +#[serial] +async fn sibling_branch_creates_exclude_at_the_schema_gate() { + let _scenario = FailScenario::setup(); + let dir = tempfile::tempdir().unwrap(); + let db = helpers::init_and_load(&dir).await; + db.branch_create("b").await.unwrap(); + let db = std::sync::Arc::new(db); + + let after_inventory = helpers::failpoint::Rendezvous::park_first( + &catalog::BRANCH_CREATE_POST_INVENTORY_PRE_NATIVE, + ); + let first_db = std::sync::Arc::clone(&db); + let first = tokio::spawn(async move { first_db.branch_create("feature").await }); + after_inventory.wait_until_reached().await; + + let second_db = std::sync::Arc::clone(&db); + let mut second = tokio::spawn(async move { + second_db + .branch_create_from(ReadTarget::branch("b"), "feature/x") + .await + }); + let overtaking = tokio::time::timeout(std::time::Duration::from_secs(1), &mut second).await; + assert!( + overtaking.is_err(), + "a sibling create must wait at the schema gate while the first create sits \ + between its inventory and its native create" + ); + + after_inventory.release(); + first + .await + .unwrap() + .expect("the first create must land after release"); + let err = second + .await + .unwrap() + .expect_err("the sibling create must refuse on the collision the first create left"); + assert!( + err.to_string().contains("shares its physical Lance path"), + "unexpected refusal: {err}" + ); + let branches = db.branch_list().await.unwrap(); + assert!( + branches.iter().any(|name| name == "feature"), + "{branches:?}" + ); + assert!( + !branches.iter().any(|name| name == "feature/x"), + "{branches:?}" + ); +} + /// Reads capture their catalog under a SHARED schema permit, so a read on /// another handle no longer waits for a writer's publish hold (RFC /// 2026-09-18-shared-schema-gate). Park a writer inside its envelope; a diff --git a/crates/omnigraph/tests/schema_apply.rs b/crates/omnigraph/tests/schema_apply.rs index 656b674b8..6bb99ea74 100644 --- a/crates/omnigraph/tests/schema_apply.rs +++ b/crates/omnigraph/tests/schema_apply.rs @@ -395,12 +395,9 @@ async fn mutation_waits_for_mid_apply_schema_gate_then_reprepares() { assert_eq!(count_rows(&db, "node:Person").await, 5); } -/// The reverse-direction pin for the shared/exclusive schema gate — the test -/// that catches a mis-classified writer: a writer holding its SHARED permit -/// (parked inside its envelope after detached commits, before publish) must -/// block a schema apply's EXCLUSIVE acquisition entirely, before the apply -/// creates its sentinel or touches any file. If a writer site were wrongly -/// left off the gate, the apply would proceed mid-write and this test reds. +/// The mis-classified-writer tripwire: a writer parked inside its envelope +/// holds its SHARED permit, so the apply must not reach the seam after its +/// sentinel. Falsified by dropping the permit in `HeldWriteGates::new`: red. #[cfg(feature = "failpoints")] #[tokio::test(flavor = "multi_thread", worker_threads = 4)] #[serial_test::serial] @@ -426,26 +423,31 @@ async fn parked_writer_blocks_schema_apply() { }); in_envelope.wait_until_reached().await; + let post_sentinel = + helpers::failpoint::Rendezvous::park_first(&catalog::SCHEMA_APPLY_POST_SENTINEL); let schema_db = Arc::clone(&db); let schema_task = tokio::spawn(async move { schema_db.apply_schema(&desired).await }); - // The apply must park on the exclusive side behind the writer's shared - // permit: repeated scheduler turns, never finished. - for _ in 0..128 { - tokio::task::yield_now().await; - if schema_task.is_finished() { - break; + // Wall time: an apply past the gate makes store requests before the seam. + let crossed = tokio::time::timeout(std::time::Duration::from_secs(1), async { + while !post_sentinel.reached() { + tokio::time::sleep(std::time::Duration::from_millis(5)).await; } - } + }) + .await; assert!( - !schema_task.is_finished(), - "schema apply must wait behind a writer's held shared schema permit", + crossed.is_err(), + "schema apply must wait behind a writer's held shared schema permit; it \ + created its sentinel while the writer was parked", ); + assert!(!schema_task.is_finished()); in_envelope.release(); writer .await .unwrap() .expect("the parked writer must publish after release"); + post_sentinel.wait_until_reached().await; + post_sentinel.release(); schema_task .await .unwrap() diff --git a/docs/dev/architecture.md b/docs/dev/architecture.md index e2b26d2b0..92673f4d9 100644 --- a/docs/dev/architecture.md +++ b/docs/dev/architecture.md @@ -129,7 +129,11 @@ gates for single-writer ownership. - Reads are snapshot-isolated. A read-view capture takes a shared schema permit (so it cannot observe a contract mid-swap) and no branch or table - gate, so it does not wait for writers on other handles. It does take its + gate, so it does not wait for writers on other handles. It does wait for + an exclusive pass on any handle of the process (schema apply, the + system-column upgrade, open, refresh, settle, reload, `sync_branch`, + branch create), and, because the gate is write-preferring, from the + moment one is queued. It also takes its handle's coordinator lock, which a publish on that handle's bound branch holds for the manifest compare-and-swap: a read and a write sharing one handle and branch (for example, the server's requests on `main`) still diff --git a/docs/dev/writes.md b/docs/dev/writes.md index 6a2733707..400c630d9 100644 --- a/docs/dev/writes.md +++ b/docs/dev/writes.md @@ -98,15 +98,20 @@ handoff use the existing complete fold. This is disposable process memory; history for collision, expected-version, and lineage validation. The capture itself (`open_write_txn`) takes no permit on the common path: -when a schema-apply sentinel stands it parks on the shared side until the -apply releases, then recaptures under the promoted contract. +when a schema-apply sentinel stands and an apply in this process holds the +exclusive permit, it parks on the shared side until the apply releases, then +recaptures under the promoted contract. A sentinel under a free gate is +another process's apply and gets the typed refusal at once. Only `mutate` +and `load` reach the park: merge, index maintenance, optimize, cleanup and +repair call `ensure_schema_apply_idle` before their capture, so a standing +sentinel refuses them one call earlier. Finalization acquires the root-shared gate order: 1. the schema gate — a shared permit for ordinary writers (only a contract-lifecycle pass such as schema apply or the system-column - upgrade takes it exclusively, so cross-branch writers do not serialize - on it; see + upgrade, and branch create, whose namespace inventory no CAS covers, + take it exclusively, so cross-branch writers do not serialize on it; see [RFC 2026-09-18-shared-schema-gate](../rfcs/2026-09-18-shared-schema-gate.md)); 2. target branch; 3. touched `(table identity, physical branch)` entries in deterministic order; diff --git a/docs/releases/v0.12.0.md b/docs/releases/v0.12.0.md index 642b28849..c9fe0fc21 100644 --- a/docs/releases/v0.12.0.md +++ b/docs/releases/v0.12.0.md @@ -279,15 +279,17 @@ Unreleased. - **Writers stop serializing on the process-local schema gate (RFC 2026-09-18, shared schema gate).** Mutations, loads, merges, index - maintenance, optimize, cleanup, repair, branch control and read-view - captures take a shared permit on the schema gate; schema apply, the - system-column upgrade, open, refresh and reload take it exclusive. A + maintenance, optimize, cleanup, repair, branch delete, the system-column + upgrade's `--check` and read-view captures take a shared permit on the + schema gate; schema apply, the system-column upgrade, open, refresh, + settle, reload, `sync_branch` and branch create take it exclusive. A writer on one branch no longer waits for a writer on another, and a read no longer waits behind another handle's publish hold (a read and a write on the same handle and branch still wait for each other while it publishes); same-branch writers still serialize on the branch gate, and - a schema apply still excludes every writer for its whole pass. No flag; per-operation object-store request - counts are unchanged. + a schema apply still excludes every writer for its whole pass. No flag; + per-operation object-store request counts on the success path are + unchanged. - **Every table write is a detached Lance commit, published once and never promoted (RFC 0067 and RFC "Detached-only tables").** Each table effect is a detached commit of the pinned base, published as the pin diff --git a/docs/rfcs/0067-detached-table-commits.md b/docs/rfcs/0067-detached-table-commits.md index 508a9b099..5a71df2de 100644 --- a/docs/rfcs/0067-detached-table-commits.md +++ b/docs/rfcs/0067-detached-table-commits.md @@ -7,7 +7,7 @@ implementation: complete authors: - ragnorc created: 2026-09-14 -updated: 2026-09-25 +updated: 2026-09-29 discussion: null supersedes: [] superseded_by: [] @@ -1258,8 +1258,12 @@ claims). serializing process-wide (#643); nothing in this RFC depends on the gates for correctness. Proposed in [Shared schema gate and the write critical section](2026-09-18-shared-schema-gate.md). - Decided by the engine maintainers in the change that - closes #643. + Decided 2026-09-19 for the schema-gate half: the gate is shared for + writers, merges, maintenance, branch delete and read captures, and + exclusive for contract-lifecycle passes and branch create (that RFC, + implemented). The merge case of + [#643](https://github.com/ModernRelay/omnigraph/issues/643), two merges + overlapping, is not evidenced by that change and stays open here. - Whether merge promotion ever uses `Restore` of the chain tip instead of per-chunk replay, and above which chain length. Per-chunk replay is the only shipped path. A `Restore` makes the row stamps of restored rows diff --git a/docs/rfcs/2026-09-18-shared-schema-gate.md b/docs/rfcs/2026-09-18-shared-schema-gate.md index 9b6d77297..44e66a8e3 100644 --- a/docs/rfcs/2026-09-18-shared-schema-gate.md +++ b/docs/rfcs/2026-09-18-shared-schema-gate.md @@ -7,12 +7,11 @@ implementation: complete authors: - ragnorc created: 2026-09-18 -updated: 2026-09-19 +updated: 2026-09-29 discussion: null supersedes: [] superseded_by: [] -blocked_on: - - "RFC 0067 (detached table commits) merging: the safety argument below is stated against the detached write path and does not hold on the sidecar-era engine." +blocked_on: [] --- # RFC: Shared schema gate and the write critical section @@ -22,25 +21,31 @@ blocked_on: ## Summary -Resolve [RFC 0067](0067-detached-table-commits.md)'s open question (#643): the -process-local ***schema gate*** — the write-queue key +Resolve the schema-gate half of [RFC 0067](0067-detached-table-commits.md)'s +open question ([#643](https://github.com/ModernRelay/omnigraph/issues/643)): +the process-local ***schema gate***, the write-queue key `("__schema_apply__", None)` that every writer, every maintenance pass, every -branch control, and every read-view capture takes exclusively today — becomes a -shared/exclusive lock. A ***contract-lifecycle pass*** (schema apply, the -system-column upgrade, and every pass that installs, discards, or republishes -the accepted schema-contract view: open, refresh, settle, reload, branch sync) -takes it exclusive, exactly as it effectively does today. Everything else — -ordinary mutations and loads, merges, index maintenance, optimize, cleanup, -repair, branch control, and read-view captures — takes a ***shared permit***. +branch control, and every read-view capture took exclusively before this +change, becomes a shared/exclusive lock. A ***contract-lifecycle pass*** +(schema apply, the system-column upgrade, and every pass that installs, +discards, or republishes the accepted schema-contract view: open, refresh, +settle, reload, branch sync) takes it exclusive, exactly as it effectively did +before. Branch create and create-from take it exclusive too: their namespace +inventory changes the live branch-ref set with no CAS over the change (Design, +The lock). Everything else (ordinary mutations and loads, merges, index +maintenance, optimize, cleanup, repair, branch delete, and read-view captures) +takes a ***shared permit***. The merge case of that question, two merges +overlapping, has no evidence in this change (Evidence and tests). The branch gate stays: same-branch writers still serialize from revalidation through publication, per RFC 0067's own recommendation. The per-table gates -stay: they keep two same-table stagers from wasting one staging. Promotion — -already correct with no gate held — moves after guard release as a separate, -independently revertible sub-decision (superseded: see Promotion outside the -guards). +stay in the acquisition order, but they add no exclusion today: every +production path takes them inside the branch gate, and their key includes the +branch (decision log, 2026-09-28). -The net effect is the second step of RFC 0067's throughput path: cross-branch +The net effect is the schema-gate half of the second step of RFC 0067's +throughput path (the other half, committing the detached effects before the +gates, was measured and not adopted: decision log, 2026-09-28): cross-branch writers, independent merges, and reads stop serializing process-wide on one mutex, while a schema apply still excludes every writer and every writer still excludes a schema apply. Nothing cross-process changes; the manifest CAS @@ -53,7 +58,11 @@ The concurrent-writes benchmark scenario (the whole-run instrument RFC 0067's "Concurrent-writes throughput diagnostics") has measured what the exclusive gate costs. Sequential runs, both engines, closed-loop and therefore diagnostic rather than claim-grade (RFC 0039 Rule 1), all writers on one -branch (`--write-branches 1`): +branch (`--write-branches 1`). The table's base commit and run date are not +recorded; the 2026-09-28 decision-log entries ran on a later base +(`b14c22c5` for the attribution entry) whose one-writer rates differ (31.7 +and 0.87 commits/s there against 18.9 and 0.64 here), so the two sets do not +compare cell by cell: | regime | writers | commits/s (median) | service p50 | p95 | |---|---|---|---|---| @@ -74,16 +83,19 @@ The same gate sits on the read path: `capture_read_view` and its historical and current variants take it exclusively for the catalog-build window. Under write load on a remote store, a read can wait behind a writer's entire publish hold (about 1.5 s per commit at 30 ms RTT), because the writer's -guards are held from revalidation through promotion. The developer guide's +guards are held from revalidation through publication. The developer guide's claim that "reads are snapshot-isolated and do not take write gates" (`docs/dev/architecture.md`) is wrong today; this RFC makes it true for the write gates that matter and corrects the sentence either way. -This step is also the enabler for the path's third step. Group commit batches -N ready publications into one CAS — but under the current gates, writers -never *arrive* at the publisher concurrently, so every batch would have size -one. Shrinking the serialization to the branch is what lets a same-branch -batch form at all and lets cross-branch writers stop forming one global queue. +This step is also part of the enabler for the path's third step. Group commit +batches N ready publications into one CAS, but under the former gate writers +never *arrived* at the publisher concurrently, so every batch would have size +one. Shrinking the serialization to the branch lets cross-branch writers stop +forming one global queue. It does not yet let a same-branch batch form: +same-branch writers still serialize at the branch gate before the publisher +(`commit_all`, `crates/omnigraph/src/exec/staging.rs`), so releasing the +branch gate ahead of the publisher is step-3 work. ## User and operational behavior @@ -91,17 +103,23 @@ No API, format, wire, or configuration change. Observable differences: - Cross-branch concurrent writers scale instead of serializing process-wide; same-branch writers keep today's ordering and conflict behavior. -- A read no longer waits for another handle's publish window to capture its - catalog view; on the schema gate it waits only for an in-flight - contract-lifecycle pass, as it must. A read through the writer's own - handle still waits while that handle publishes on its bound branch: the - publish holds the handle's coordinator lock across the manifest - compare-and-swap, and the capture needs it (see Unresolved questions). +- The server's reads on `main` still wait for its writes to `main`: the + server shares one handle per graph, a publish on a handle's bound branch + holds that handle's coordinator lock across the manifest compare-and-swap, + and a read capture needs the same lock (see Unresolved questions). The gain + on the read path reaches deployments with more than one handle: there a + read no longer waits for another handle's publish window to capture its + catalog view. On the schema gate a read waits only for an in-flight + contract-lifecycle pass and, in plain mode, for a queued one: a queued + exclusive acquisition (an apply, or an `Omnigraph::open` in the same + process) parks every later shared permit until it runs, so a read can still + wait behind another handle's publish window through that queue. - A schema apply or system-column upgrade still waits for every in-flight - shared holder to drain, then excludes all of them — the same fairness as - today's mutex, made explicit by a write-preferring lock: once the exclusive - acquisition is queued, new shared permits queue behind it, so a stream of - writers cannot starve a schema apply. + shared holder to drain, then excludes all of them: the same fairness as the + former mutex. In plain mode the lock is write-preferring: once the + exclusive acquisition is queued, new shared permits queue behind it, so a + stream of writers cannot starve a schema apply. Installed (DST) mode has no + such preference (Design, DST determinism). ## Design @@ -114,29 +132,98 @@ in-repo precedent for a shared/exclusive permit pair (`ExportCutPermit` / `ExportDestructivePermit`). Two typed permits replace the anonymous `QueueGuard` at the schema position: -- `SchemaSharedPermit` — read side; taken by commit_all, the no-op - conditional-mutation CAS, merge, ensure_indices and the full-text rebuild, - optimize, cleanup, repair, branch create/create-from/delete, and the - read-view captures. The write capture (`open_write_txn`) takes it only - while a schema-apply sentinel stands: it parks on the shared side until - the apply releases, then recaptures; the common path takes no permit - there. -- `SchemaExclusivePermit` — write side; taken by schema apply, the - system-column upgrade, and the contract-lifecycle passes on the handle: - open, refresh, `settle_pending_schema_install`, `reload_schema_if_source_changed`, - and `sync_branch`'s coordinator swap. +- `SchemaSharedPermit`, the read side: taken by `commit_all`, the no-op + conditional-mutation CAS, `branch_merge_impl`, `ensure_indices` and the + full-text rebuild, optimize, cleanup, repair, `branch_delete_as`, + `ensure_no_pending_recovery`, the system-column upgrade under + `options.check`, the read-view captures, and the test-only + `reconcile_orphaned_branches`. The write capture (`open_write_txn`) takes + it only while a schema-apply sentinel stands and the gate is busy: it + parks on the shared side until the apply releases, then recaptures; the + common path takes no permit there (Errors and budgets below). +- `SchemaExclusivePermit`, the write side: taken by schema apply, the + system-column upgrade, the contract-lifecycle passes on the handle (open, + refresh, `settle_pending_schema_install`, + `reload_schema_if_source_changed`, and `sync_branch`'s coordinator swap), + and `branch_create_as` and `branch_create_from_impl`. The classification rule is auditable in one sentence: **a pass that only reads the accepted contract/catalog view shares; a pass that can change which -view is accepted excludes.** That is the same division the key's own -documentation already draws ("serializes every graph-global schema writer … -against the passes that install or discard a staged schema contract", -`crates/omnigraph/src/db/manifest.rs`). - -Acquisition order is unchanged: schema permit → branch gate → sorted -per-table gates, and the permits ride in `CommittedMutation.guards` with the -same caller-held drop discipline. The gate remains non-reentrant in both -modes; the existing self-deadlock notes carry over. +view is accepted excludes, and so does a pass that changes the live +branch-ref set with no CAS over the change.** The second clause is branch +create: `create_branch_recoverably` +(`crates/omnigraph-core/src/branch_control.rs`) lists the refs, refuses a +`path_collision`, may reclaim a ref-less tree, then creates, and no CAS +covers that inventory, so two creates whose branch gates are disjoint +(`feature` from `main`, `feature/x` from `b`) both passed it under shared +permits. The rule lives on `SchemaGateSlot` +(`crates/omnigraph/src/db/write_queue.rs`); read-only open is the one site on +the exclusive side by conservatism rather than by the rule. + +**Lock order and reentrancy.** The order is the schema permit, then the +branch gate or gates, then the sorted per-table gates, then the handle's +coordinator lock (`Arc>`); +`merge_authority_cache` is taken after the branch gate and before any +coordinator open (`branch_delete_as`, merge capture). A writer's permits ride +in `CommittedMutation.gates`, a `HeldWriteGates` (`_schema`, `_branch`, +`_tables`), with the same caller-held drop discipline. The gate is +non-reentrant in both modes on one task: an exclusive holder re-acquiring +either side self-deadlocks, and a shared holder re-acquiring the shared side +can park forever behind a queued exclusive in plain mode. No call path takes +the gate twice: `refresh_coordinator_only` serves holders that need a +coordinator refresh, and `refresh_for_reprepare` keeps the reprepare off the +exclusive side (decision log, 2026-09-29). `refresh` itself holds the +exclusive side twice in sequence, once around the coordinator refresh and +once inside `reload_schema_if_source_changed`. + +**Errors and budgets.** Only `mutate` and `load` reach the write capture's +park; `branch_merge_impl`, `ensure_indices`, optimize, cleanup and repair +call `ensure_schema_apply_idle` first and get the typed `manifest_conflict` +one call earlier. In `open_write_txn` a standing sentinel +(`schema_apply_locked`, a live listing of `__manifest`'s `_refs/branches/`, +not an in-memory flag) is classified by the gate. If an exclusive permit is +held or queued in this process (`try_acquire_schema_shared` returns `None`), +the apply is local: the capture parks on the shared side, then recaptures +under the promoted contract. If the gate is free, the sentinel belongs to +another process's apply or a dead one: the capture returns the typed +`manifest_conflict` refusal once a second listing confirms the sentinel still +stands (an apply in this process may have released between the two). Parks +do not count against `MAX_CAPTURE_RETRIES` (8), which bounds torn captures +only; exhausting it returns `manifest_read_set_changed`. Nothing is emitted +while a capture is parked. Cost: the common path (no sentinel) takes no +permit and no extra request; each park costs one more listing, and the +refusal path one more again. + +**Permit sites.** The 24 acquisition sites on this tree, one row per call +site (the write capture's park counts once; the system-column upgrade's two +sides count once each). Paths are under `crates/omnigraph/src/`. + +| file | function | side | +|---|---|---| +| `db/omnigraph.rs` | `open_with_storage_and_mode` (read-write and read-only open) | exclusive | +| `db/omnigraph.rs` | `ensure_no_pending_recovery` | shared | +| `db/omnigraph.rs` | `open_write_txn` (the park, sentinel standing only) | shared | +| `db/omnigraph.rs` | `sync_branch` | exclusive | +| `db/omnigraph.rs` | `refresh` | exclusive | +| `db/omnigraph.rs` | `settle_pending_schema_install` | exclusive | +| `db/omnigraph.rs` | `reload_schema_if_source_changed` | exclusive | +| `db/omnigraph.rs` | `capture_read_view` | shared | +| `db/omnigraph.rs` | `capture_current_read_view` | shared | +| `db/omnigraph.rs` | `capture_historical_read_view` | shared | +| `db/omnigraph.rs` | `branch_create_as` | exclusive | +| `db/omnigraph.rs` | `branch_create_from_impl` | exclusive | +| `db/omnigraph.rs` | `branch_delete_as` | shared | +| `db/omnigraph/schema_apply.rs` | `apply_schema` | exclusive | +| `db/omnigraph/system_column_upgrade.rs` | `upgrade_system_columns`, `options.check` | shared | +| `db/omnigraph/system_column_upgrade.rs` | `upgrade_system_columns`, the upgrade | exclusive | +| `db/omnigraph/optimize.rs` | `optimize_all_datasets` | shared | +| `db/omnigraph/optimize.rs` | `cleanup_all_datasets` | shared | +| `db/omnigraph/optimize.rs` | `reconcile_orphaned_branches` (test-only) | shared | +| `db/omnigraph/repair.rs` | `repair_all_datasets` | shared | +| `db/omnigraph/table_ops.rs` | `maintain_indices_for_branch` (`ensure_indices`, the full-text rebuild) | shared | +| `exec/staging.rs` | `commit_all` | shared | +| `exec/mutation.rs` | `mutate_one_attempt` (the no-op conditional-mutation CAS) | shared | +| `exec/merge.rs` | `branch_merge_impl` | shared | ### DST determinism @@ -146,24 +233,12 @@ The new slot keeps that protocol on both modes: shared and exclusive acquisition pass through the same `scheduled_lock`-style turn point, and both permit drops take a turn and bump the slot epoch. Without this, seeded interleavings silently stop replaying; with it, the DST arbiter sees the same -event vocabulary it sees today. - -### Promotion outside the guards -> Superseded. The implementation measured this move as harmful (decision -> log, 2026-09-25: the successor writer promotes the same pin concurrently -> and the two replays serialize on Lance's commit path), and -> [Detached-only tables](2026-09-21-detached-only-tables.md) then removed -> promotion altogether, so there is nothing left to move. The -> `HeldWriteGates` envelope helper this section introduced stays. - - -`promote_held_all` runs today inside all three gates although its own -contract states the write is already durable and graph-visible and a failure -is left for the next writer. This RFC moves promotion after guard release: -the branch gate frees one Lance-replay earlier, which at object-store -latency is a measurable slice of the hold. The sub-decision is independently -revertible (move one call back inside the guard scope) and is listed -separately in the evidence plan. +event vocabulary it sees today. Fairness differs by mode: plain mode +inherits tokio's write-preferring FIFO; installed mode never enters the +native waiter queue, both sides use `try_read_owned` and `try_write_owned` +inside the turn loop, so grant order and starvation-freedom are properties of +the seed, and the write preference is asserted only by the plain-mode unit +test `queued_schema_exclusive_blocks_later_shared`. ### Why the old rationale no longer binds @@ -181,21 +256,32 @@ misclassified site can no longer strand a moved HEAD behind a schema lock; the worst a wrong shared classification yields is a stale-authority publish attempt, which the manifest CAS refuses exactly as it refuses any other stale writer. Correctness never rested on the gate; RFC 0067 says so explicitly -("nothing in this RFC depends on the gates for correctness"). - -Second, the classification surface is enumerable and closed. The 22 -acquisition sites divide cleanly under the one-sentence rule above, the rule -is enforceable in review (a new site must name its permit type), and the -typed permit pair makes the choice visible at the call site instead of -implicit in a mutex. +("nothing in this RFC depends on the gates for correctness"), and the code +carries it: `commit_all` (`crates/omnigraph/src/exec/staging.rs`) commits +every effect detached and revalidates the complete authority token under the +gates, and the publisher's `PublishPrecondition::ExactGraphHead` +(`crates/omnigraph/src/db/omnigraph/table_ops.rs`) refuses a stale head +whatever gates were held. + +Second, the classification surface is enumerable and closed. The 24 +acquisition sites (the permit table above) divide under the rule above, with +read-only open the one exclusive site held by conservatism; the rule is +enforceable in review (a new site must name its permit type), and the typed +permit pair makes the choice visible at the call site instead of implicit in +a mutex. ## Invariants - One publication door (invariant 2) is untouched: the CAS, its preconditions, and publisher OCC are unchanged. - One coherent accepted view (invariant 3) is what the exclusive side - protects: a contract-lifecycle pass still swaps the accepted view with no - reader or writer in flight, because shared holders drain first. + protects: shared holders drain before a contract-lifecycle pass swaps the + accepted view, so no permit holder is in flight across the swap. Two + readers of the contract stay outside the gate: the write capture + (`open_write_txn` reads the contract and builds its catalog with no permit, + safe through the sentinel probe, the trailing schema-state re-read, and + revalidation under the shared permit at commit) and `resolved_target`, + which reads the contract with no permit at all. - Recovery and pin semantics (RFC 0067, as amended by detached-only tables) are unchanged. - The deny-list line "process-local locks presented as distributed writer @@ -232,31 +318,61 @@ total, no flag or staged rollout is warranted. The acceptance bar for the implementation PR, mapped to existing owners: -- **Whole-run instrument, before/after.** The concurrent-writes scenario at +- **Whole-run instrument, as measured.** The concurrent-writes scenario at `--writers 8 --write-branches {1,8}`, local FS and RustFS with the 30 ms - injected-latency recipe. Success: w8×B8 scales toward B× the single-writer - rate while w8×B1 stays branch-gate-bound and unchanged; - `authority_conflicts` falls at B>1; the record shows the critical section - shrank to revalidate-and-publish. The B-axis baselines are recorded before - the change lands. + injected-latency recipe, three builds run back to back per cell on + `b14c22c5`; no baseline was recorded before the change landed. The + decision log is the record; no bench record lands in the tree. Measured: + w8×B8 reached 154.0 commits/s locally and 2.13 at +30 ms, 4.9× and 2.4× + the one-writer rate (31.7 and 0.87, 2026-09-28 entry) against B = 8; + w8×B1 stayed branch-gate-bound locally (20 to 25 commits/s in every + build; the +30 ms `main` cell is not recorded); `authority_conflicts` is + not reported; the critical section did not shrink to + revalidate-and-publish, since the detached commit stays inside the hold + (2026-09-28 entry). - **Existing pins that survive as-is:** the branch-gate exclusivity cell (`failpoints.rs`, cross-handle branch-gate serialization), optimize's main-gate hold over disjoint tables, and the 16-handle strict-load race in `consistency.rs`, which is gate-scope-agnostic by construction. -- **Pins to re-derive:** the mid-apply blocking test in `schema_apply.rs` - (its assertion — a mutation stays pending while an apply is in flight — - survives; its mechanism becomes reader-behind-writer and is re-verified), - and the read-only-open and refresh gate-hold tests, which split along the - new classification (capture shares; install/publication excludes). -- **New coverage this RFC commissions:** a DST concurrent-universe arm that - races ordinary writers against a schema apply. Today's concurrent universe - has no schema apply in its op mix, so the shared/exclusive boundary would - otherwise land with zero deterministic-simulation coverage. The existing - seeded seam scheduler and the write queue's turn/epoch hooks are the - mechanism. -- **Cost contracts unchanged:** `write_cost.rs` per-op counts (manifest - ceiling, the schema fence's exact read/exists counts) are asserted - byte-identical — gate scope changes overlap, never operation counts. +- **Pins re-derived and added.** The mid-apply blocking test in + `schema_apply.rs` + (`mutation_waits_for_mid_apply_schema_gate_then_reprepares`) keeps its + assertion, a mutation stays pending while an apply is in flight, with a + comment-only re-derivation to reader-behind-writer. The read-only-open and + refresh gate-hold tests + (`read_only_open_holds_schema_gate_through_catalog_capture`, + `refresh_holds_schema_gate_through_catalog_publication`) are unchanged, + because both sites stay exclusive. Added: `parked_writer_blocks_schema_apply`, + the mis-classification tripwire, asserts that the apply does not reach + `SCHEMA_APPLY_POST_SENTINEL` while a writer holds its shared permit (with + the writer's permit dropped in `HeldWriteGates::new` it is red); + `cross_branch_writers_overlap_inside_schema_gate` and + `read_capture_proceeds_while_writer_parked` pin the two shared-side gains; + `sibling_branch_creates_exclude_at_the_schema_gate` (`failpoints.rs`) pins + branch create on the exclusive side, red with both create sites shared; + `concurrent_reprepare_refresh_blocks_read.gqt` pins that a same-branch + reprepare no longer parks the process's readers; four write-queue unit + tests pin the slot, among them the plain-mode fairness pin + `queued_schema_exclusive_blocks_later_shared`. +- **DST coverage, as shipped.** The plain-mode pin + `dst_schema_apply_racing_writers_first_contact` (`seam_schedule: false`) + races two writers against a schema actor on three seeds and asserts, per + seed, `alternations >= 1` and `era_commits_between_data >= 1`. The fleet + `dst_concurrent_fleet` gains the plain-mode arms `schema` and + `schema+maint` (added; their first run is not yet recorded). The scheduler + run `dst_schema_apply_racing_writers_hunt` is `#[ignore]` and in no + workflow; strict replay for this arm is the hunt's claim, not a CI pin. + `dst_seam_scheduler_bite_and_replay` has no schema actor: it pins the + shared permit's turn/epoch protocol, and the exclusive side under the + scheduler is pinned only by the ignored hunt. +- **Cost contracts.** The `write_cost.rs` per-op counts (manifest ceiling, + the schema fence's exact read/exists counts) stay green unchanged; gate + scope changes overlap, not the common-path operation counts. The write + capture's park adds one `_refs/branches/` listing per park, on the + sentinel-standing path only. +- **Not evidenced here:** independent merges overlapping. `branch_merge_impl` + takes the shared permit, and no instrument or pin in this change runs two + merges at once; the merge case of RFC 0067's question stays open there. - **Docs that move in the same change:** the gate-order section of `docs/dev/writes.md`, and the read-path sentence in `docs/dev/architecture.md`. @@ -264,22 +380,30 @@ The acceptance bar for the implementation PR, mapped to existing owners: ## Rollout One implementation PR after acceptance: lock + permits + classification + -test re-derivations + the DST arm + doc updates, with the -before/after instrument runs in the PR body. No flag; reversibility is a -one-commit revert. +the write capture's park behind an in-flight apply + the branch-create and +reprepare follow-ups + the pins listed under Evidence and tests + the DST +arm + doc updates. The mid-apply pin's re-derivation is comment-only; the +open and refresh pins are unchanged. The instrument runs are recorded in the +decision log, not in the tree. No flag; reversibility is a one-commit +revert. ## Unresolved questions - Group commit (RFC 0067 path step 3) restructures the same critical section - at the publisher; it composes with this change (it needs concurrent - arrivals, which this change creates) and is a separate proposal. + at the publisher; it composes with this change (cross-branch arrivals are + concurrent now; same-branch arrivals still serialize at the branch gate, + which step 3 must release ahead of the publisher) and is a separate + proposal. Forcing event: that proposal opening. No decider is named. - Same-handle reads still wait on publication. A publish on a handle's bound branch holds that handle's coordinator lock across the manifest compare-and-swap (`commit_updates_on_branch_with_expected`), and a read capture takes the same lock. The server shares one handle per graph, so its reads on `main` still wait for its writes to `main`. Removing that wait means publishing without holding the coordinator lock and installing the new view - afterwards; it is a separate change. + afterwards; it is a separate change. Forcing event: the server's + same-handle wait on `main` measured as a bottleneck, or the group-commit + proposal restructuring the publisher, whichever comes first. No decider is + named. ## Decision log @@ -287,9 +411,13 @@ one-commit revert. concurrent-writes instrument's sequential cross-engine baselines as the motivating evidence. Blocked on RFC 0067 merging. - 2026-09-19 — Implemented. Optimize and cleanup take the shared side (they - hold every table gate they touch, so they already serialize with writers - at table grain); read-only open and reload stay exclusive as recorded - conservatism. The mis-classification tripwire is + serialize with writers at branch grain: optimize holds main's branch gate, + cleanup holds every listed branch's gate, and `plan_collection` re-checks + cleanup's branch listing under those gates and refuses on drift; corrected + 2026-09-29 from "table grain"); read-only open stays exclusive as recorded + conservatism, and reload takes the exclusive side by the rule, since it + republishes the accepted view (corrected 2026-09-29). The + mis-classification tripwire is `parked_writer_blocks_schema_apply`; the plain-mode fairness pin is `queued_schema_exclusive_blocks_later_shared`; the DST concurrent universe gained the schema-apply-racing-writers arm (strict replay: see @@ -314,13 +442,16 @@ one-commit revert. hunt claim, not a CI pin: an apply's table rewrite runs on the single `lance-cpu` pool thread, invisible to the arbiter, so its stall budget trips under load; the plain-mode seeds are the pin and - `dst_seam_scheduler_bite_and_replay` pins the permits' turn/epoch protocol. + `dst_seam_scheduler_bite_and_replay` pins the shared permit's turn/epoch + protocol (it has no schema actor; the exclusive side under the scheduler + is pinned only by the ignored hunt). - 2026-09-28 — Ported onto main after [detached-only tables](2026-09-21-detached-only-tables.md) removed promotion, which retires the promotion sub-decision outright, and after the `omnigraph-core` extraction. The 22 acquisition sites, the classification and the DST arm carry over unchanged; the write capture's sentinel probe - now uses the coordinator's `schema_apply_locked`. + now uses the coordinator's `schema_apply_locked`, a live listing of + `__manifest`'s `_refs/branches/`, not an in-memory flag. - 2026-09-28 — An independent review (Codex, `gpt-6-astra`) found three gaps, each verified in the code. (1) The documentation overstated the read path: a read on the writer's own handle still waits for that handle's coordinator @@ -332,9 +463,9 @@ one-commit revert. barrier opened, and the write-preferring lock ran every apply before the first writer commit. The actor now pauses a seeded 2–31 ms before each apply, and every apply of every seed lands between writer commits. (3) The hunt ran - reader actors whose read-only opens take the exclusive side with no arbiter - hook, outside the turns `sched_escapes == 0` certifies; it runs without - readers now. + reader actors whose read-only opens take the exclusive side; their gate + transitions fall outside the turns `sched_escapes == 0` certifies, so it + runs without readers now. - 2026-09-28 — The second half of RFC 0067's step 2, committing the detached table effects before taking the gates, was measured and not adopted. With timing probes in the write path, a single writer on one branch holds the @@ -346,9 +477,10 @@ one-commit revert. with eight writers, and 2.9 at +30 ms, where each failure holds the gate for about 300 ms before giving up (about 19 s of a 30 s window, more than the successful writes used). Committing detached before the gates would make - each of those losers write a commit that is dead on arrival — roughly seven - times the table writes, most of them garbage for the collector — to save at - most 15% of the hold. The per-table gate plan in RFC 0067 would also invert + each of those losers write a commit that is dead on arrival, 7.6 times the + table writes locally and 3.9 times at +30 ms (one publication plus 6.6 and + 2.9 failed revalidations), most of them garbage for the collector, to save + at most 15% of the hold. The per-table gate plan in RFC 0067 would also invert the lock order: every production path takes table gates inside the branch gate, and the key includes the branch, so they add no exclusion today. The levers the measurement points to are a cheap in-process check that fails a @@ -373,9 +505,9 @@ one-commit revert. 46% of that. Raising the ceiling requires same-branch writes that do not invalidate each other, which is the footprint admission of RFC 0067's group commit (step 3), not further work on the gate. -- 2026-09-28 — Attribution, after review asked for one effect at a time. - This change has two effects: - - read-view captures stop waiting behind another writer's publication; +- 2026-09-28 — Attribution, one effect at a time. This change has two + effects: + - read-view captures stop waiting behind another handle's publication; - publications on different branches overlap. A diagnostic build separated them: this change plus one process-wide lock @@ -403,4 +535,55 @@ one-commit revert. dominated by re-prepares, with 116 to 178 exhausted re-prepare budgets per cell in every build. The capture overlap stays a mechanism that `read_capture_proceeds_while_writer_parked` pins, not a throughput claim. +- 2026-09-29 — Five changes from the implementation, and one not made. + (1) Branch create and create-from moved to the exclusive side + (`branch_create_as`, `branch_create_from_impl`); branch delete stays + shared. `create_branch_recoverably` lists the refs, checks + `path_collision`, may reclaim a ref-less tree, then creates, with no CAS + over that inventory, so two creates with disjoint source and target gates + both passed it under shared permits. The rule gained its second clause (a + pass that changes the live branch-ref set with no CAS over the change + excludes); the seam `BRANCH_CREATE_POST_INVENTORY_PRE_NATIVE` and the pin + `sibling_branch_creates_exclude_at_the_schema_gate` show it: green on this + tree, red with both sites put back on the shared side. Cleanup and optimize + stay shared: cleanup's pre-gate branch listing is re-checked under its + gates by `plan_collection`, which refuses on drift. (2) The reprepare no + longer takes the exclusive side. `refresh_for_reprepare` runs + `refresh_coordinator_only`, then one `inspect_staged_contract` probe, and + calls `refresh` only on a published staging + (`StagedContract::Marked { published: true, .. }`); the reprepare loops in + `mutation.rs` and `loader/mod.rs` call it. Before, every `ReadSetChanged` + reprepare called `refresh`, which holds the exclusive permit twice, so + under the write-preferring lock each same-branch reprepare parked every + new shared acquisition in the process. Pin: + `concurrent_reprepare_refresh_blocks_read.gqt`, starved at entry `r1` on + the earlier head, green now. (3) The write capture's park changed shape: + with a standing sentinel and a free gate, the sentinel is another + process's apply or a dead one, and the capture returns the typed refusal + after a second listing confirms it; with a busy gate it parks on the + shared side and recaptures. Parks no longer count against + `MAX_CAPTURE_RETRIES`. This replaces the earlier shape, eight probes per + cross-process apply and a `manifest_read_set_changed` when the sentinel + cleared during the last park. (4) `upgrade_system_columns` under + `options.check` takes the shared permit, exclusive otherwise. + (5) `parked_writer_blocks_schema_apply` now asserts that + `SCHEMA_APPLY_POST_SENTINEL` is not reached while the writer is parked; + with the writer's permit dropped in `HeldWriteGates::new` it is red, where + the earlier assertion stayed green. Not made: a debug assertion for + non-reentrancy was tried and removed, because `tokio::task::try_id()` + gives no identity under `block_on` and task identity is not call-path + identity (`join!` inside one task would trip it); the rule stays + documented on `SchemaGateSlot`. Body sentences this entry supersedes: the + Summary's "Promotion, already correct with no gate held, moves after guard + release as a separate, independently revertible sub-decision" and "The + per-table gates stay: they keep two same-table stagers from wasting one + staging"; the Motivation's "held from revalidation through promotion" (now + "through publication") and "Shrinking the serialization to the branch is + what lets a same-branch batch form at all"; the section "Promotion outside + the guards", cut whole (the `HeldWriteGates` helper is now named under The + lock); the shared list's "branch create/create-from/delete"; the Evidence + bullet that promised the open and refresh pins "split along the new + classification"; and invariant 3's "swaps the accepted view with no reader + or writer in flight". The 2026-09-19 entry's grain and conservatism + clauses are corrected in place and marked.