diff --git a/docs/architecture/rfcs/shared-goal-authority-state-provider-v0.md b/docs/architecture/rfcs/shared-goal-authority-state-provider-v0.md index 693f728e5d..bf7c936b33 100644 --- a/docs/architecture/rfcs/shared-goal-authority-state-provider-v0.md +++ b/docs/architecture/rfcs/shared-goal-authority-state-provider-v0.md @@ -1280,11 +1280,19 @@ budget, one checkpoint history read and 1,048,576 retained projection bytes plus This is still not completion of lane L. File and NoKV continue to retain and decode their complete journal on every load, so bounded recovery is a property of the embedded candidate rather than cross-provider parity; the SQLite profile -also still retains receipts and events without pruning. Unavailable logical/WAL -traffic and the <=15x cumulative write-growth budget, 1 MiB and 300k headroom, -full-domain workload, large-history recovery, fenced backup/restore, supported -upgrade/rollback, OS/runtime coverage and the >=10-day elapsed soak remain -holds, and runner completion cannot claim them. See the +also still retains receipts and events without pruning. The split +storage-traffic measurements landed with the matched-capacity entrypoint +(#4224 batch 1): the formal 64 KiB 10k/100k profile on the reference runtime +(Node 22.22.3/SQLite 3.51.3, declared local host) measures logical writes at +70,326 vs 70,324 bytes per commit, WAL traffic at 22,611 vs 22,623 bytes per +commit (exact frame counts over pinned-read-mark windows), and an app-observed +lock-wait p95 of 244 ms under a 200 ms held write lock — cumulative growth +ratios of 10.00x and 10.01x against the <=15x budget, with whole-run WAL +totals, pure busy-handler time and physical device writes still unmeasured. +Remaining holds: 1 MiB and 300k headroom, full-domain workload, +large-history recovery, fenced backup/restore, supported upgrade/rollback, +OS/runtime coverage and the >=10-day elapsed soak; runner completion cannot +claim them. See the [SQLite qualification commands](../../reference/sqlite-authority-store.md#reproduce-validation). The public minimum remains Node 22.18 for File; SQLite additionally requires synchronous finalization and the WAL-reset fix, with Node 22.22.3/SQLite 3.51.3 diff --git a/docs/reference/sqlite-authority-store.md b/docs/reference/sqlite-authority-store.md index bdd6674e25..6dbe99b067 100644 --- a/docs/reference/sqlite-authority-store.md +++ b/docs/reference/sqlite-authority-store.md @@ -278,15 +278,28 @@ the isolated fixture. `--python` chooses the Python executable. These figures include process startup but do not drop the OS file cache. Cold Node-only load and warm actual-provider calls are separate. The provider's normal per-call connection open/close remains inside warm timing. CLI mutations happen after -the fixed-history measurement; their extra commits are reported separately. +the fixed-history measurement; CLI, traffic-window and lock-probe commits all +extend past the fill target and are counted separately, so target-state rows +keep their meaning. Reports carry p50/p95/p99 and counts, parent-process RSS, application request JSON bytes and separate DB/WAL/SHM sizes at the target history. Resource-usage peak RSS is process-lifetime across both groups; sampled axis RSS is separate, and CLI child RSS is not measured. Application bytes, final files, SQLite logical writes, cumulative WAL traffic and physical device writes are different -metrics. The unavailable write-traffic and pure busy-wait metrics remain -`missing`; a final WAL size of zero proves no cumulative-write bound. +metrics and are never substituted for one another. Logical write volume is +measured from the filled database as the serialized bytes each commit hands to +SQLite (commits row plus the full-projection head rewrite plus amortized +checkpoint rows); page, index and compaction overhead belong to the other +columns. Cumulative WAL traffic is measured over one bounded commit window per +axis: read marks pinned by two observer connections make every WAL reset +impossible, so frame growth over the window is exact, and the per-commit +traffic at both depths carries the <=15x cumulative-growth budget. Lock wait is +app-observed: a probe process holds the write lock for a controlled interval +and the end-to-end store commit wait is reported against the uncontended +baseline. Whole-run WAL totals, pure busy-handler time and physical device +writes remain `missing`; a final WAL size of zero still proves no +cumulative-write bound. Each axis reserves 5 GiB free space, caps its database at 16 GiB and checks a 2,400-second fill budget. All data are generated in a new temporary directory; @@ -348,12 +361,17 @@ command, migration manifest and reverse export remain separate deliverables. ### Qualification holds / 资格保留项 The report's `passed` rows apply only to their named axis and sample counts. -`failed` measurements remain failed; `missing` rows include cumulative storage -writes, pure lock wait, steady-state RSS proof, the full domain profile, 1 MiB -and 300k headroom, 24-hour consumer lag, large-history recovery, fenced -backup/restore, supported upgrades/rollback, OS/runtime coverage and a real ->=10-day soak. Those holds still block profile promotion. Accelerated volume -never substitutes for elapsed time, and running this command starts no soak. +`failed` measurements remain failed. The split storage-write rows — +`logical_write_growth`, `wal_traffic_growth` and `lock_wait_observed` — carry +the <=15x cumulative-growth budget as per-commit traffic measured at both +depths, and an invalidated window or missing probe is missing evidence, never a +pass from the surviving columns. `missing` rows still include whole-run WAL +totals, pure busy-handler time, physical device writes, steady-state RSS proof, +the full domain profile, 1 MiB and 300k headroom, 24-hour consumer lag, +large-history recovery, fenced backup/restore, supported upgrades/rollback, +OS/runtime coverage and a real >=10-day soak. Those holds still block profile +promotion. Accelerated volume never substitutes for elapsed time, and running +this command starts no soak. Retained state is measured where the formal profile runs: an axis reports its checkpoint count, replay budget, recovery tail and retained projection/delta diff --git a/examples/coordination/sqlite-capacity-report.ts b/examples/coordination/sqlite-capacity-report.ts index 8271c448c5..cd8f30e25c 100644 --- a/examples/coordination/sqlite-capacity-report.ts +++ b/examples/coordination/sqlite-capacity-report.ts @@ -18,6 +18,46 @@ export function latency(samples: readonly number[]): Latency { return {n: sorted.length, p50_ms: at(.5), p95_ms: at(.95), p99_ms: at(.99)}; } +/** + * Exact WAL traffic over one bounded commit window. A held read mark blocks + * every WAL reset, so the file only appends and frame growth is exact. The + * window measures per-commit traffic at one history depth; it is not a + * whole-run total. + */ +export type WalTrafficWindow = + | {status: "measured"; warmup_commits: number; window_commits: number; page_size_bytes: number; + frame_bytes: number; wal_bytes: number; frames: number; wal_bytes_per_commit: number} + | {status: "invalid"; reason: string}; + +/** + * Logical write volume from the filled database itself: the serialized bytes + * each commit hands to SQLite (commits row plus the full-projection head + * rewrite, with checkpoint rows amortized). Page, index and compaction + * overhead are deliberately excluded; they belong to WAL traffic and file + * growth, which are reported separately. + */ +export interface LogicalWriteAccounting { + commits_rows_sampled: number; + commits_row_bytes_mean: number; + checkpoints: number; + checkpoint_row_bytes_mean: number; + head_projection_bytes: number; + per_commit_logical_bytes: number; + cumulative_logical_bytes: number; + formula: string; +} + +/** + * App-observed lock wait: end-to-end store commit latency while a probe + * process holds the database write lock for a controlled interval. The + * node:sqlite driver does not expose busy-handler internals, so this is the + * application-observed wait, not pure busy time. + */ +export type LockWaitProbe = + | {status: "measured"; samples: number; held_write_lock_ms: number; + uncontended_commit_p50_ms: number; observed_wait: Latency} + | {status: "invalid"; reason: string}; + export interface CapacityAxis { target_commits: number; completed_commits: number; @@ -33,6 +73,9 @@ export interface CapacityAxis { bounded_profile: SqliteAuthorityBoundedProfile | null; /** Linear archive audit, only requested where its cost is affordable. */ history_audit: {status: string; commits: number; checkpoints: number} | null; + wal_traffic_window: WalTrafficWindow | null; + logical_writes: LogicalWriteAccounting | null; + lock_wait: LockWaitProbe | null; sampled_peak_rss_bytes: number; resource_peak_rss_bytes: number; fill_seconds: number; @@ -106,10 +149,37 @@ export function capacityLedger(axes: readonly CapacityAxis[], formal: boolean): scope: "retained checkpoint and delta bytes against one full copy per retained commit", observed: retained, budget: Math.floor(perCommitCopy / 8), unit: "bytes"}); } + // Logical writes, WAL traffic and final file size are three separate + // measurements; none may substitute for another. Each growth row is the + // cumulative 10k -> 100k growth implied by per-commit traffic measured at + // both depths under the identical matched workload, so a per-commit cost + // that grows with history depth fails the <=15x budget. + const perCommitGrowth = (id: string, baselinePerCommit: number | undefined, + finalPerCommit: number | undefined, method: string) => { + const ratio = baselinePerCommit !== undefined && finalPerCommit !== undefined && + baselinePerCommit > 0 && finalPerCommit > 0 ? 10 * (finalPerCommit / baselinePerCommit) : undefined; + if (!ready || ratio === undefined || !Number.isFinite(ratio)) { + rows.push({id, status: "missing", scope: `requires the complete matched profile and a measured window at both depths (${method})`}); + } else rows.push({id, status: ratio <= 15 ? "passed" : "failed", + scope: `cumulative ${method} growth from 10k to 100k commits at fixed live state and delta sizes`, + observed: ratio, budget: 15, unit: "ratio"}); + }; + perCommitGrowth("logical_write_growth", + baseline?.logical_writes?.per_commit_logical_bytes, final?.logical_writes?.per_commit_logical_bytes, + "logical write"); + perCommitGrowth("wal_traffic_growth", + baseline?.wal_traffic_window?.status === "measured" ? baseline.wal_traffic_window.wal_bytes_per_commit : undefined, + final?.wal_traffic_window?.status === "measured" ? final.wal_traffic_window.wal_bytes_per_commit : undefined, + "WAL traffic"); + const lock = final?.lock_wait; + if (!ready || lock?.status !== "measured" || lock.observed_wait.n !== (formal ? 12 : 3)) { + rows.push({id: "lock_wait_observed", status: "missing", + scope: "requires the matched profile's held-write-lock probe at the 100k axis"}); + } else rows.push({id: "lock_wait_observed", status: "passed", + scope: "app-observed store commit wait while a probe process holds the write lock; driver busy-handler internals remain unexposed", + observed: lock.observed_wait.p95_ms, unit: "ms"}); const scope: Record = { domain_workload: "eight agents, four writers, leases/capture/archive and the production-scale fixture remain separate", - cumulative_storage_writes: "application input bytes and final files cannot qualify logical writes, WAL traffic or the <=15x budget", - lock_wait_distribution: "no pure busy-handler timing is exposed by this node:sqlite driver", steady_state_rss: "sampled RSS and per-process peak are observations, not a proof across steady-state windows", large_history_recovery: "small fault regressions do not qualify bounded recovery of a 100k history; the linear archive audit is only launched in the rehearsal profile", payload_and_headroom: "1 MiB, 300k and bursts are not launched by this profile", diff --git a/examples/coordination/sqlite-capacity.ts b/examples/coordination/sqlite-capacity.ts index 865d8d9a8d..0e0fdd7d40 100644 --- a/examples/coordination/sqlite-capacity.ts +++ b/examples/coordination/sqlite-capacity.ts @@ -1,10 +1,11 @@ /** Disposable SQLite qualification. No live goal or caller-supplied runtime. */ import assert from "node:assert/strict"; import {createHash} from "node:crypto"; -import {spawnSync} from "node:child_process"; +import {spawn, spawnSync} from "node:child_process"; import {existsSync, mkdtempSync, readFileSync, readdirSync, rmSync, statfsSync, statSync, writeFileSync} from "node:fs"; import {cpus, platform, release, tmpdir, totalmem} from "node:os"; +import {DatabaseSync} from "node:sqlite"; import {delimiter, dirname, join, relative, sep} from "node:path"; import {performance} from "node:perf_hooks"; import {fileURLToPath} from "node:url"; @@ -48,12 +49,17 @@ const report: Record = { fill_read_write_ratio: "5:1", read_mix: "three head, oldest receipt, deterministic middle receipt", records: "one native synthetic Todo; fixed padding isolates history growth; not the full domain profile", sampling: "last 1000 commits (or entire smaller rehearsal); nearest-rank quantiles", + traffic_window: "8 warmup commits, then one bounded window (1000 formal / 100 rehearsal) with a held read mark that blocks WAL resets; exact frames from WAL file growth", + lock_probe: "a probe process holds the write lock for 200 ms per sample (12 formal / 3 rehearsal); the end-to-end store commit wait is reported", cold_cli: options.cli ? "new Python process and newly started managed Effect runtime per sample; shutdown outside timing" : "not_requested", cold_node: "new Node process and import plus first load; OS file cache is not dropped", warm: "same process, actual provider opens and closes each connection"}, durability: {journal_mode: "WAL", synchronous: "FULL", altered_for_measurement: false}, budgets: {per_axis_fill_seconds: 2400, database_bytes: 16 * 1024 ** 3, minimum_free_bytes: 5 * 1024 ** 3}, - metric_limits: {logical_storage_writes: "missing", cumulative_wal_traffic: "missing", + metric_limits: { + logical_storage_writes: "measured per commit as serialized bytes handed to SQLite (commits row delta+events+receipts, full-projection head rewrite, amortized checkpoint row); page, index and compaction overhead excluded", + cumulative_wal_traffic: "measured over one bounded commit window per axis with a held read mark that blocks WAL resets; whole-run WAL totals remain unmeasured", + lock_wait: "app-observed end-to-end commit wait while a probe process holds the write lock; includes connection open, excludes busy-handler internals (not exposed by node:sqlite)", pure_busy_wait: "missing", physical_device_writes: "missing", rss_scope: "Node parent sampled per axis; resourceUsage peak is process-lifetime across both axes; CLI child RSS is not measured", application_byte_scope: "fixed-history fill requests only; excludes CLI requests and storage/checkpoint work"}, @@ -105,12 +111,29 @@ async function measureAxis(count: number): Promise { const axis: CapacityAxis = {target_commits: count, completed_commits: 0, projection_json_bytes: 65536, sample_window: Math.min(1000, count), status: "failed", warm: null, cold_node: null, cold_cli: null, application_request_json_bytes: 0, files_at_target: null, sampled_peak_rss_bytes: process.memoryUsage().rss, - bounded_profile: null, history_audit: null, + bounded_profile: null, history_audit: null, wal_traffic_window: null, logical_writes: null, lock_wait: null, resource_peak_rss_bytes: 0, fill_seconds: 0, cli_commits: 0, cleanup_verified: false}; const commits: number[] = [], heads: number[] = [], receipts: number[] = []; const timed = async (fn: () => Promise, samples?: number[]): Promise => { const start = performance.now(), result = await fn(); samples?.push(performance.now() - start); return result; }; + // Instrumentation commits (traffic window, lock probe) extend the same + // matched workload past the fill target; they are counted separately so the + // ledger's target-state rows keep their meaning. + let revision: string | null = null; + let extraCommits = 0, instrumentOrdinal = count + 1; + const commitProbe = async (operationId: string): Promise => { + const ordinal = instrumentOrdinal++; + const input: AuthorityStoreCommit = {expected_provider_revision: revision, operation_id: operationId, + next_projection: projection, + events: [{kind: "synthetic", ordinal, data: "e".repeat(1800)}], + receipts: [{operation_id: operationId, ordinal, data: "r".repeat(1800)}]}; + const started = performance.now(); + const result = await store.commitAuthority(input); + if (result.status !== "applied") throw new Error(`instrumentation commit rejected: ${operationId}`); + revision = result.provider_revision; extraCommits++; + return performance.now() - started; + }; const environment = {...process.env, PATH: dirname(process.execPath) + delimiter + (process.env.PATH ?? ""), NODE_OPTIONS: "--experimental-sqlite", TMPDIR: root, TMP: root, TEMP: root}; let phase = "setup", managedRuntimeUsed = false; @@ -141,7 +164,6 @@ async function measureAxis(count: number): Promise { repo: root, state_file: "state.md", status: "active", domain: "synthetic-storage-qualification", adapter: {kind: "read_only_project_map_v0", status: "connected-read-only"}, coordination: {registered_agents: ["agent-a"]}}]})); - let revision: string | null = null; phase = "matched_fill"; const start = performance.now(); for (let i = 1; i <= count; i++) { @@ -200,6 +222,114 @@ async function measureAxis(count: number): Promise { : {status: audit.status, commits: 0, checkpoints: 0}; assert.equal(audit.status, "verified"); } + phase = "logical_writes"; + // Logical write volume comes from the filled database itself, so the + // numbers stay separated from WAL traffic and final file size. + { + const db = new DatabaseSync(store.path); + try { + const head = db.prepare("SELECT length(CAST(projection AS BLOB)) AS bytes FROM head WHERE singleton = 1") + .get() as {bytes: number}; + const commitRow = db.prepare(`SELECT count(*) AS n, + avg(length(CAST(delta AS BLOB)) + length(CAST(events AS BLOB)) + length(CAST(receipts AS BLOB))) AS mean + FROM (SELECT delta, events, receipts FROM commits ORDER BY cursor DESC LIMIT 1000)`).get() as + {n: number; mean: number}; + const checkpointRow = db.prepare( + "SELECT count(*) AS n, avg(length(CAST(projection AS BLOB))) AS mean FROM checkpoints").get() as + {n: number; mean: number}; + assert(commitRow.n > 0 && head.bytes > 0 && Number.isFinite(commitRow.mean)); + const interval = checkpointRow.n > 0 ? axis.completed_commits / checkpointRow.n : Number.POSITIVE_INFINITY; + const perCommit = commitRow.mean + head.bytes + + (Number.isFinite(interval) ? checkpointRow.mean / interval : 0); + axis.logical_writes = {commits_rows_sampled: commitRow.n, + commits_row_bytes_mean: Math.round(commitRow.mean), checkpoints: checkpointRow.n, + checkpoint_row_bytes_mean: Math.round(checkpointRow.mean), head_projection_bytes: head.bytes, + per_commit_logical_bytes: Math.round(perCommit), + cumulative_logical_bytes: Math.round(perCommit * axis.completed_commits), + formula: "per commit = commits row (delta+events+receipts) + full-projection head rewrite + checkpoint row amortized over its interval"}; + } finally { db.close(); } + } + phase = "wal_traffic_window"; + // WAL traffic is measured over one bounded window with read marks pinned + // so no checkpoint can reset the WAL: the file only appends and frame + // growth is exact. Two observers are needed because a store connection + // checkpoints and resets the WAL on every close: the first observer + // blocks that reset while one align commit leaves frames behind, then the + // second observer pins a read mark inside the now non-empty WAL, where a + // reset is impossible. The window measures per-commit traffic at this + // history depth; it is not a whole-run total. + { + const windowWarmup = 8, windowCommits = formal ? 1000 : 100; + const pin = (observer: DatabaseSync) => { + observer.exec("BEGIN"); + assert.equal((observer.prepare("SELECT count(*) AS n FROM commits").get() as {n: number}).n, + axis.completed_commits + extraCommits); + }; + const reset = (observer: DatabaseSync) => { + try { observer.exec("ROLLBACK"); } catch { /* the transaction already ended */ } + observer.close(); + }; + const blocker = new DatabaseSync(store.path); + const marker = new DatabaseSync(store.path); + try { + blocker.exec("PRAGMA busy_timeout = 5000"); + for (let i = 1; i <= windowWarmup; i++) await commitProbe(`wal-warmup-${count}-${i}`); + const pageSize = (blocker.prepare("PRAGMA page_size").get() as {page_size: number}).page_size; + pin(blocker); + await commitProbe(`wal-align-${count}`); + pin(marker); + const walPath = store.path + "-wal"; + const startSize = statSync(walPath).size, startHeader = Buffer.from(readFileSync(walPath).subarray(0, 32)); + assert(startSize > 32, "WAL must hold the align commit when the read mark is pinned"); + for (let i = 1; i <= windowCommits; i++) await commitProbe(`wal-window-${count}-${i}`); + const endSize = statSync(walPath).size, endHeader = Buffer.from(readFileSync(walPath).subarray(0, 32)); + const frameBytes = 24 + pageSize, growth = endSize - startSize; + if (!startHeader.equals(endHeader) || growth <= 0 || growth % frameBytes !== 0) { + axis.wal_traffic_window = {status: "invalid", reason: startHeader.equals(endHeader) + ? "WAL growth is empty, negative or not frame aligned" : "WAL header changed; a reset happened despite the held read marks"}; + } else { + const frames = growth / frameBytes; + axis.wal_traffic_window = {status: "measured", warmup_commits: windowWarmup, + window_commits: windowCommits, page_size_bytes: pageSize, frame_bytes: frameBytes, + wal_bytes: growth, frames, wal_bytes_per_commit: growth / windowCommits}; + } + } finally { + try { reset(blocker); } catch { /* cleanup only */ } + try { reset(marker); } catch { /* cleanup only */ } + } + } + phase = "lock_wait"; + // The node:sqlite driver exposes no busy-handler timing, so the probe + // measures the application-observed wait: a separate process holds the + // write lock for a controlled interval while one store commit runs. + { + const heldMs = 200, lockSamples = formal ? 12 : 3; + const probe = `const {DatabaseSync} = require("node:sqlite"); + const db = new DatabaseSync(process.argv[1]); + db.exec("PRAGMA busy_timeout = 3000"); + db.exec("BEGIN IMMEDIATE"); + process.stdout.write("READY\\n"); + setTimeout(() => { db.close(); process.exit(0); }, Number(process.argv[2]));`; + const waits: number[] = []; + for (let i = 1; i <= lockSamples; i++) { + const child = spawn(process.execPath, ["--no-warnings", "--experimental-sqlite", "-e", probe, + store.path, String(heldMs)], {stdio: ["ignore", "pipe", "inherit"]}); + const {stdout} = child; + assert(stdout, "lock probe stdout must be piped"); + const ready = new Promise((resolveReady, rejectReady) => { + const guard = setTimeout(() => rejectReady(new Error("lock probe did not hold the write lock in time")), 10000); + stdout.once("data", () => { clearTimeout(guard); resolveReady(); }); + child.once("exit", () => { clearTimeout(guard); + rejectReady(new Error("lock probe process exited before holding the write lock")); }); + }); + await ready; + waits.push(await commitProbe(`lock-probe-${count}-${i}`)); + await new Promise(resolveExit => child.once("exit", () => resolveExit())); + } + assert(axis.warm?.commit, "lock wait needs the uncontended commit baseline"); + axis.lock_wait = {status: "measured", samples: lockSamples, held_write_lock_ms: heldMs, + uncontended_commit_p50_ms: axis.warm.commit.p50_ms, observed_wait: latency(waits)}; + } phase = "cold_node"; const cold: number[] = [], samples = formal ? 20 : 3; for (let i = 0; i < samples; i++) { @@ -235,7 +365,7 @@ async function measureAxis(count: number): Promise { } axis.cold_cli = {mutation: latency(mutation), status: latency(status), quota: latency(quota)}; const final = await store.loadAuthority(); assert.equal(final.status, "loaded"); - if (final.status === "loaded") assert.equal(final.cursor, String(count + samples)); + if (final.status === "loaded") assert.equal(final.cursor, String(count + extraCommits + samples)); } axis.status = "passed"; } catch (error) { diff --git a/tests/control_plane_ts/sqlite_capacity.test.ts b/tests/control_plane_ts/sqlite_capacity.test.ts index 35baea1320..80d3244719 100644 --- a/tests/control_plane_ts/sqlite_capacity.test.ts +++ b/tests/control_plane_ts/sqlite_capacity.test.ts @@ -23,6 +23,14 @@ function axis(count: number): CapacityAxis { replay_budget_commits: 63, recovery_tail_commits: 0, retained_projection_bytes: 1024, retained_delta_bytes: 1024, retained_payload_bytes: 0, database_bytes: 4096, wal_bytes: 0, shm_bytes: 0}, history_audit: {status: "verified", commits: count, checkpoints: Math.ceil(count / 64)}, + wal_traffic_window: {status: "measured", warmup_commits: 8, window_commits: 1000, + page_size_bytes: 4096, frame_bytes: 4120, wal_bytes: 41200000, frames: 10000, + wal_bytes_per_commit: 41200}, + logical_writes: {commits_rows_sampled: 1000, commits_row_bytes_mean: 4096, + checkpoints: Math.ceil(count / 64), checkpoint_row_bytes_mean: 65536, head_projection_bytes: 65536, + per_commit_logical_bytes: 73728, cumulative_logical_bytes: 73728 * count, formula: "fixture"}, + lock_wait: {status: "measured", samples: 12, held_write_lock_ms: 200, + uncontended_commit_p50_ms: 1, observed_wait: sample(12)}, application_request_json_bytes: 0, files_at_target: {database_bytes: 0, wal_bytes: 0, shm_bytes: 0}, sampled_peak_rss_bytes: 0, resource_peak_rss_bytes: 0, fill_seconds: 0, cli_commits: 20}; } @@ -33,13 +41,50 @@ test("budget failure remains failed; small rehearsals and unavailable metrics st const rows = capacityLedger([baseline, final], true); assert.equal(rows.find(row => row.id === "head_history_growth")?.status, "failed"); assert.equal(rows.find(row => row.id === "head_p95")?.status, "passed"); - assert.equal(rows.find(row => row.id === "cumulative_storage_writes")?.status, "missing"); + assert.equal(rows.find(row => row.id === "logical_write_growth")?.status, "passed"); + assert.equal(rows.find(row => row.id === "wal_traffic_growth")?.status, "passed"); + assert.equal(rows.find(row => row.id === "lock_wait_observed")?.status, "passed"); assert.equal(rows.find(row => row.id === "elapsed_soak")?.status, "missing"); assert(capacityLedger([baseline, final], false).every(row => row.status === "missing")); final.status = "failed"; assert.equal(capacityLedger([baseline, final], true)[0]?.status, "failed"); }); +test("split storage-write rows cannot stand in for each other or hide per-commit growth", () => { + const baseline = axis(10000), final = axis(100000); + const measured = baseline.wal_traffic_window; + assert(measured?.status === "measured"); + // Per-commit WAL traffic that doubles with history depth is a 10x2 = 20x + // cumulative growth: beyond the 15x budget even though every latency row is + // healthy and final file sizes stay plausible. + final.wal_traffic_window = {...measured, wal_bytes_per_commit: 82400}; + const amplified = capacityLedger([baseline, final], true); + assert.equal(amplified.find(row => row.id === "wal_traffic_growth")?.status, "failed"); + assert.equal(amplified.find(row => row.id === "logical_write_growth")?.status, "passed"); + // The same amplification hidden inside logical writes must fail there too. + const logicalBaseline = axis(10000), logicalFinal = axis(100000); + logicalFinal.logical_writes = {...logicalBaseline.logical_writes!, per_commit_logical_bytes: 147456}; + assert.equal(capacityLedger([logicalBaseline, logicalFinal], true) + .find(row => row.id === "logical_write_growth")?.status, "failed"); + // An invalidated window, absent accounting or missing probe is missing + // evidence, never a pass from the surviving columns. + final.wal_traffic_window = {status: "invalid", reason: "fixture"}; + const lost = capacityLedger([baseline, final], true); + assert.equal(lost.find(row => row.id === "wal_traffic_growth")?.status, "missing"); + assert.equal(lost.find(row => row.id === "logical_write_growth")?.status, "passed"); + final.logical_writes = null; + final.lock_wait = null; + const stripped = capacityLedger([baseline, final], true); + assert.equal(stripped.find(row => row.id === "logical_write_growth")?.status, "missing"); + assert.equal(stripped.find(row => row.id === "lock_wait_observed")?.status, "missing"); + // A probe with the wrong sample count is not qualification evidence. + const shortProbe = axis(100000); + shortProbe.lock_wait = {status: "measured", samples: 11, held_write_lock_ms: 200, + uncontended_commit_p50_ms: 1, observed_wait: {n: 11, p50_ms: 1, p95_ms: 2, p99_ms: 3}}; + assert.equal(capacityLedger([axis(10000), shortProbe], true) + .find(row => row.id === "lock_wait_observed")?.status, "missing"); +}); + test("incomplete, wrong-size or malformed evidence cannot satisfy matched budgets", () => { for (const change of [ (value: CapacityAxis) => {value.completed_commits--;}, @@ -82,7 +127,19 @@ test("small capacity entrypoint exercises real SQLite and never claims a full qu assert.deepEqual(report.axes.map((row: CapacityAxis) => row.completed_commits), [100, 1000]); assert(report.axes.every((row: CapacityAxis) => row.cleanup_verified && row.status === "passed")); assert(report.ledger.every((row: {status: string}) => row.status === "missing")); - assert.equal(report.metric_limits.cumulative_wal_traffic, "missing"); + assert(report.axes.every((row: CapacityAxis) => row.wal_traffic_window?.status === "measured" && + row.wal_traffic_window.window_commits === 100 && + row.wal_traffic_window.wal_bytes === row.wal_traffic_window.frames * row.wal_traffic_window.frame_bytes && + row.wal_traffic_window.wal_bytes_per_commit === row.wal_traffic_window.wal_bytes / 100)); + assert(report.axes.every((row: CapacityAxis) => row.logical_writes !== null && + row.logical_writes.per_commit_logical_bytes === + Math.round(row.logical_writes.commits_row_bytes_mean + row.logical_writes.head_projection_bytes + + row.logical_writes.checkpoint_row_bytes_mean / (row.completed_commits / row.logical_writes.checkpoints)))); + assert(report.axes.every((row: CapacityAxis) => row.lock_wait?.status === "measured" && + row.lock_wait.samples === 3 && row.lock_wait.observed_wait.n === 3 && + row.lock_wait.observed_wait.p50_ms >= row.lock_wait.uncontended_commit_p50_ms)); + assert.match(report.metric_limits.cumulative_wal_traffic, /held read mark/); + assert.match(report.metric_limits.lock_wait, /app-observed/); assert.equal(report.workload.cold_cli, "not_requested"); assert.equal(report.runtime.sqlite_version.length > 0, true); });