Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
36 changes: 26 additions & 10 deletions docs/architecture/rfcs/shared-goal-authority-state-provider-v0.md
Original file line number Diff line number Diff line change
Expand Up @@ -1205,15 +1205,22 @@ Keep live-state size fixed when isolating history growth, then grow live state
separately. No goal-wide unbounded list of completed Todos or receipts may be
hidden inside the supposedly fixed live projection.

For the current `FileAuthorityStore`, a fixed projection of P bytes retained in
each of N transactions costs approximately P*N final history bytes and
P*N*(N+1)/2 cumulative document-publication bytes, before head, event, receipt,
and envelope overhead. Normal reads also decode and validate the full chain.
With P=15 KiB, the renewal-only case gives about **534 GiB** of cumulative
publication at day 10 and **4.69 TiB** at day 30. The former 380 MiB estimate was
only N*P at day 30, not the cumulative rewrite of retained projections. These
are analytical payload estimates, not physical SSD writes or measured latency;
growing receipt indexes inside every projection can make the model worse.
The original full-projection File journal retained approximately P*N payload
bytes and republished approximately P*N*(N+1)/2 bytes across N commits. That
historical model must not be applied to the current checkpoint/delta format:
#5102 retired that layout from ordinary reads and writes.

The current File provider retains a checkpoint every 64 commits plus deltas,
events and original receipts in one envelope. Its approximate retained bytes
are `H(N) = ceil(N/64)*P + sum(delta/event/receipt/metadata bytes)`, before the
live head and envelope overhead. Each commit still durably replaces the whole
envelope, so cumulative application publication is `sum(H(n))`. Warm reads
read/hash the envelope and may reuse its verified view; cold reads reconstruct
and verify the history. SQLite instead updates transactional indexed rows and
bounded checkpoint windows. These mechanisms motivate a matched experiment;
neither a formula nor a cache hit establishes a short-term default choice.
Report application publication separately from physical disk writes, and
compare current code on equal state, history, durability and cold/warm workload.

#### Preferred local direction and compatibility boundary

Expand Down Expand Up @@ -3201,7 +3208,16 @@ Qualify **one** long-lived local default profile. SQLite is the current D2
candidate; File remains the real reference/explicit profile and migration
rehearsal backend. Do not publish two ambiguous defaults, declare the current
File history layout long-horizon-qualified, or silently fall back from a
selected SQLite store. The final profile decision must cite its D2 evidence.
selected SQLite store. Release activation must cite its D2 evidence. The September 27 matched
short-history experiments also select SQLite as the **short-term default
implementation target**: mutation/restart costs beat current checkpoint/delta
File, while warm read tradeoffs depend on the projection. PR #4931 now shares
privately owned TS replay between SQLite proofs and provider-neutral archive
recovery, retaining exact byte proofs and isolating returned rows. Matched
Linux evidence reduces many-field receipt/scan p95 by 88%/68%, but large-state
budgets and sustained-memory qualification remain open. This is not permission
to enable the default now.
[Measurements, reproduction and D2/D3/L9 dependencies](../../reference/sqlite-authority-store.md#short-term-default-decision-and-matched-experiment).
PostgreSQL shares the TS semantic contracts but has independent service,
tenant, restore and capacity qualification; its deployment must not delay the
local profile's work.
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -944,12 +944,17 @@ provider 已通过;经评审切换前,已交付规则仍是 `retain_all_v0`
一个模型 Turn 可能产生多次 commit。验证历史增长时固定 live state,之后单独增加
live state。不能把所有已完成 Todo 或 receipt 的无界列表藏在所谓固定的 live projection。

当前 `FileAuthorityStore` 每笔保留 P 字节 projection,N 笔约产生 P*N 最终历史字节,
累计文档发布量约 P*N*(N+1)/2,尚未计 head、event、receipt 和 envelope。普通读取还会
解码、验证完整链。P=15 KiB 时,仅 renew 的例子在第 10 天累计发布约 **534 GiB**,
第 30 天约 **4.69 TiB**。旧文中的 380 MiB 只算了第 30 天的 N*P,并非保留历次
projection 后的累计重写。这是 payload 解析估算,不是 SSD 物理写入或实测延迟;
若每个 projection 自身还包含不断增长的 receipt index,成本可能更高。
原先全量 projection 的 File journal 约保留 P*N 字节,并在 N 笔提交中累计发布
P*N*(N+1)/2 字节。这个历史模型不能用于当前 checkpoint/delta 格式:#5102
已从普通读写路径退役旧布局。

当前 File 每 64 笔保存 checkpoint,其余保存 delta、event 和原始 receipt,仍放在一个
完整 envelope 中。保留量近似为 `H(N) = ceil(N/64)*P + sum(delta/event/receipt/metadata 字节)`,
另加 live head 和 envelope 开销;每次提交仍完整替换该 envelope,因此累计应用发布量为
`sum(H(n))`。热读会读取/散列 envelope 并可能复用已验证视图,冷读需要重建和验证历史。
SQLite 则更新事务化索引行并验证有界 checkpoint 窗口。这些机制是对照实验的依据,
不是短期默认选择的结论。必须在当前代码、相同状态/历史/持久性及冷热负载下实测,
并区分应用发布字节与物理磁盘写入。

#### 本地优先方向与兼容边界

Expand Down Expand Up @@ -2475,8 +2480,13 @@ provider 确认;权威空集合不回退到陈旧 Markdown。Legacy 与预览

长程默认应选定**一个**合格本地 profile。SQLite 是当前 D2 候选;File 保留为真实
对照、显式可选 profile 和迁移演练后端。不能发布两个含混的默认项,不能把现有 File
历史布局直接称为长程合格,也不能从选定 SQLite 静默回退。最终选择必须引用 D2
证据。PostgreSQL 复用 TS 语义合同,但 service、tenant、restore 和 capacity 单独
历史布局直接称为长程合格,也不能从选定 SQLite 静默回退。发布启用必须引用 D2
证据。9 月 27 日同负载短历史实验也选择 SQLite 作为**短期默认实现目标**:写入、
重启成本优于现有 checkpoint/delta File;热读取取舍取决于数据形状。#4931 现让
SQLite 证明与 provider-neutral 归档恢复共用持有私有状态的 TS 重放组件,保留完整
摘要字节,隔离返回记录与内部状态。同机 Linux 对照中,多字段 receipt/scan p95
降低 88%/68%,但大状态预算与持续内存资格仍未闭合,不能立即启用默认。
[实测、复现命令及 D2/D3/L9 依赖](../../reference/sqlite-authority-store.md#short-term-default-decision-and-matched-experiment)。PostgreSQL 复用 TS 语义合同,但 service、tenant、restore 和 capacity 单独
资格化;其部署不阻塞本地路线。

核对基线:#4286(命令回执/归档)、#4289(typed 工作/归属 intent)、#4292
Expand Down
129 changes: 127 additions & 2 deletions docs/reference/sqlite-authority-store.md
Original file line number Diff line number Diff line change
@@ -1,11 +1,118 @@
# SQLite authority provider

SQLite is an **opt-in local conformance candidate**, behind the existing
TypeScript `AuthorityStore` interface. File remains the default. This slice
TypeScript `AuthorityStore` interface. File remains the canonical-provider
fallback when no selector is present; this does not make every new Goal
canonical or migrate an existing legacy Goal. This slice
does not promote a goal, run a live cutover, enable cross-host writes, or
qualify ten elapsed days of operation. It does provide the explicit
version-1 to version-2 database migration described below.

## Short-term default decision and matched experiment

The September 27 decision is to target **SQLite for the next qualified local
new-Goal default**, rather than first defaulting to File and moving again.
This is an implementation direction, **not default activation or completed D2
qualification**. File remains an explicit provider, a conformance reference and
an export/recovery destination. Existing Goals retain their selected authority;
a rejected SQLite runtime must never silently open File instead.

The decision compares current File checkpoint/delta storage, not its retired
full-projection-per-commit layout. The final matched experiment ran sequentially
on Linux x86_64 (16 vCPUs), Node 22.22.3 / SQLite 3.51.3, with baseline
`96ce9efb3` and candidate `758db221e`. For each workload the baseline SQLite
arm ran immediately before its candidate arm; candidate File arms followed.
The same runner and qualified runtime served every arm. Each uses 20 warm
read samples, the final 100 writes and five fresh-process head reads. Values
below are p95 milliseconds. Cold process includes module loading and does not
clear the OS page cache. These are bounded observations, not population
estimates or formal capacity/soak qualification.

| Projection / commits | SQLite receipt before | Receipt after | SQLite scan 100 before | Scan after |
| --- | ---: | ---: | ---: | ---: |
| Mixed 20 KiB / 128 | 67.69 | 16.04 | 184.69 | 71.37 |
| Mixed 20 KiB / 512 | 75.70 | 17.18 | 187.22 | 67.61 |
| 464 Todos, 64 leases, 220 KiB / 128 | 786.01 | 91.52 | 1964.03 | 623.49 |
| Changing 1 MiB / 128 | 456.97 | 288.76 | 1190.10 | 1114.94 |
| Fixed 64 KiB / 128 | 62.96 | 9.57 | 170.34 | 56.36 |

The many-field receipt/scan costs fall by 88%/68%, but still exceed the
50/250 ms targets on this host. Changing-1-MiB scan improves only about 6%:
returning 100 complete large states remains expensive. Mixed and fixed-state
reads also improve; the optimization is no longer limited to large strings.
Write costs remain close to baseline because this is primarily a replay change.
No frozen qualification threshold has been increased.

The candidate-provider comparison retains the tradeoffs instead of declaring
one provider faster for every operation:

| Projection / commits | SQLite write / warm head / cold head | File write / warm head / cold head | File receipt / scan 100 |
| --- | ---: | ---: | ---: |
| Mixed 20 KiB / 128 | 10.26 / 2.11 / 146.14 | 21.01 / 4.81 / 296.79 | 3.28 / 97.72 |
| Mixed 20 KiB / 512 | 10.56 / 2.59 / 143.09 | 51.87 / 6.08 / 702.25 | 5.55 / 78.13 |
| 464 Todos, 64 leases, 220 KiB / 128 | 64.24 / 16.44 / 152.50 | 72.75 / 7.69 / 1494.86 | 3.48 / 939.29 |
| Changing 1 MiB / 128 | 97.50 / 19.51 / 164.87 | 761.97 / 115.46 / 1892.69 | 154.34 / 441.47 |
| Fixed 64 KiB / 128 | 12.27 / 2.84 / 144.58 | 14.85 / 3.19 / 329.43 | 3.22 / 85.83 |

SQLite has lower write and restart costs in every measured shape. File has
cheaper warm receipts and, for the many-field fixture, a cheaper warm head.
SQLite now scans ordinary/many-field history faster, while File wins the
changing-1-MiB scan. This supports SQLite as the implementation target for
mutation/restart-heavy local operation, not a universal read-performance claim.

The 1-MiB SQLite post-fill RSS rises from about 118 to 160 MiB; other measured
SQLite shapes remain near their baseline. This includes fixture and verification
allocations and is not a steady-state or leak measurement. Sustained-memory
qualification remains open. Baseline/candidate SQLite final store bytes match
in all five shapes; no persisted format is changed. Earlier macOS measurements
motivated the initial target, but a new local paired run encountered load above
40 and unstable unchanged-baseline timings. Its correctness checks were kept;
its timing samples are not used to claim improvement or qualification.

These tests change a deterministic observation field in a retained full state;
they exercise storage, not the complete CLI/Turn workflow. In `changing-1m`,
padding is resized to retain exactly 1 MiB, so the large string itself can
change. It must not be reported as a stable-payload cache benchmark.

Reproduce each arm from the same checkout and qualified Node runtime:

```sh
node --experimental-sqlite --experimental-strip-types \
examples/coordination/local-provider-comparison.ts \
--provider sqlite --workload mixed --commits 128 --samples 20
# Repeat with --provider file. Workloads: mixed, full, fixed-64k, changing-1m.
# Use --commits 512 to cross more checkpoint windows; --output writes JSON.
```

The runner creates and removes its own temporary store, checks complete
projections (including Todo metadata), original receipts and reopened state,
and records the source revision, runtime, runner hash and tracked source diff
hash. It does not open a selected live Goal. RSS includes fixture/checking
allocations; File publication bytes are application bytes, not physical disk
writes. Use the existing SQLite capacity runner for WAL traffic and D2 history
sizes. Keep performance experiments separate from concurrent test suites.

Before changing release defaults, the existing owners must close these gaps:

1. **L6 / D2:** rerun the unchanged reference capacity profiles after read-path
optimization; qualify many-field/changing-state history, crash/restore,
consumer lag, supported runtimes/platforms and the >=10-day elapsed soak.
PR #4931 contributes read-proof optimization, not a D2 pass.
2. **L8 / D3:** qualify the integrated Goal command/projection and migration
path. Reuse merged reviewed migration/retained audit (#5173), managed-host
protection (#5144) and obsolete Todo-source retirement (#5054), rather than
count them as new work. A storage benchmark does not certify long-running
execution or authorize migration of existing Goals.
3. **L9:** make new-Goal creation/onboarding choose that qualified profile,
including installed runtime admission, settings/readback and packaged entry
points. Keep explicit provider selection and reviewed backup/rollback.

Both current local providers require Node >=22.22.3. SQLite uses built-in
`node:sqlite`: it adds no database service or external SQLite package. Its
actual embedded SQLite/finalization probe remains required, and a pre-existing
managed runtime must be restarted on the qualified executable as documented
below. PostgreSQL deployment is independent of this local default decision.

## Placement and persistence

The provider belongs to the existing shared-coordination authority boundary
Expand Down Expand Up @@ -69,10 +176,28 @@ layered so that each layer pays only for what it returns:

| Layer | Proves | Cost |
| --- | --- | --- |
| Live head (`loadAuthority`, `commitAuthority`) | Head row digest over the live projection, the retained transaction at that cursor reproducing its exact commit digest, parent linkage, `min=1`/`count=max=head` cursor continuity, and the presence of the checkpoint that covers the head | One head row, one retained row, one parent digest and index lookups; independent of retained history |
| Live head (`loadAuthority`, `commitAuthority`) | Head row digest over the live projection, the retained transaction at that cursor reproducing its exact commit digest, parent linkage, `min=1`/`count=max=head` cursor continuity, and the presence of the checkpoint that covers the head | One head row, one retained row and one parent digest; no projection replay. The indexed continuity count still depends on retained cursor count |
| Materialized history (`scanCommitted`, `readReceipt`) | Every row from the covering checkpoint through the requested span, including each delta, state digest and parent lineage; paged scans also prove the lookahead row used for `has_more` | At most one checkpoint window plus the requested span |
| Archive audit (`verifyAuthorityHistory`) | The complete delta chain from the empty root, every checkpoint against retained history, and the final state against the head | Linear in retained history; qualification and recovery only |

Historical reconstruction uses the shared TS `AuthorityStateReplay` owner.
It validates and privately copies the initial state and each delta, copies only
changed object paths, and reuses exact canonical encodings for unchanged
subtrees. Both proofs still hash the complete original v0 byte sequence; this
is not a new Merkle proof or a persisted digest format. Cached subtree keys are
weak, and encoded long-string buffers are capped separately; no cache survives
the provider read/audit call. A receipt query verifies its entire covering span
without materializing projections it does not return. Scans return independent
JSON objects, so editing one row cannot alter another row or a later read.

The same owner now reconstructs provider-neutral archives. Previously a
consumer could edit a yielded transaction's projection and change the decoder's
next replay basis, causing a later digest or Goal-identity failure. The decoder
now retains private state; callers receive detached projections. Existing
File/SQLite/PostgreSQL archives and terminal seals keep the same bytes. Rejected
delta batches never advance the replay frontier, and sparse protocol arrays
are rejected rather than bypassing validation of their absent elements.

A missing head, rolled-back head, internal cursor gap, rewritten receipt/event,
orphaned parent digest or mismatched state digest is rejected as
`provider_protocol_violation` before returning authority or accepting a write.
Expand Down
Loading
Loading