Conversation
atomic-server's build.rs runs `pnpm run build`, which runs `build:wasm` -> `wasm-pack` -> a nested cargo. That child cargo blocked forever on the workspace `target/.cargo-lock` held by the outer cargo running the build script (e.g. an editor's `cargo check`), parking everything at 0% CPU. Give the nested frontend/wasm build its own CARGO_TARGET_DIR so it takes a separate lock and can't deadlock against the outer build. wasm artifacts are copied out via --out-dir, so the isolated dir holds only intermediates.
Immutable, signatureless, content-addressed schema resources, end to end.
A frozen id is the BLAKE3 of the RFC 8785 (JCS) canonicalization of the
machine-meaning only; descriptions/labels are presentation, excluded from the
hash, so cosmetic edits never churn ids. Resources are verified by re-hash,
never by signature.
lib:
- frozen.rs: frozen_id / verify_frozen (serde_jcs + blake3).
- subject.rs: is_frozen_did / frozen_hash_hex / from_frozen_hash.
- Tree::Frozen storage wired through both backends.
- db.rs#materialize_frozen: get_resource resolves did:ad:frozen by re-hash +
read-only parse (SaveOpts::DontSave, no class-requires validation).
- commit.rs: reject commits to frozen subjects.
server:
- GET/PUT /frozen/{hash} (PUT verifies frozen_id(body) == url hash).
browser/lib:
- Store.fetchFrozenResource resolves did:ad:frozen (GET /frozen -> verify ->
materialize); registerFrozenSchema materializes locally + optionally PUTs.
- getProperty treats description as optional (presentation, not identity).
Cross-language contract pinned by test-vectors/frozen.json (TS jcs.ts and Rust
serde_jcs produce identical ids). Capstone integration test publishes from a
producer and resolves from a fresh consumer against a real server.
Register an app-bundled *.schema.lock.json into the store: verify every frozen object by re-hash, then materialize each as a read-only Resource. A schema shipped with the code (the lockfile) now resolves with zero network — the "available without a host" path. Symmetric with buildSchemaLock on the producer.
The mutable half of the two-layer model: a normal signed Ontology (genesis DID)
on the author's drive whose classes/properties reference the immutable frozen
ids. Its stable subject is the durable "name" for the latest version, and its
signed commit history is the version log — re-running it on a schema change
records a new version while old frozen ids stay permanently resolvable. With
{ save: true } it is signed and committed.
Rust can now author frozen resource graphs that produce ids byte-for-byte identical to the TypeScript producer. Ports the full content-addressing primitive — Tarjan SCC, color-refinement canonical ordering, topological hashing, and one-unit-per-cycle — mirroring browser/lib/src/freeze.ts. Validated by a new cross-language fixture (test-vectors/freeze-resources.json): a Rust test freezes the same acyclic DAG and 2-cycle the TS side produced and asserts identical local_id -> frozen_id maps.
Schemas can now be *authored* in Rust, not just consumed: an order-preserving SchemaDef -> freeze_schema produces did:ad:frozen ids byte-for-byte identical to the TS freezeSchema, closing the multi-language-authoring goal. Builds identity- only Ontology/Class/Property bodies (presentation excluded) and content-addresses them via freeze_resources. Mirrors TS normalization (keys sorted) so the order-sensitive ontology id matches. Validated by test-vectors/freeze-schema.json: Rust freeze_schema of the same schema reproduces the TS ontology/class/property ids exactly.
Backs the data-browser "Freeze" action: freezes a resource and, by default, the
structure it references (Ontology + Classes + Properties, Document + Elements,
etc.) into content-addressed did:ad:frozen JSON-AD. References between included
resources are rewritten to frozen ids; references outside (core schema, drives,
agents) stay as subjects. Hierarchy/server metadata (parent, lastCommit) is
stripped. { closure: false } freezes only the root; { save: true } publishes to
/frozen. Returns { root, bySubject, frozen } for display/download.
Adds a Freeze item to the resource context menu (any resource) that opens a dialog showing the content-addressed did:ad:frozen JSON-AD for the resource and the structure it references (via Store.freezeStructure). Copy / Download the artifact, or Publish it to the server's /frozen store. Format toggle: JSON-AD (reproducible, no history) active; Loro (keeps CRDT history, binary) shown as coming-soon. An "Include referenced structure" checkbox toggles the closure.
ResourceContextMenu: a did:ad:frozen resource is never writable (canWrite is forced false), so Edit / Add child / Import / Delete are hidden — frozen resources are content-addressed and immutable. browser/e2e/tests/frozen.spec.ts: drives the full browser flow against a real server — freeze the dev drive (dialog shows a did:ad:frozen JSON-AD body), Publish to the server, resolve the frozen id back, and assert the frozen resource has no Edit affordance. Passes end to end.
A content-addressed resource now shows a "❄ Frozen" badge in the top bar (title: "Content-addressed and immutable — verified by hash"), the positive counterpart to hiding the edit affordances. The e2e spec asserts the badge is visible on a resolved frozen resource.
Bring did:ad:frozen content-addressed schemas and code-first schema authoring onto current develop. Resolve conflicts while keeping develop store/search/CLI behavior and wiring Freeze UI into the action registry. Co-authored-by: Joep Meindertsma <joep@ontola.io>
… freeze Unify content-addressed schema loading behind Store.useSchema, materialize from the defineSchema body registry without a network hop (including offline), add immutable Cache-Control on GET /frozen, and speed up cycle color refinement by precomputing blanked content. Expand @tomic/lib/schema exports and cover the new paths with unit tests. Co-authored-by: Joep Meindertsma <joep@ontola.io>
This was referenced Sep 1, 2026
Member
Author
|
Decision (accepted 2026-09-01, |
This was referenced Sep 1, 2026
joepio
added a commit
that referenced
this pull request
Sep 1, 2026
…tus lines (#1320) Deleted (superseded): llm-wasm-gui-plugins.md, importers.md, habits-app.md (folded into plugins.md), on-device-atomic-daemon.md (Android half lives in android-data-reuse.md, desktop remainder is a note in virtual-drive.md), rust-dependency-upgrade-audit.md (done). plugins.md is imported from feat/plugin-model (PR #1307) and annotated: everything it calls built lives on that branch, not develop. Corrected status lines: json-schema-code-first (PR #1262 in flight), virtual-drive (shipped as desktop NFS mount), table-view-filters (multi-view switcher shipped), actions (step 1 shipped), unified-sync (Phase 0b / OQ4 / OQ6 drift), encryption (opfs + vault v1 shipped), cloud-sync-managed-node (verified against LocalProcessNodeProvider only), sync.md (Flutter WS session shipped), structural-problems-index #1 (still open, compiler on). Dated "Current (...)" lines reconcile: same-agent vs rights-based peer trust (683a25d, 2026-07-17); DID = sig(genesis cert) not sig(genesis commit) (0b1b13b / 232aca8); cross-user QR pairing; legacy strokeData string parsing split by layer; multi-constraint queries vs single-pair WS subscriptions. README index rows updated to match; silent-failures.md indexed. Claude-Session: https://claude.ai/code/session_019asLKBrBWY5ovyeCgtmdSd
1 of 3 tasks
joepio
added a commit
that referenced
this pull request
Sep 19, 2026
…against develop #1500 merged on 2026-09-17, so its two e2e handoffs, the website handoff and the start-server script leave planning/ per the folder's own rules. Their open items and durable lessons move to the owning plans: the website follow-ups and the inline-RTE flake diagnosis to website-publishing.md, the CI-versus-local worker topology to e2e-concurrency.md, the remaining red specs and the useDriveClass gap to e2e-diagnostic-hygiene.md, the harness traps (redb lock outliving the socket, desynced compiled i18n catalogs, a borrowed node_modules, the server-root fetch) to silent-failures.md, and the two rules every agent needs (.wuchale regeneration, HTTP-200 gating) to AGENTS.md. The README's "Biggest gaps" list keeps only open gaps, each checked against code: the query index dropping rows (ignored test in lib/src/db/test.rs) and store growth (full snapshot per commit, CLI-only compaction) are added, both with a PR in flight; the plugin model is on develop since #1500, so only code-first schemas (#1262) remain PR-bound. Status rows corrected: #1355 (npm job) and #1347 (MCP plan) merged 2026-09-18, #1500 merged, #1498 still a draft, and the addResource / CommitBuilder counts recounted (0 non-test call sites in data-browser, 31 non-test references all in browser/lib/src). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H1U7TNdjwqXKv8MQYproiq
michielbdejong
added a commit
that referenced
this pull request
Sep 22, 2026
* fix: watch wasm/src and lib/src in server/build.rs (#1525)
`cargo run`'s mtime check decided whether to rerun the JS/WASM build by
comparing browser/*/src (and a few config files) against dist, but never
looked at the top-level wasm/ or lib/ crates. A Rust-only change under
either was invisible to it, so cargo run kept serving the previously
built atomic_wasm_bg.wasm until something else also touched a watched
JS/TS path and masked the bug.
Fixes #1517
Co-authored-by: Michiel de Jong <michielbdejong@ontola.io>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
* fix(server,lib): error-handling hygiene on the commit and read paths (#1545)
Four small defects found in a data-flow code review, one PR:
- lib/src/db.rs, sled `Db::init` (`db-sled` feature): `migrate_maybe(..)`
was followed by `.map(..)` where `.map_err(..)` was meant. The error
message was built from the Ok value and thrown away, and a migration
failure was not wrapped in it.
- lib/src/db.rs, `Db::get_resource`: a stored Loro snapshot that cannot
be read or applied still falls back to the stored propval projection,
but the failure is now a `tracing::warn!` carrying the subject and the
error instead of a silently discarded `Result`. A permanently corrupt
row warns once per read; there is no dedupe pattern in the crate and a
warn is cheap next to the read itself.
- server/src/appstate.rs: drop the "empty store" bootstrap branch. Its
`!config.store_path.exists()` was evaluated after `Db::init_redb_file`
had already `create_dir_all`ed the store directory, so it was always
false, and `Db::init_redb_file` runs `populate::bootstrap` itself on
every open (fingerprint-guarded), which made the branch redundant even
when `--initialize` turned it on. The vector-index reindex now keys off
`config.initialize`, which `build_config` computes before the store is
opened (`--initialize` or "store did not exist yet").
- lib/src/sync/engine.rs, `ingest_commit`: the legacy `set`/`push`/
`remove` rejection substring-matched the raw request body, so a commit
whose subject *is* one of those Property resources (editing the `set`
Property's description on atomicdata.dev, say) or whose body quotes
those URLs was refused. The check now runs on the parsed commit's
properties, with the same error message. Also removes a verbatim
duplicated doc line on `collect_drive_subjects`.
Tests: `ingest_commit_rejects_legacy_field_commits` now sends a signed
commit that carries the `set` property (the old bare unsigned JSON would
fail at "No signature field" before reaching the check);
`ingest_commit_accepts_values_that_mention_legacy_fields` applies a
commit whose value quotes the deprecated URLs and proves a commit on the
`set` Property reaches the ownership gate (fails on the old check);
`fresh_store_gets_core_models_without_initialize` opens `AppState` on a
brand-new store without `--initialize` and finds the core models.
Claude-Session: https://claude.ai/code/session_01KxExJcGw13DnHknGTMAhXY
Co-authored-by: Claude <noreply@anthropic.com>
* fix(table): one Escape returns the grid to keyboard navigation (#1538)
Cell editors that render their own surface — the relation picker's popover,
the markdown and JSON dialogs — opt out of the table's Escape handling via
`KeyboardInteraction.ExitEditMode` and dismiss themselves. Nothing then took
the grid out of Edit mode, where arrow keys are ignored, so the cursor sat
still until a second Escape.
- `TableEditorContext` owns `exitEditMode()`; those editors close through it
instead of poking `tableRef.focus()`.
- The grid leaves Edit mode on its own when the opt-out clears. The opt-out
only exists while such a surface is open, so its disappearance means the
surface is gone — however it was dismissed.
- Arrow keys are routed from `document` while the grid owns a selected cell
and focus has fallen to `<body>`, handing focus back to the grid. Arrow
keys only, so a stray character cannot start an edit from off-grid.
Two render bugs surfaced while writing the test:
- `onRowExpand` / `onCellResize` defaulted to inline arrows, giving the `Row`
component handed to react-window a new identity every render, which
unmounted and remounted every row (the scroll reset already warned about in
this file) and re-ran each cell's mount effects.
- `useCellOptions` published a fresh `Set` on every mount, which combined with
the above into an endless remount/render loop. It now keeps the existing Set
when the contents match.
Coverage: the package's vitest ran node-only, so the one grid test rendered
static markup and could not press a key; the e2e suite presses Escape only on
cells the table itself handles, and navigates by clicking rather than arrowing
afterwards. Adds jsdom and Testing Library, and a keyboard test that fails on
every behaviour above without this change.
Claude-Session: https://claude.ai/code/session_016e7h57D9UYUNLNMcNkn5WG
Co-authored-by: Claude <noreply@anthropic.com>
* ci(release): publish @tomic/* packages to npm on a v* tag (#1355)
* ci(release): publish @tomic/* packages to npm on a v* tag
Tag releases published crates.io and GitHub assets but left npm as a
manual `pnpm publish -r`. That is why latest is still 0.40.0 and the
beta tag is stuck on 0.41.0-beta.0. Pre-releases use the beta/rc
dist-tag so they cannot replace latest.
Co-authored-by: joepmeindertsma <joepmeindertsma@gmail.com>
* ci(release): build @tomic/lib first and skip crashing attw on publish
A single filter list compiles cli in parallel with lib, which fails
TS2307 on a fresh checkout. lib's prepublishOnly attw also crashes
(undefined filename), so publish uses --ignore-scripts after the
explicit build.
Co-authored-by: joepmeindertsma <joepmeindertsma@gmail.com>
---------
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
* ci: publish the :develop image from its own job (#1550)
The publish was the second-to-last step of the CI job, so it inherited
whatever that step had already spent: a wedged Dagger engine on the runner
burned the ten minute connect timeout and the publish was skipped without
ever attempting a build. As a step it was also unre-runnable on its own.
It is now a job, `images`, with `needs: ci` — same gate, its own Dagger
connect, and its own entry under "Re-run failed jobs".
publish-develop-image.yml is a workflow_dispatch-only publish for when the
pipeline cannot go green for reasons unrelated to whether the binary
builds. It refuses `latest` and `v*` tags, which release.yml owns.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017VVSjwfYArWmabUqCJKYoM
* Pin Rust builds and enforce Server/SaaS dependency alignment (#1499)
* Pin Rust builds and enforce Server/SaaS dependency alignment
* Run trusted CI orchestration directly on Mancave
* fix(server,lib): security hygiene from the September 2026 audit (#1558)
Six small items from planning/security-audit-2026-09.md (D, C16, F), each
verified against current code and covered by a test, plus the C20
verification recorded in the audit doc.
- CORS: any origin may still read, but Access-Control-Allow-Credentials is
sent only to origins this server answers for (configured domain, base
domain tenants, loopback, the desktop webview origins), by the same rule
RequestContext uses for Host. server/src/cors.rs, gate outside the
permissive Cors layer.
- Client errors as client errors: ParseError maps to 400, malformed commit
bodies are parse errors, signature and auth-header failures are typed
Unauthorized (AtomicError::into_unauthorized) and the server keeps that
type instead of stringifying it. 400/401 instead of 500, so they no
longer reach Sentry as crashes.
- C16: the importing flag and import source are a tokio task-local
(ws_apply::import_scope) instead of process-wide statics; the live push
loop skips the one peer a change came from via the DbEvent source_id,
which now also covers COMMIT frames applied for a live peer.
CommitIngestOpts::suppress_live_echo is removed. The Flutter listener
reads the event's source_id instead of the flag.
- collections::sort_resources is a total order (missing last, numbers
numeric with NaN placed, numbers before strings); the old comparator
panicked sort_by from 21 elements up on Rust 1.81+.
- sync::peer: LIVE_PEERS recovers from a poisoned lock instead of
unwrapping it.
- /upload answers 400 on a multipart error instead of 200 with a partial
list.
- Drive-by for the clippy gate on stable 1.94: an unnecessary to_string in
a lib/src/db/website.rs test, pre-existing on develop.
Claude-Session: https://claude.ai/code/session_01H1U7TNdjwqXKv8MQYproiq
Co-authored-by: Claude <noreply@anthropic.com>
* planning: retire the #1500 session artifacts and reconcile the index against develop
#1500 merged on 2026-09-17, so its two e2e handoffs, the website handoff and
the start-server script leave planning/ per the folder's own rules. Their
open items and durable lessons move to the owning plans: the website
follow-ups and the inline-RTE flake diagnosis to website-publishing.md, the
CI-versus-local worker topology to e2e-concurrency.md, the remaining red
specs and the useDriveClass gap to e2e-diagnostic-hygiene.md, the harness
traps (redb lock outliving the socket, desynced compiled i18n catalogs, a
borrowed node_modules, the server-root fetch) to silent-failures.md, and the
two rules every agent needs (.wuchale regeneration, HTTP-200 gating) to
AGENTS.md.
The README's "Biggest gaps" list keeps only open gaps, each checked against
code: the query index dropping rows (ignored test in lib/src/db/test.rs) and
store growth (full snapshot per commit, CLI-only compaction) are added, both
with a PR in flight; the plugin model is on develop since #1500, so only
code-first schemas (#1262) remain PR-bound. Status rows corrected: #1355
(npm job) and #1347 (MCP plan) merged 2026-09-18, #1500 merged, #1498 still
a draft, and the addResource / CommitBuilder counts recounted (0 non-test
call sites in data-browser, 31 non-test references all in browser/lib/src).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H1U7TNdjwqXKv8MQYproiq
* Reconcilable host-to-Drive mappings, for hosted vanity subdomains (#1539)
Adds Db::sync_drive_mappings, list_drive_mappings and managed_alias_hosts so a
control plane that holds no user signing keys can install and reconcile host to
Drive mappings without /bind-drive. Removal is scoped to the hosts a previous
run installed, so the localhost entries and hand-made bindings survive.
Mapping keys are normalized, since the Host header arrives in whatever case the
client sent. A host bound to a missing Drive now 404s instead of falling through
to the store root, which on a multi-tenant node answered one tenant's hostname
with another namespace's content.
--served-domain-suffix makes a server answer for *.<suffix> without taking on
the rest of --base-domain, which would also change how subjects are normalized
on a box that already holds customer data.
* fix(lib): write datatype tags in the shared export path, after sealing the edit (#1553)
Datatype tags are now written inside exportLoroDeltaInternal, after the
user's ops are sealed, so a drained commit carries them and undo no
longer steps back into the tag write.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KxExJcGw13DnHknGTMAhXY
* perf(lib): coalesce cold-load OPFS reads into one worker round trip per tick (#1554)
Store batches the render-phase local-db cache misses of a tick into a
single getResourcesWithSnapshots call, chunked at 200 subjects, and
shares an in-flight read between duplicate requests for one subject.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KxExJcGw13DnHknGTMAhXY
* fix(lib,react): emit LocalChange from Resource.remove/push; route useArray.push through the save scheduler (#1541)
Resource.push and Resource.remove now emit ResourceEvents.LocalChange
like set does, so subscribed hooks re-render; useArray.push goes through
the same save scheduler as the other writes; StoreContext no longer
hands out a phantom Store when a provider is missing.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KxExJcGw13DnHknGTMAhXY
* refactor(lib): one server-managed property list and one client-db serializer (#1542)
Server-managed properties are declared once (server-managed-props.ts)
and derived from that list everywhere, and the OPFS row is written
through a single Resource.toClientDbJsonAd/Store.recordPersistedState
path instead of two hand-rolled serializers that had drifted.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KxExJcGw13DnHknGTMAhXY
* fix(lib): route destroy through the durable outbox so offline deletes survive (#1556)
A destroy is signed into the durable outbox and drained like every other
commit, instead of being POSTed straight to the server, so a delete made
offline (or during a reload) is retried rather than lost. Incoming
updates and local-db hydration skip a subject with a pending destroy.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KxExJcGw13DnHknGTMAhXY
* fix(lib): apply destroy removal in the same transaction as the commit envelope (#1544)
A signed destroy now removes the resource, its Loro snapshot, index and
search rows and its cascade-deleted children in the same redb write that
stores the envelope and commit row, with tombstones recorded only after
that write lands. Malformed destroy commits return an error instead of
panicking the request.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KxExJcGw13DnHknGTMAhXY
* fix(lib): emit DbEvent::Destroyed for cascaded children only after the removal is applied (#1561)
recursive_remove no longer announces anything; it collects each removed
subject with its drive. remove_resource and the destroy branch of
apply_commit announce them through Db::emit_destroyed once the removal
has actually landed, so a subscriber never learns of a deletion that
could still roll back.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KxExJcGw13DnHknGTMAhXY
* fix(lib): classify commit refusals by error code, not message text (#1557)
"Commits cannot be edited" now carries error_code::IMMUTABLE_COMMIT (10)
on both transports, and the client decides whether a terminal drop is
benign from the code, falling back to the legacy message match only when
a response carries no recognised code.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KxExJcGw13DnHknGTMAhXY
* fix(query-index): list every isA encoding, check the index against the store, scope DID rows by drive stamp
A drive-scoped, sorted, class-filtered collection query returned fewer
rows than the store held, with nothing to show for it
(planning/silent-failures.md, "A query index silently disagreed with the
data it indexes"). The ignored reproduction
`is_a_encodings_all_match_the_class_constraint` found 2 of 4 rows.
Two causes, both in the read that only this query shape takes:
- Rows whose `isA` read back as a plain `String` were in the index and
were read, then hidden: the built-in collection class extender did
`is_a.to_subjects()?` on the value, and `Db::resolve_query_member`
turned any extender error into "drop the row" — from the page and from
the count. Class extenders now read `isA` through the same reference
strings the index uses, and an extender that cannot decide is skipped
with a `warn!` instead of hiding a row the reader may see.
- Rows whose `isA` was a `String` holding the JSON array (the legacy Loro
encoding of a `ResourceArray`) were never candidates: the index keyed
them by the literal `["…"]`, and the matcher rejected them for the same
reason. `Value::to_reference_index_strings` and `Value::contains_value`
now read the array's elements from one helper, so index and matcher
cannot disagree.
What should have shouted now does:
- `Db::check_query_index(&Query)` compares a query's member index with a
scan of the store and reports the missing and stale subjects
(`QueryIndexReport`). Used by the tests; a full scan, so not on the
query path.
- A first build of a filter cross-checks the equality constraints the
planner did not scan, bounded by the planner's 512-entry scan cap,
warns with the filter and the subjects it missed, and files them
(`Db::cross_check_first_build`). The hot path is untouched.
Security audit C17: a DID resource cannot be routed to a drive by its
subject, so every drive's watched filters saw it on commit. The `Db` now
resolves each watched filter's `drive` subject to its drive root
(`filter_drive_roots`, filled in `register_watched_query` and
`populate_watched_queries_cache`) and compares a DID row's `drive` stamp
with that root on the commit path and in the index build. The filter's
identity and encoding are unchanged, which is what broke the first
attempt. A row without a stamp, or a filter whose subject is not stored
here, is not excluded; read rights still apply at query time.
Also: `lib/src/db/website.rs` test — `to_hex().to_string()` is
`clippy::unnecessary_to_owned` on the current toolchain, which fails the
workspace clippy gate on develop; one-line fix so the gate is green.
Tests: the reproduction is un-ignored;
`replicated_rows_reach_a_watched_scoped_sorted_query` (5 rows watched,
17 replicated through the sync import path, 22 from every shape),
`first_build_cross_checks_the_unscanned_constraint`,
`check_query_index_names_missing_and_stale_members`,
`did_rows_stamped_into_another_drive_stay_out_of_a_watched_query`.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H1U7TNdjwqXKv8MQYproiq
* feat(server,lib): startup store-size diagnostics and automatic redb compaction
Opening the file store now logs its size (on disk and length) and open
duration, a warning above 1 GiB naming `atomic-server compact`, and redb's
`DatabaseStats` breakdown: allocated/leaf/branch pages, stored, metadata
and fragmented bytes, and the reclaimable bytes (on disk minus pages in
use, which is what `compact()` gives back; redb's sparse growth headroom
does not count).
When the file is at least 256 MiB on disk and at least 30% of it is dead,
`Db::init_redb_file` compacts it on the handle it already holds, before
any table is read and before the server listens, logs before/after sizes
and duration, and keeps the outcome in `Tree::PluginMeta`
(`Db::last_compaction`). Any failure is a warning; startup never fails
on it. Not applied to the OPFS or sled backends.
Server flags `--auto-compact` (`ATOMIC_AUTO_COMPACT`, default true),
`--auto-compact-min-mb` (256) and `--auto-compact-min-reclaimable-percent`
(30) build the `CompactionPolicy`; library callers pass one to
`Db::init_redb_file_with_policy`. Documented in the installation guide.
Test: a real redb file churned through the same open path (doubling
overwrites plus throwaway resources deleted mid-file) is compacted on
reopen, gives back most of the measured free space, reads every kept
resource back, keeps the record across the next open, and is left alone
with the policy disabled. Config tests cover the flags.
Also fixes a pre-existing `unnecessary_to_owned` lint in
`db::website` tests so `clippy -D warnings` passes.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H1U7TNdjwqXKv8MQYproiq
* fix(data-browser): make page transitions opt-in until every browser works (#1563) (#1566)
View transitions were on by default with only a user opt-out. They are
still broken outside Chromium on desktop: the card-to-page morph blows up
into an overlay on Firefox, and #1563 reports the same on Android.
Flip the setting to opt-in. The stored key is new (`viewTransitionsEnabled`,
default false) rather than a flipped default on `viewTransitionsDisabled`:
that key already holds `false` for anyone who toggled the old checkbox back
off, so reusing it would have left transitions on for exactly the people who
had noticed them misbehaving.
The transition machinery and CSS are untouched, so re-enabling by default is
a one-line change once each browser is verified.
- Settings > Accessibility now reads "Enable page transition animations".
- Playwright config and fixtures seed the new key.
- New test pins both sides of the default in useNavigateWithTransition.
Claude-Session: https://claude.ai/code/session_01Pxhq3hXh8AgAk4NZzNdnkR
Co-authored-by: Claude <noreply@anthropic.com>
* Default the navbar to the bottom on touch devices (#1567)
The navbar started at the top everywhere. On a phone or tablet that puts
the sidebar toggle, the search button and the context menu out of thumb
reach, which matters more than a navbar that sits in the same place on
every device.
The default now keys on a touch-primary pointer rather than viewport
width: a landscape tablet is as wide as a laptop, and a desktop window
dragged narrow is still mouse-operated. New `isTouchPrimary()` helper,
which `useHover` now shares instead of inlining the same media query.
Only the default changes. `useLocalStorage` falls back to it solely when
the key is absent, so a position picked in Settings is preserved.
Claude-Session: https://claude.ai/code/session_01VT2jpYQ73mPfiueW25ty3y
Co-authored-by: Claude <noreply@anthropic.com>
* Keep toasts clear of the bottom navbar, drop the old sidebar spacer (#1570)
Toasts stack from the bottom right, which is exactly where the navbar
sits when it is at the bottom, so a toast covered the search button and
the context menu. The toast container is now lifted by the navbar's
height in that case; with the navbar at the top nothing changes.
The sidebar's OverlapSpacer goes with it. It added 3.8rem beneath the
App panel below 950px with a bottom navbar, dating from the floating
navbar mode. The overlap it guarded cannot happen: SideBarWrapper's
height is the viewport minus the navbar, so the bar never covers the
sidebar.
* docs: rewrite the feature list (#1572)
Every line is one short clause. Ordered core-first: local-first, then
real-time collaboration and presence, then the things you make with it,
then the platform underneath.
Adds the features that were missing: websites, the kanban / calendar /
dashboard / timer views, table templates, the virtual drive, apps,
plugins and integrations, meetings and canvas. Drops the passkey
recovery line, folds invites into authorization and the serialization
formats into the API line.
docs/src/atomic-server.md carries the same table and is updated with it.
* Converge JS and WASM plugins: one host, one manifest, one install path (#1546) (#1571)
* Generalize PluginRelease to both runtimes and add Release, Installation and Listing classes
* Expose validate_plugin_zip so publish can check a package before storing it
* Add one host core for sandboxed plugins
One HostCore carries the effective grant (caller, never sudo, bounded by
the installation identity), the fetch policy, secret substitution, the
pinned egress check and a streaming response cap, plus get_config,
get_plugin_agent and commit. commit is refused for proposal-only
(extension world) installations. One shared wasmtime Engine and one
fuel/memory policy keyed by runtime and capability live here too.
* Make the JS StoreHost a thin delegate to HostCore
Reads and queries now go through the shared grant, fetch through the
shared policy, and the runtime uses the shared engine and limits. The
PluginHost trait moves to host_core and is re-exported here.
* Make the wasm class-extender host a thin delegate to HostCore
The host impl only translates WIT types now. The server WIT copy is
synced with atomic-plugin/wit so get-plugin-agent, which the guest SDK
already exposes, is served.
* Install both runtimes through an Installation commit and publish wasip2 releases from zip bytes
* Add the plugin runtime convergence plan for #1546
* Add manifest schemaVersion 2 shared by JS and WASM plugins
One versioned manifest carries runtime, world, entrypoints, capabilities
and network origins. v1 upgrades with defaults and serializes unchanged,
so existing release ids are stable. plugin.json translates at the boundary.
Fixtures under testdata/plugin-manifest are shared with the TS mirror.
* Mirror manifest v2 parsing in @tomic/lib against the shared fixtures
* Document manifest v2 and mark plugin.json as the translated legacy form
* Add Release, Installation and Listing to the browser server ontology
* Add installRelease, publishZipRelease and the install review reader to @tomic/lib
* Publish uploaded zips as releases and install them through the Installation review
* Install Store listings through the Installation review, keeping drafts as a secondary action
* Add the Installation view with pause, revoke and uninstall, and list installations on the drive
* Cover the Installation review screen in the plugin and Store e2e specs
* Fix padding-line lint in the manifest v2 mirror
* Derive plugin resource limits from v2 capabilities for both hosts
* Publish wasip2 releases with a v2 manifest and enforce the Installation grant set
* Serve wasip2 releases and their zip from /plugin-package
* Record every published release as a Release resource and serve private releases to its readers
* Record the plugin convergence implementation checkpoint
* Treat schemaVersion 2 as valid in the app package malformed-release test
* Box the manifest in FetchPolicy to keep the enum small
* browser: remove the legacy Plugin view and zip reader
Legacy Plugin resources are migrated into Installations by the server at
startup, so the browser no longer renders or manages class Plugin:
PluginPage, PluginCard, UpdatePluginButton, useCreatePlugin
(uninstall/update) and the client-side zip reader go. The server validates
the uploaded zip and translates plugin.json, so NewPluginButton no longer
inspects it.
AssignRights, PermissionRow and ConfigReference move under
views/Installation, their only user. PluginPermissions and pluginUtils fold
into a shared CapabilityList (extracted from InstallationReviewDialog) that
also renders the pluginPermissions the server writes on an Installation.
InstallationPage shows the manifest author the server fills in.
* browser: drive plugin list reads Installations only
The drive's plugins property was the legacy Plugin list; every plugin is
an Installation found by class under the drive now.
* lib: drop unused manifest and install exports
resolveManifest/ResolvedManifest had no caller outside its test; the
server resolves defaults, the browser validates and reviews. Also removes
the unused RUNTIME_WASIP2, WORLD_SERVER_EXTENSION and INSTALLATION_STATUSES
constants and un-exports the two helper types nothing imported. The Store
compares against RUNTIME_JS instead of a literal.
* e2e: align plugin spec with the Installation page
* Replace the legacy Plugin install hook with a startup migration and one manifest per plugin
PluginMeta now holds a single unified manifest (JSON-encoded records; legacy
MessagePack records upgrade on read and are translated from plugin.json the
first time the plugin loads). Legacy Plugin + pluginFile resources are rewritten
at startup into active Installations pinned to a Release published from their
zip, keeping subject and on-disk files. Re-activating an already materialized
release no longer re-extracts.
* Replace the KV plugin catalog with Listing resources and collapse the publish paths
A public publish now records a publicly readable Listing at <server>/listings/<id>
next to the Release; /plugin-catalog is a class query over Listings and answers
with the Listing fields plus release URL, id, runtime and world. The three
publish handlers share release::publish_release for recording, listing and
world checks. CatalogEntry and the plugin-catalog/v1 keys are gone; the KV
release store stays as the cache it is.
* Document the Installation-only plugin flow and record the cleanup checkpoint
Docs describe Release, Listing and Installation, the Store and zip-upload
paths that both end in the Installation review, and how legacy Plugin
resources are migrated. The planning checkpoint lists what this cleanup
removed and what is still open.
* Read the flat Listing shape from /plugin-catalog in the Store
* Pin a materialized install to its release id, and make pausing stop a plugin
Two things an Installation got wrong once it was the single install path.
An activation decided it had nothing to materialize by comparing the
installed manifest with the release's. Two releases of one plugin agree on
every manifest field and differ in their package bytes, so repointing an
Installation at a new release skipped the extraction and left the previous
code serving under the new release's id. `PluginMeta` now records the release
whose code is on disk and that is what the check compares; a record from
before the field existed says nothing about the bytes, so it materializes
again. The legacy migration records it too, so it still re-extracts nothing.
`paused` and `draft` did nothing at all. The hook handled `active` and
`revoked` and fell through for the rest, and no runtime path read the status,
so Pause wrote a property while the plugin kept serving its hooks. Pausing
now unregisters the class extender, which is not an uninstall: the files and
`PluginMeta` stay, so resuming re-registers the same agent rather than
minting a new identity. `installation::resolve` refuses anything but
`active`, which stops JS runs, and the startup loader skips extenders whose
Installation is not active, so a pause survives a restart.
The wasip2 install test asserted the old behaviour, that pausing and
resuming leaves the extender registered. It now asserts what pausing is for,
and the 'must not re-extract' case it was really about moved to a config
change, which is where it happens in practice.
* Update an installed plugin in place, rather than uninstalling it first
The Installation page showed its pinned release read-only and the only way
to get newer code onto a drive was Uninstall then install again. That
retires the plugin's agent, so everything it owned stopped being its own,
and a second install of the same namespace and name is refused while the
first one is there.
Uploading a zip on the Installation page now publishes it as a private
release, reviews it through the dialog the Store and the zip upload already
use, and repoints this same resource at it. The subject, the config and the
agent stay; only the release, the grants and the version change, and they
change in one commit, because the server checks the grants against the new
release's manifest and would refuse a release that arrived on its own.
The review dialog takes the wording as a prop and starts its config editor
from the config in use rather than the release default, since for an update
that is what the installer is deciding about. A zip for a different plugin
is refused in the page with the names in the message, ahead of the server
refusing it on the identifier comparison.
* Record the materialized release from the install hook, and cover the encoding
Which release's code is on disk is the Installation's business, not the zip
extractor's, so `activate` writes it after `install_or_update_plugin`
returns rather than passing it down. A failed extraction then leaves no id,
and no id means materialize again, which is the safe direction. It also keeps
that function at seven arguments, so clippy gains nothing new.
`PluginMeta`'s encoding test constructed the struct literally and so did not
compile. It now round-trips the id, and a second test pins the compatibility
that matters: a record written before the field existed reads back as `None`
rather than as some release, and is written out without the key.
* Format the touched browser files with oxfmt
The browser workspace formats with oxfmt, not prettier, and `pnpm lint` runs
`oxfmt --check` in each package. Four files had been written by prettier,
which reformatted them whole, so `lib lint` and `data-browser lint` both
failed on the format check. Running the repo's own formatter over them puts
them back, and the net change against the previous head is once again only
the lines this branch actually adds.
`pnpm run -r lint` now exits 0 across the workspace.
* fix(lib,react): make "destroyed" one fact on the store (#1573)
* fix(lib,react): make "destroyed" one fact on the store
Deleting a resource left a "Resource with error" row in the sidebar.
A `parent=` query keeps answering with a destroyed child for a moment: the
answer was computed before the destroy landed, and `Collection.fetchPage`
deliberately races the local index against the server's `/query`. Every reader
of such an answer that had not itself witnessed the delete put the row back,
and that is not cosmetic — rendering the row asks the store for the resource,
which re-creates the entry the destroy had removed and then 404s.
The knowledge that a subject was deleted only ever lived in caches scoped to
one instance: `hasPendingDestroy` (an outbox entry, dropped the moment the
server acks), `Collection._removedSubjects` and `useChildren`'s `removedRef`.
A reader created after the delete — a second view on the same query, a folder
re-expanded — had none of them, which is why the row came back.
The store now owns it, as the in-memory half of the tombstone `removeResource`
already wrote to the local database:
- `Store.isDestroyed` / `clearDestroyed`, backed by an insertion-ordered set
capped at 10k subjects.
- `Collection` and `useChildren` read it; their own sets are gone. A destroyed
subject is never a member, whatever state arrives for it, and is kept out of
`_queriedMembers` so a lifted tombstone can re-admit it.
- `applyIncoming` refuses state for a destroyed subject after the ack as well
as before it.
- `getResourceLoading` answers for one from the tombstone instead of firing a
request that can only 404.
Only creating a resource again under the same subject lifts a tombstone
(`newResource`); a stale answer no longer can. That is a behaviour change:
`applyResourceChange` used to treat any matching resource event as a restore,
which is the resurrection this fixes.
* test(data-browser): give the Markdown schema suite its own timeout
fileContentsToTiptapJson awaits a lazily imported collaborative Markdown
schema. Whichever of the two tests in that suite runs first pays the
import cost, which exceeds vitest's 5s default whenever the test threads
are oversubscribed. On the shared runner that is routine: a recent run
reported 243s of import time for 86s of wall time, and the suite timed
out there twice in a row while passing locally every time.
Give the suite the same 60s budget the recovery-enrollment suite already
uses for the same reason.
* Create a drive's plugin schema terms in one round, not nineteen (#1574)
* Create a drive's plugin schema terms in one round, not nineteen
Asking for a new plugin on a drive that has no plugin schema yet means
waiting for `ensureSchema` to write all nineteen terms of `pluginSchema()`
first, and `ensureAll` saved them one after another. Sixteen properties and
three classes, each its own awaited round trip, is most of the time between
clicking "New plugin" and the page appearing, and it is why
`plugins.spec.ts` "a plugin proposes changes, and nothing is written until
you approve" times out waiting for the New plugin heading on a cold drive.
An earlier pass already batched the orphan lookups; the saves were still
sequential, which is the rest of the wait.
The terms are independent: they write different resources under the same
ontology and none of them reads another's subject, so the order they land
in does not matter. Resolve every spec first, collect the ones that have to
be created, and send those creates together. The ontology's own list is
still assembled in spec order, because that order is what someone reading
the ontology sees.
* Note the plugin schema batching in the browser changelog
* Give the Rust checks container the plugin zip its tests include
`server/src/plugins/plugin.rs` and `server/src/handlers/plugin_release_test.rs`
`include_bytes!` `browser/e2e/tests/fixtures/test-plugin.zip`, but
`rustChecksContainer` mounts only the crate directories and `testdata`, so
`/code/browser` does not exist and both test targets fail to compile:
error: couldn't read `server/src/handlers/../../../browser/e2e/tests/fixtures/test-plugin.zip`
Mount the one file. Nothing else under `browser/` belongs in that container,
and mounting the tree would make every front-end edit invalidate the Rust
layer. This is the same treatment `testdata/pairing-request.json` already gets
a few lines above, for the same reason: a fixture both the Rust and the
JavaScript side test against.
The two call sites are behind `cfg(all(test, feature = "wasm-plugins"))`, so
only the checks container needs it; the build containers are unaffected.
* Stop filing every server error in Sentry twice (#1576)
* Stop filing every server error in Sentry twice
Two things reported the same failure. `sentry_actix` captures a handler's
5xx with the request attached, and `tracing_actix_web` separately logs
"Error encountered while processing the incoming HTTP request" at `error!`,
which the Sentry tracing layer turned into a second event.
The staging floods of September showed it plainly: the issue groups pair
up, 6313 events on one side and 6312 on the other, for the same two
incidents. On a free-tier quota that is half the budget spent on duplicates.
Give the tracing layer an event filter that ignores the `tracing_actix_web`
target and delegates everything else to `default_event_filter`, so
background work still reports as before. The log line is untouched; only
the duplicate report goes away.
Also stop the data-browser reporting from a dev server. Vite's hot reload
throws while swapping modules ("_s is not a function", "Cannot access X
before initialization", all with `@react-refresh` frames), which is not a
defect in anything shipped, and those arrived rated above every real bug in
the backlog. Setting VITE_SENTRY_ENVIRONMENT still reports from a local
build on purpose.
* Add changelog entries for the Sentry reporting fixes
* Change video link in README
Updated video link to a private user image.
* Tell people the assistant can build apps and websites (#1577)
* Tell people the assistant can build apps and websites
Before: the new-resource page listed tables, documents, dashboards and the
rest, but nothing on it said apps or websites were possible. Both classes are
minted per drive, so they only appear under "Your resource types" once the
drive already has one — which never happens to someone who was never told.
After: a row of suggestions sits under the AI composer. Clicking one puts a
half-written request in the composer with the caret at the end, so the user
finishes the sentence instead of sending "build me an app" and being asked
what they meant. The row hides once they write their own words, because
setting a textarea from code does not go on the browser's undo stack.
How: AI_BUILD_SUGGESTIONS in the creation catalog, rendered as a gradient-
bordered row in NewRoute under the composer. Unit tests cover the seeds and
the hide rule; an e2e test covers clicking, swapping and hiding.
* Honour "Enable AI Features" on the new-resource page
Before: unchecking Enable AI Features hid the navbar button and the AI chat
entry, but left the whole "Build with AI" composer and the Ask AI button in
the search field on the page. Submitting either one silently ticked the
setting back on, so the only way to get rid of the composer was to not press
it.
After: the composer, its suggestions and the Ask AI shortcut are gone while
the setting is off, and nothing turns it back on behind the user's back. The
blank buttons, templates and upload area stay, since none of them are AI
features. The no-matches line stops offering the assistant too.
How: the section and the Ask AI button render on enableAI, and both
setEnableAI(true) calls are dropped now that nothing reaching the assistant
is rendered while it is off. An e2e test sets the stored setting and checks
what remains.
* Make the build suggestions look like the class buttons
Before: the suggestions were pills with a sparkle icon and labels written as
"An app", "A website". They read as a different kind of control from the
blank class buttons right below them, which they are not.
After: App, Website, Dashboard, Custom table, on the same button as the blank
ones, with a faint rainbow border in place of the grey one. The border comes
up to full colour on hover, focus and while a suggestion is in the composer,
so it still reads as the assistant's row rather than a second Start blank.
How: Suggestion extends BasicChoice and paints its border with a two-layer
background-image, so the geometry can never drift from the blank buttons.
The sparkle is gone.
* Give the build suggestions their class icons
Before: the suggestions were the only buttons on the page with a bare label.
Next to a row of blank buttons that each carry their class's icon, they read
as a different kind of thing.
After: each one wears the icon of the class it makes. Dashboard and Custom
table take theirs from the class subject, so they are literally the same icon
as the blank Dashboard and Table buttons. Website takes the globe already
mapped to website-project. App gets the launcher grid, added to the icon map
so an app shows the same glyph wherever it appears.
How: a suggestion names either a class subject or a drive-minted class
shortname, and the row resolves it through getIconForClass like every other
class button does. The puzzle piece a new app carries as its placeholder
emoji was the obvious choice for App and is deliberately not used: plugins
and installations already have it, and an app sits beside both.
* Close three gaps the plugin convergence left behind (#1575)
A publish refused for claiming the wrong world had already stored the
package bytes and cached the release record: the check ran in the handler,
after `release::publish_package` had written both. The claim now goes into
that function, which asks straight after reading the manifest and before the
first write, so a refusal leaves nothing on the node. `expect_world` reads
the world from the manifest rather than from the release, because at that
point there is no release to read it from.
`installation_grants` fell back to the capabilities the manifest declares
whenever any step of finding the Installation's approved grants failed. That
is the one direction a fallback must not take: the declared set is what the
plugin asked for, so a lookup going wrong rewarded asking for more than was
approved. Only the absence of an Installation, which is a legacy draft that
never had grants, still reads the manifest. An Installation that cannot be
read, or whose grants cannot be, now grants nothing and logs why. The class
check also went through `Value::to_subjects`, which errors on the scalar
`isA` encodings, so an encoding alone could hand over the declared set;
`Resource::class_subjects` and `Resource::has_class` are now the one
encoding-tolerant reader, shared with `ClassExtender`.
An Installation's `release` is declared an atomicURL and held bare `blake3:`
ids, which meant every writer set it with datatype validation switched off.
Every publish records a `Release` resource and the response carries its
subject; the browser's two zip paths were throwing that away and storing the
id. `publishZipRelease` returns `subject` and the callers pass it, so the
property holds what it is declared to hold and nothing skips validation.
`release::resolve` still accepts a bare id, because Installations written
before this carry one and have to keep working.
One consequence worth knowing: an Installation whose `release` is a URL is
resolved with a read check that a bare id skipped. `/releases/<id>` is keyed
on the content hash but parented to whichever drive published first, so an
agent publishing byte-identical bytes on a second drive gets the first
drive's Release resource back and may not be able to read it. They now get a
rights error at install time rather than installing from a record they
cannot see. That a content-addressed subject is parented to one drive is a
deeper wart, left alone here.
* Stop a stale auth proof from failing requests that needed no auth (#1578)
A browser keeps its authentication proof in the `atomic_session` cookie.
`setCookieAuthentication` signed one proof and stored it for a day, and
`checkAuthenticationCookie` only asked whether a cookie existed, so nothing
ever re-signed it. Since `AUTH_MAX_AGE_MS` landed in 0.41 the server refuses
a proof older than five minutes, which made a stale proof the ordinary state
of any tab left open, and every request such a tab made was answered 401 --
public ones included.
On staging one tab polling the public `GET /server` endpoint produced 4,215
rejections in a row, every one of them carrying the same `signed at`
timestamp while the server's clock walked past the window.
Two halves:
- The browser refreshes before it goes stale. The cookie now lives no longer
than the proof it carries, and the freshness check reads the signed
timestamp out of the cookie instead of only looking for the cookie's
presence, so the request path installs a fresh proof well inside the
window.
- The server treats an aged-out proof as no proof. An HTTP request, and the
headers a socket is opened with, did not ask to be authenticated; a proof
that has expired says the caller *was* this agent and is no longer, which
is not by itself a reason to refuse. They continue as the public agent and
the rights check decides. A signature that does not verify is still
refused outright, and so is a stale proof in an `AUTH` frame or a peer
handshake, where the caller asked to be authenticated and is owed the
answer.
* Give the plugin specs a budget that matches what creating a plugin costs (#1579)
Several tests in plugins.spec.ts have been failing on develop since
16 September, all at the same point: they click `New plugin` and the page
never arrives inside the suite's 10s `expect` budget.
It is not a hang. The first plugin on a drive materializes that drive's
plugin schema first, and `pluginSchema()` is nineteen terms, each its own
resource and its own commit. The browser already sends those writes
together, but the server applies commits one at a time: instrumented, all
sixteen property saves start within 30ms of each other and finish 130 to
150ms apart, so the cost is additive whatever the client does. Measured
against a debug build, the step takes 5.2s on a fresh store and 9.4 to
11.3s once the store holds a handful of drives, which is where this spec
runs in a full-suite pass. That is why it passes alone on a laptop and
fails in CI.
So the wait gets a budget that matches the work, here rather than for the
whole suite, and the test gets room for what comes after it. The two
places that had the helper's steps copied inline now call the helper, so
there is one budget rather than three.
This is the budget, not the cost. Nineteen sequential writes to open a
page is its own problem and is being looked at separately.
* Format the tests that came with the stale-auth-proof fix (#1580)
`cargo fmt --all --check` fails on develop's head in three places, all in
tests added by #1578: the `sign_message` call in `signed_at`, the `assert_eq!`
on `error.error_type`, and the `get_client_agent` binding in
`a_cookie_whose_proof_aged_out_is_the_public_agent`.
Nothing caught it because #1578's own develop run was concurrency-held behind
an earlier push and then cancelled before it ran, so the format gate never
executed against that commit.
The cost is larger than one red check. The gate runs first and dies about
twenty seconds in, before any test starts, so develop and every branch that
merges develop produce no test information at all until this lands. Run 4180
on a feature branch already failed this way.
This is `cargo fmt --all` output and nothing else: line wrapping plus one
trailing comma, no behaviour change.
* Put back 184 UI strings that render blank in a production build (#1583)
Before: a good deal of the app's text was simply not there once built. The
page transition toggle in Settings was an unlabelled checkbox. So were both
plugin visibility toggles on the Integrations page. The Installation review
had no "Release" tab to click, the new-resource page's assistant suggestions
had no names, and 180-odd other strings were missing the same way.
After: they render. Verified by rebuilding the frontend and running the specs
that were failing on the missing text.
Why it happened: `wuchale` compiles the .po catalogs into the bundle at build
time and does not extract during a production build, so any string added
without re-running extraction has no entry, and a message with no entry
renders as nothing. The catalogs had drifted 184 strings behind the source.
Two things here. The catalogs are regenerated, which also drops 40 msgids for
text that no longer exists, most of them fragments of sentences that were
since rewritten as one. And `wuchale.config.js` now points the extractor at
the integration screens as well.
That second part matters: Connect GitHub, Google Calendar, Notion and the
other connect screens are React components inside the pinned `devonian`
package, aliased in as `@localthought/atomic-integrations`. The adapter's
default `src/**` never reaches them, so a regeneration without that line
would have taken 22 strings those screens still use out of the catalog and
blanked them, which is how this was found.
Three e2e specs were failing on the blank text, and pass with it back:
`settings.spec.ts:7`, `integration-store-install.spec.ts:14` and
`integration-visibility.spec.ts:119`, along with both `new-resource-catalog`
failures. `settings.spec.ts` also asked for the toggle by its old name; #1566
turned that setting into an opt-in and renamed it, so it now matches on the
words the setting's own search keywords use.
This is worth a guard of its own: nothing today notices when the catalog
falls behind, and the symptom is invisible to everything except a person
looking at the screen.
* Answer the integration spec's catalog mock in the shape the endpoint uses
The "integration categories default off" spec stubbed /plugin-catalog with
`{ metadata: { ... }, verification }`, which is the publish payload, not the
catalog one. `plugin_release::catalog` answers with one flat object per
Listing: name, description, publisher, domains, standards, release,
releaseId.
Reading `entry.domains` off the nested object gave undefined, and spreading it
in the store's search filter threw "domains is not iterable" during render, so
the error boundary replaced the whole integrations page. The spec then failed
on the first thing it looked for after the reload, `[data-integration="mt940"]`,
which made it read like a bundled-integration or a preference-persistence bug
rather than a bad mock.
* Fix twenty-five e2e failures, and two real bugs hiding among them (#1582)
* Fix seventeen e2e failures that were never the code under test
Before: develop's e2e reported 38 failures. Seventeen of them had nothing to
do with the features they name; they were the test setup asking for things
the pipeline does not provide.
After: all seventeen pass. Verified locally against a built bundle and a
real server, the same shape CI runs.
Three separate causes.
Eleven specs loaded app modules by source path. Seven website specs plus
installation-recovery ran `await import('/src/chunks/Website/websiteModel.ts')`
(and `'/@fs' + resolve(...)` for `browser/lib/src/ws-v2.ts`) inside
`page.evaluate`. Vite's dev server resolves those; the bundle atomic-server
embeds does not, so every one of them died on "Failed to fetch dynamically
imported module". They were added on 15 September against a dev server and
have never passed in CI; `pnpm typecheck` in browser/e2e has been reporting
them as TS2307 the whole time. An E2E build now hangs those modules on
`window.atomicE2E` through the new `helpers/e2eModules.ts`, attached behind
`isE2E()` so a release build neither exposes the registry nor pulls the
chunks into its entry graph. Their types are written out in
`browser/e2e/global.d.ts`: the e2e project cannot import data-browser's own
types without dragging the whole app graph across its rootDir.
Website publishing was never switched on for the test server. Without
`ATOMIC_WEBSITE_ORIGIN` the server answers every hosting call with "Website
hosting is disabled", which surfaces two different ways: `website-versions`
and `website.spec.ts` throw out of `hostingRequest`, and `website-inline-rte`
and `website-errors` fail on a click, because the resulting toast renders
over the preview iframe and intercepts pointer events.
`http://sites.localhost:9883` satisfies the separate-domain rule the server
enforces, and is set for both the Dagger e2e service and the local script.
The apps specs hit the wall #1579 just fixed for plugins. `createApp` calls
`ensureSchema(pluginSchema())`, the same nineteen sequential writes, against
the suite's 10s budget. Same fix in the same shape: one `newApp` helper with a
budget that matches the work, replacing five copies of the same three lines.
One more, found on the way: `newResource` looked for a class button by name
anywhere under `main`, and #1577 added an assistant suggestion row carrying
some of the same words. Two buttons named "Dashboard" is a strict mode
violation, which is what `dashboard.spec.ts:287` was failing on. The search is
now scoped to the page's sections, where the class buttons live.
Not fixed here, and still failing: website-preview-stability, which is a real
bug rather than plumbing, since opening "AI edit" tears down the website
preview the test asserts must survive. Also integration-workspace, which does
not merely import a source path but fetches the module's text and regexes
another path out of it.
* Fix six more e2e failures, two of them real crashes
The integration cluster on develop was three separate things wearing the
same clothes.
**The mock proxy was at an address the browser cannot reach.** It runs
inside the atomic-server container, so from the server's process it is on
127.0.0.1:19090 -- and the e2e build handed that same value to the browser,
which runs in the playwright container, where nothing answers there. Every
page that lists integrations then carried a second `role="alert"` reading
"TypeError: Failed to fetch", so the specs asserting on an alert read the
wrong one or failed strict mode. Pointing the bundle at
`http://atomic.localhost:19090` reaches the same container through the host
mapping chromium is already given, and the server keeps its own loopback
URL, which is right for it. Nothing server-side reads a configured proxy
origin, and the origin a LocalThought connection stores is used only by the
browser, so this value is the browser's alone. The per-test forwarding route
in the Pets spec was a workaround for exactly this and goes with it.
**Opening an AI chat could take the page down.** `useEditor` replaces the
tiptap editor when its dependencies change and destroys the old one, which
nulls its `commandManager` while its last state stays readable. So
`editor.isEmpty` still answers and `editor.commands` throws. A prefill
arriving in that same breath -- which is what "New automation" and "AI edit"
do -- left the user at the error boundary with "Cannot read properties of
null (reading 'commands')". Both effects now check `isDestroyed`, which also
covers an editor whose view has not mounted yet. The prefill is not marked
as applied when it is skipped, so the new editor still receives it.
**A release is reviewed before it is used.** The store card offers "Open",
which raises the installation review, and the draft is one of the choices
there. The spec still clicked a "Create draft" button on the card itself.
Two specs were reaching into the app by source path, which only a vite dev
server can answer: one imported `githubInstaller.ts`, read the served text
and regexed the provider bundle's path out of it, the other imported
`runScript.ts`. Both now take the modules from the `window.atomicE2E`
registry, which gains the installer and `@tomic/lib`.
Last, `integration-visibility` was sending `/plugin-catalog` the shape it
had before the catalog was flattened. `IntegrationStore` spreads
`entry.domains`, so an entry without one took the Integrations page to its
error boundary rather than failing an assertion.
Verified against a locally built bundle and a real server: all five
integration specs and `plugins.spec.ts:121` pass, and Notion and Clockify
pass as soon as the proxy is reachable and fail exactly as CI does when it
is not.
* Put a plugin's release in the commit that creates its Installation
Installing a plugin has been refused since the Installation class started
requiring `release`: the server answers the commit with "Property
.../properties/release missing. Is required in class Installation", the
outbox drops it as terminal, and nothing is installed. Both paths are
affected, the integration store and a zip upload, since both go through
`installRelease`.
`store.newResource` signs the genesis commit from the propvals it is given,
right there at creation, and holds it on the resource until `save()` moves
it into the outbox. A property set between those two calls lands in a later
commit, so the genesis commit the server validates is missing it. `release`
was set that way.
So it moves into the propvals, where the rest of the Installation's fields
already are, and the separate `set` goes.
Verified against a locally built bundle and a real server: the install now
gets through the review dialog and past the commit. The spec still fails
here, on a fetch of a class hosted at atomicdata.dev that this sandbox has
no route to, which CI does.
* Note the plugin install and AI chat fixes in the browser changelog
* Check the e2e port before wiping the store, and offer to run the mock proxy
Two ways the local e2e server told you the product was broken when it was not.
The store was wiped before the port was checked. With an older server still
listening, the script printed "Something is already listening" and exited,
leaving that process serving a store which had just been deleted underneath
it. Runs against that look ordinary and mean nothing: new-resource-catalog
lost its search box entirely, which reads like a UI regression and is not one.
The port is now checked first, nothing is wiped when it fails, and the message
says how to find the stale process. It also warns against `pkill -f
atomic-server`, which matches the caller's own command line and kills the
shell that ran it, so a restart chained after it silently never happens.
The integration specs need an integration proxy, which CI runs beside the
server and this script did not run at all. Without it Notion, Clockify and
GitHub raise "TypeError: Failed to fetch", and that second role="alert" breaks
their own strict-mode alert assertions, so four specs fail for a reason that
has nothing to do with them. plugins.spec.ts:1270, red on develop across two
CI runs and recorded as a product bug in the automation toggle, is green as
soon as the browser can reach a proxy.
`--mock-proxy` runs it. It waits for /catalog to answer rather than assuming a
started process is a working one, since a proxy that failed to bind produces
exactly the silent wrong results this is meant to prevent, and it warns when
the built bundle does not mention the port, because VITE_INTEGRATION_PROXY_URL
is read at build time and nothing the script does afterwards can rescue a
bundle built without it.
It gives the server no proxy environment of its own. Passing what .dagger
gives its e2e service took plugins.spec.ts from two failures to four, adding
:109 and :514, neither of which touches an integration provider. Those
variables belong to a container this is not.
* Give the plugin sandbox round trips a budget they can actually meet
Three of the plugin specs waited on a real trip through a real sandbox with
the suite's 10s action budget. The browser posts to the server, the server
starts a sandbox, and the plugin's discover phase answers out of it; none of
that fits 10s on a loaded box. Measured here, each of the three passes on its
own and times out on exactly that step when the suite runs it beside another,
which is what CI does on every shard.
So the waits are widened where the sandbox is, not for the suite:
- GitHub install. When the URL the test waits for appears, the install is
still running: the button reads "Connecting…" and is disabled, so
"Connections" does not exist yet. The click succeeds at 45s and the test
takes 48s, against a 60s per-test default with nothing left over, so the
test gets 120s too. This one now passes in the full file.
- Both Clockify workspace discoveries.
Same shape as the wait `newApp` documents in apps.spec.ts: the budget was
never achievable on this hardware, and each assertion is about what came back,
not about how fast it came.
This does not make the two Clockify specs green. It moves them off the
discovery step they never got past and on to later assertions — "skips
repeats" at the end of the apply flow, and the transport-error path behind
Preview import — which are about behaviour and want their own look.
* Let the Clockify apply and its aborted run finish before asserting
Both Clockify specs waited ten seconds on something that has to reach the
server and come back. The apply writes the plugin source and then pins the
release over the network, and the review button only unmounts once both have
landed; the transport-error path has to get the aborted run's refusal back
before the page can report it.
Neither is stuck, which is what made them read as behaviour failures. The
review button's count sits at 1 for the whole wait and then the assertion
gives up, and the page Playwright captures afterwards has no button and no
error alert on it: by then the apply had finished. Run on its own, with the
stock ten second budget, the apply spec passes in 53s. It is contention with
the other worker, not latency.
So both waits get 45s, and the file goes from three failures to one under the
same load that produced them.
* Let the Notion connection fail properly before reading the alert
Pressing Enter starts the installation rather than just validating it: the
connection is created against the server, and only when the secret write comes
back refused does the page say so. The spec read the alert ten seconds later.
That is a round trip, and it does not fit ten seconds while another worker is
running, which is every shard in CI. On its own the spec passes in 21s.
With this the whole file passes under the load that was failing it.
* New video
* Revise README for better package descriptions
* Update README.md for clarity and formatting
Corrected formatting and removed incomplete sentences in the README.
* Show a shared chatroom's messages to the person it was shared with (#1581)
Before: accepting an invite to a chatroom on someone else's drive opened the
thread, with its title, its place in the sidebar and its composer, and said
"No messages yet". The messages were there and the guest was allowed to read
them; they were simply never fetched.
After: the guest sees the conversation.
The message list comes from `useChatMessages`, which builds a collection query
for children of the chatroom with class Message. `CollectionBuilder` defaults a
query's drive scope to `store.getDrive()`, the drive the viewer currently has
selected, and the query index is keyed by drive. For the owner that is the same
drive the chatroom lives on, so it works; for a guest arriving through an invite
it is their own drive, so the query asked the wrong drive and got zero rows back
every time.
Instrumented against a local reproduction of `chatroom @smoke`: for the invited
agent the server's chatroom endpoint, which scopes from the chatroom's own
subject, found the one message, while the collection query the view actually
uses returned `total_items=0` for the same agent in the same run. With the scope
taken from the thread's subject it returns 1 and the spec passes.
How: `useChatMessages` passes the thread's subject as the collection's drive,
leaving the default in place while the subject is still the `unknown-subject`
placeholder. The hook also backs the comments panel, which had the same problem
on any resource shared from another drive.
* Build the local e2e server the way CI builds it
The script pointed at `target/debug/atomic-server` and told you to build with a
bare `cargo build -p atomic-s…
joepio
added a commit
that referenced
this pull request
Sep 22, 2026
…end guide) (#1552) * docs: tell the local-first story (URLs, Atomic Sync, Flutter) The docs still read as if AtomicServer were an HTTP origin you post commits to. This refresh sells what the code now does, without rewriting the spec pages. New pages: - urls.md: one map of every identifier shape (did:ad resource / agent / commit / blob / node, https, internal:, atomic://pair, ?drive= hints, HTTP aliases of a DID), placed early in the book. - local-first.md: data ownership as "the bytes are on your device", what the server is for now, and the honest trade-offs. - sync.md: the transport-agnostic sync model (signed commits, Loro, version vectors, outbox, rights on every transport) with a transport table. websockets.md and browser-peer-sync.md now sit under it as the wire-format reference and the WebRTC path. - flutter.md: the Dart SDK on top of atomic_lib via flutter_rust_bridge, its API groups, a minimal flow, and how it compares to the other clients. Edits: overview, motivation, when-to-use, Extended index and table, DID status line, API page (fetching did:ad over HTTP, /commit not /commits), tooling, rust-lib feature list, personal-data-store use case, Subject field in core concepts, README realtime bullet. Commits are described as the signed envelope around a write rather than the versioning mechanism. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018ZmptEWmdz92GXTuP77vrb * docs: local-first guide, server-optional Store docs, always-on framing Second pass on the docs refresh. - New end-to-end guide "Build a local-first app with @tomic/lib": identity from a keypair, Store and Drive with no server, OPFS persistence via the WASM client database, promoting the Drive to an always-on device, and pairing a phone through the Flutter SDK. Every call is one the library exposes today; where a step still needs a checkout (the WASM build) the guide says so. - AtomicServer chapter opens on "the always-on device" instead of on the feature list. - js.md and js-lib/store.md: serverUrl is optional, createDrive with localOnly, promoteLocalDrive, sync status, and did:ad subjects in getResource; save() no longer described as "commit to the server". - Solid comparison: local-first device store vs POD, DIDs vs WebID, CRDT sync vs document writes. - Astro guide: note that it is the HTTP path and where the local-first guide fits. - planning/dart-sdk-package.md: the plan for extracting the Dart SDK into a pub.dev package, referenced from flutter.md, tooling.md and the roadmap. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018ZmptEWmdz92GXTuP77vrb * Pin Rust builds and enforce Server/SaaS dependency alignment (#1499) * Pin Rust builds and enforce Server/SaaS dependency alignment * Run trusted CI orchestration directly on Mancave * fix(server,lib): security hygiene from the September 2026 audit (#1558) Six small items from planning/security-audit-2026-09.md (D, C16, F), each verified against current code and covered by a test, plus the C20 verification recorded in the audit doc. - CORS: any origin may still read, but Access-Control-Allow-Credentials is sent only to origins this server answers for (configured domain, base domain tenants, loopback, the desktop webview origins), by the same rule RequestContext uses for Host. server/src/cors.rs, gate outside the permissive Cors layer. - Client errors as client errors: ParseError maps to 400, malformed commit bodies are parse errors, signature and auth-header failures are typed Unauthorized (AtomicError::into_unauthorized) and the server keeps that type instead of stringifying it. 400/401 instead of 500, so they no longer reach Sentry as crashes. - C16: the importing flag and import source are a tokio task-local (ws_apply::import_scope) instead of process-wide statics; the live push loop skips the one peer a change came from via the DbEvent source_id, which now also covers COMMIT frames applied for a live peer. CommitIngestOpts::suppress_live_echo is removed. The Flutter listener reads the event's source_id instead of the flag. - collections::sort_resources is a total order (missing last, numbers numeric with NaN placed, numbers before strings); the old comparator panicked sort_by from 21 elements up on Rust 1.81+. - sync::peer: LIVE_PEERS recovers from a poisoned lock instead of unwrapping it. - /upload answers 400 on a multipart error instead of 200 with a partial list. - Drive-by for the clippy gate on stable 1.94: an unnecessary to_string in a lib/src/db/website.rs test, pre-existing on develop. Claude-Session: https://claude.ai/code/session_01H1U7TNdjwqXKv8MQYproiq Co-authored-by: Claude <noreply@anthropic.com> * planning: retire the #1500 session artifacts and reconcile the index against develop #1500 merged on 2026-09-17, so its two e2e handoffs, the website handoff and the start-server script leave planning/ per the folder's own rules. Their open items and durable lessons move to the owning plans: the website follow-ups and the inline-RTE flake diagnosis to website-publishing.md, the CI-versus-local worker topology to e2e-concurrency.md, the remaining red specs and the useDriveClass gap to e2e-diagnostic-hygiene.md, the harness traps (redb lock outliving the socket, desynced compiled i18n catalogs, a borrowed node_modules, the server-root fetch) to silent-failures.md, and the two rules every agent needs (.wuchale regeneration, HTTP-200 gating) to AGENTS.md. The README's "Biggest gaps" list keeps only open gaps, each checked against code: the query index dropping rows (ignored test in lib/src/db/test.rs) and store growth (full snapshot per commit, CLI-only compaction) are added, both with a PR in flight; the plugin model is on develop since #1500, so only code-first schemas (#1262) remain PR-bound. Status rows corrected: #1355 (npm job) and #1347 (MCP plan) merged 2026-09-18, #1500 merged, #1498 still a draft, and the addResource / CommitBuilder counts recounted (0 non-test call sites in data-browser, 31 non-test references all in browser/lib/src). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H1U7TNdjwqXKv8MQYproiq * Reconcilable host-to-Drive mappings, for hosted vanity subdomains (#1539) Adds Db::sync_drive_mappings, list_drive_mappings and managed_alias_hosts so a control plane that holds no user signing keys can install and reconcile host to Drive mappings without /bind-drive. Removal is scoped to the hosts a previous run installed, so the localhost entries and hand-made bindings survive. Mapping keys are normalized, since the Host header arrives in whatever case the client sent. A host bound to a missing Drive now 404s instead of falling through to the store root, which on a multi-tenant node answered one tenant's hostname with another namespace's content. --served-domain-suffix makes a server answer for *.<suffix> without taking on the rest of --base-domain, which would also change how subjects are normalized on a box that already holds customer data. * fix(lib): write datatype tags in the shared export path, after sealing the edit (#1553) Datatype tags are now written inside exportLoroDeltaInternal, after the user's ops are sealed, so a drained commit carries them and undo no longer steps back into the tag write. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KxExJcGw13DnHknGTMAhXY * perf(lib): coalesce cold-load OPFS reads into one worker round trip per tick (#1554) Store batches the render-phase local-db cache misses of a tick into a single getResourcesWithSnapshots call, chunked at 200 subjects, and shares an in-flight read between duplicate requests for one subject. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KxExJcGw13DnHknGTMAhXY * fix(lib,react): emit LocalChange from Resource.remove/push; route useArray.push through the save scheduler (#1541) Resource.push and Resource.remove now emit ResourceEvents.LocalChange like set does, so subscribed hooks re-render; useArray.push goes through the same save scheduler as the other writes; StoreContext no longer hands out a phantom Store when a provider is missing. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KxExJcGw13DnHknGTMAhXY * refactor(lib): one server-managed property list and one client-db serializer (#1542) Server-managed properties are declared once (server-managed-props.ts) and derived from that list everywhere, and the OPFS row is written through a single Resource.toClientDbJsonAd/Store.recordPersistedState path instead of two hand-rolled serializers that had drifted. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KxExJcGw13DnHknGTMAhXY * fix(lib): route destroy through the durable outbox so offline deletes survive (#1556) A destroy is signed into the durable outbox and drained like every other commit, instead of being POSTed straight to the server, so a delete made offline (or during a reload) is retried rather than lost. Incoming updates and local-db hydration skip a subject with a pending destroy. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KxExJcGw13DnHknGTMAhXY * fix(lib): apply destroy removal in the same transaction as the commit envelope (#1544) A signed destroy now removes the resource, its Loro snapshot, index and search rows and its cascade-deleted children in the same redb write that stores the envelope and commit row, with tombstones recorded only after that write lands. Malformed destroy commits return an error instead of panicking the request. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KxExJcGw13DnHknGTMAhXY * fix(lib): emit DbEvent::Destroyed for cascaded children only after the removal is applied (#1561) recursive_remove no longer announces anything; it collects each removed subject with its drive. remove_resource and the destroy branch of apply_commit announce them through Db::emit_destroyed once the removal has actually landed, so a subscriber never learns of a deletion that could still roll back. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KxExJcGw13DnHknGTMAhXY * fix(lib): classify commit refusals by error code, not message text (#1557) "Commits cannot be edited" now carries error_code::IMMUTABLE_COMMIT (10) on both transports, and the client decides whether a terminal drop is benign from the code, falling back to the legacy message match only when a response carries no recognised code. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KxExJcGw13DnHknGTMAhXY * fix(query-index): list every isA encoding, check the index against the store, scope DID rows by drive stamp A drive-scoped, sorted, class-filtered collection query returned fewer rows than the store held, with nothing to show for it (planning/silent-failures.md, "A query index silently disagreed with the data it indexes"). The ignored reproduction `is_a_encodings_all_match_the_class_constraint` found 2 of 4 rows. Two causes, both in the read that only this query shape takes: - Rows whose `isA` read back as a plain `String` were in the index and were read, then hidden: the built-in collection class extender did `is_a.to_subjects()?` on the value, and `Db::resolve_query_member` turned any extender error into "drop the row" — from the page and from the count. Class extenders now read `isA` through the same reference strings the index uses, and an extender that cannot decide is skipped with a `warn!` instead of hiding a row the reader may see. - Rows whose `isA` was a `String` holding the JSON array (the legacy Loro encoding of a `ResourceArray`) were never candidates: the index keyed them by the literal `["…"]`, and the matcher rejected them for the same reason. `Value::to_reference_index_strings` and `Value::contains_value` now read the array's elements from one helper, so index and matcher cannot disagree. What should have shouted now does: - `Db::check_query_index(&Query)` compares a query's member index with a scan of the store and reports the missing and stale subjects (`QueryIndexReport`). Used by the tests; a full scan, so not on the query path. - A first build of a filter cross-checks the equality constraints the planner did not scan, bounded by the planner's 512-entry scan cap, warns with the filter and the subjects it missed, and files them (`Db::cross_check_first_build`). The hot path is untouched. Security audit C17: a DID resource cannot be routed to a drive by its subject, so every drive's watched filters saw it on commit. The `Db` now resolves each watched filter's `drive` subject to its drive root (`filter_drive_roots`, filled in `register_watched_query` and `populate_watched_queries_cache`) and compares a DID row's `drive` stamp with that root on the commit path and in the index build. The filter's identity and encoding are unchanged, which is what broke the first attempt. A row without a stamp, or a filter whose subject is not stored here, is not excluded; read rights still apply at query time. Also: `lib/src/db/website.rs` test — `to_hex().to_string()` is `clippy::unnecessary_to_owned` on the current toolchain, which fails the workspace clippy gate on develop; one-line fix so the gate is green. Tests: the reproduction is un-ignored; `replicated_rows_reach_a_watched_scoped_sorted_query` (5 rows watched, 17 replicated through the sync import path, 22 from every shape), `first_build_cross_checks_the_unscanned_constraint`, `check_query_index_names_missing_and_stale_members`, `did_rows_stamped_into_another_drive_stay_out_of_a_watched_query`. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H1U7TNdjwqXKv8MQYproiq * feat(server,lib): startup store-size diagnostics and automatic redb compaction Opening the file store now logs its size (on disk and length) and open duration, a warning above 1 GiB naming `atomic-server compact`, and redb's `DatabaseStats` breakdown: allocated/leaf/branch pages, stored, metadata and fragmented bytes, and the reclaimable bytes (on disk minus pages in use, which is what `compact()` gives back; redb's sparse growth headroom does not count). When the file is at least 256 MiB on disk and at least 30% of it is dead, `Db::init_redb_file` compacts it on the handle it already holds, before any table is read and before the server listens, logs before/after sizes and duration, and keeps the outcome in `Tree::PluginMeta` (`Db::last_compaction`). Any failure is a warning; startup never fails on it. Not applied to the OPFS or sled backends. Server flags `--auto-compact` (`ATOMIC_AUTO_COMPACT`, default true), `--auto-compact-min-mb` (256) and `--auto-compact-min-reclaimable-percent` (30) build the `CompactionPolicy`; library callers pass one to `Db::init_redb_file_with_policy`. Documented in the installation guide. Test: a real redb file churned through the same open path (doubling overwrites plus throwaway resources deleted mid-file) is compacted on reopen, gives back most of the measured free space, reads every kept resource back, keeps the record across the next open, and is left alone with the policy disabled. Config tests cover the flags. Also fixes a pre-existing `unnecessary_to_owned` lint in `db::website` tests so `clippy -D warnings` passes. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H1U7TNdjwqXKv8MQYproiq * docs: address review on the overview and URLs pages - URLs and identifiers is Core, not Extended: the pointer now sits in the Core paragraph of the overview. - Tools list leads with Atomic Cloud (atomicserver.eu); self-hosting is one line without the docker incantation; the Tauri desktop and mobile apps are listed with where to download them. - urls.md: add localId (a name within a parent, used by imports and plugins with the atomic:… namespace convention), the atomic://open deep link next to atomic://pair, and a short list of strings that look like identifiers but are not (Loro commit origins, old placeholders). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018ZmptEWmdz92GXTuP77vrb * fix(data-browser): make page transitions opt-in until every browser works (#1563) (#1566) View transitions were on by default with only a user opt-out. They are still broken outside Chromium on desktop: the card-to-page morph blows up into an overlay on Firefox, and #1563 reports the same on Android. Flip the setting to opt-in. The stored key is new (`viewTransitionsEnabled`, default false) rather than a flipped default on `viewTransitionsDisabled`: that key already holds `false` for anyone who toggled the old checkbox back off, so reusing it would have left transitions on for exactly the people who had noticed them misbehaving. The transition machinery and CSS are untouched, so re-enabling by default is a one-line change once each browser is verified. - Settings > Accessibility now reads "Enable page transition animations". - Playwright config and fixtures seed the new key. - New test pins both sides of the default in useNavigateWithTransition. Claude-Session: https://claude.ai/code/session_01Pxhq3hXh8AgAk4NZzNdnkR Co-authored-by: Claude <noreply@anthropic.com> * Default the navbar to the bottom on touch devices (#1567) The navbar started at the top everywhere. On a phone or tablet that puts the sidebar toggle, the search button and the context menu out of thumb reach, which matters more than a navbar that sits in the same place on every device. The default now keys on a touch-primary pointer rather than viewport width: a landscape tablet is as wide as a laptop, and a desktop window dragged narrow is still mouse-operated. New `isTouchPrimary()` helper, which `useHover` now shares instead of inlining the same media query. Only the default changes. `useLocalStorage` falls back to it solely when the key is absent, so a position picked in Settings is preserved. Claude-Session: https://claude.ai/code/session_01VT2jpYQ73mPfiueW25ty3y Co-authored-by: Claude <noreply@anthropic.com> * Keep toasts clear of the bottom navbar, drop the old sidebar spacer (#1570) Toasts stack from the bottom right, which is exactly where the navbar sits when it is at the bottom, so a toast covered the search button and the context menu. The toast container is now lifted by the navbar's height in that case; with the navbar at the top nothing changes. The sidebar's OverlapSpacer goes with it. It added 3.8rem beneath the App panel below 950px with a bottom navbar, dating from the floating navbar mode. The overlap it guarded cannot happen: SideBarWrapper's height is the viewport minus the navbar, so the bar never covers the sidebar. * docs: rewrite the feature list (#1572) Every line is one short clause. Ordered core-first: local-first, then real-time collaboration and presence, then the things you make with it, then the platform underneath. Adds the features that were missing: websites, the kanban / calendar / dashboard / timer views, table templates, the virtual drive, apps, plugins and integrations, meetings and canvas. Drops the passkey recovery line, folds invites into authorization and the serialization formats into the API line. docs/src/atomic-server.md carries the same table and is updated with it. * Converge JS and WASM plugins: one host, one manifest, one install path (#1546) (#1571) * Generalize PluginRelease to both runtimes and add Release, Installation and Listing classes * Expose validate_plugin_zip so publish can check a package before storing it * Add one host core for sandboxed plugins One HostCore carries the effective grant (caller, never sudo, bounded by the installation identity), the fetch policy, secret substitution, the pinned egress check and a streaming response cap, plus get_config, get_plugin_agent and commit. commit is refused for proposal-only (extension world) installations. One shared wasmtime Engine and one fuel/memory policy keyed by runtime and capability live here too. * Make the JS StoreHost a thin delegate to HostCore Reads and queries now go through the shared grant, fetch through the shared policy, and the runtime uses the shared engine and limits. The PluginHost trait moves to host_core and is re-exported here. * Make the wasm class-extender host a thin delegate to HostCore The host impl only translates WIT types now. The server WIT copy is synced with atomic-plugin/wit so get-plugin-agent, which the guest SDK already exposes, is served. * Install both runtimes through an Installation commit and publish wasip2 releases from zip bytes * Add the plugin runtime convergence plan for #1546 * Add manifest schemaVersion 2 shared by JS and WASM plugins One versioned manifest carries runtime, world, entrypoints, capabilities and network origins. v1 upgrades with defaults and serializes unchanged, so existing release ids are stable. plugin.json translates at the boundary. Fixtures under testdata/plugin-manifest are shared with the TS mirror. * Mirror manifest v2 parsing in @tomic/lib against the shared fixtures * Document manifest v2 and mark plugin.json as the translated legacy form * Add Release, Installation and Listing to the browser server ontology * Add installRelease, publishZipRelease and the install review reader to @tomic/lib * Publish uploaded zips as releases and install them through the Installation review * Install Store listings through the Installation review, keeping drafts as a secondary action * Add the Installation view with pause, revoke and uninstall, and list installations on the drive * Cover the Installation review screen in the plugin and Store e2e specs * Fix padding-line lint in the manifest v2 mirror * Derive plugin resource limits from v2 capabilities for both hosts * Publish wasip2 releases with a v2 manifest and enforce the Installation grant set * Serve wasip2 releases and their zip from /plugin-package * Record every published release as a Release resource and serve private releases to its readers * Record the plugin convergence implementation checkpoint * Treat schemaVersion 2 as valid in the app package malformed-release test * Box the manifest in FetchPolicy to keep the enum small * browser: remove the legacy Plugin view and zip reader Legacy Plugin resources are migrated into Installations by the server at startup, so the browser no longer renders or manages class Plugin: PluginPage, PluginCard, UpdatePluginButton, useCreatePlugin (uninstall/update) and the client-side zip reader go. The server validates the uploaded zip and translates plugin.json, so NewPluginButton no longer inspects it. AssignRights, PermissionRow and ConfigReference move under views/Installation, their only user. PluginPermissions and pluginUtils fold into a shared CapabilityList (extracted from InstallationReviewDialog) that also renders the pluginPermissions the server writes on an Installation. InstallationPage shows the manifest author the server fills in. * browser: drive plugin list reads Installations only The drive's plugins property was the legacy Plugin list; every plugin is an Installation found by class under the drive now. * lib: drop unused manifest and install exports resolveManifest/ResolvedManifest had no caller outside its test; the server resolves defaults, the browser validates and reviews. Also removes the unused RUNTIME_WASIP2, WORLD_SERVER_EXTENSION and INSTALLATION_STATUSES constants and un-exports the two helper types nothing imported. The Store compares against RUNTIME_JS instead of a literal. * e2e: align plugin spec with the Installation page * Replace the legacy Plugin install hook with a startup migration and one manifest per plugin PluginMeta now holds a single unified manifest (JSON-encoded records; legacy MessagePack records upgrade on read and are translated from plugin.json the first time the plugin loads). Legacy Plugin + pluginFile resources are rewritten at startup into active Installations pinned to a Release published from their zip, keeping subject and on-disk files. Re-activating an already materialized release no longer re-extracts. * Replace the KV plugin catalog with Listing resources and collapse the publish paths A public publish now records a publicly readable Listing at <server>/listings/<id> next to the Release; /plugin-catalog is a class query over Listings and answers with the Listing fields plus release URL, id, runtime and world. The three publish handlers share release::publish_release for recording, listing and world checks. CatalogEntry and the plugin-catalog/v1 keys are gone; the KV release store stays as the cache it is. * Document the Installation-only plugin flow and record the cleanup checkpoint Docs describe Release, Listing and Installation, the Store and zip-upload paths that both end in the Installation review, and how legacy Plugin resources are migrated. The planning checkpoint lists what this cleanup removed and what is still open. * Read the flat Listing shape from /plugin-catalog in the Store * Pin a materialized install to its release id, and make pausing stop a plugin Two things an Installation got wrong once it was the single install path. An activation decided it had nothing to materialize by comparing the installed manifest with the release's. Two releases of one plugin agree on every manifest field and differ in their package bytes, so repointing an Installation at a new release skipped the extraction and left the previous code serving under the new release's id. `PluginMeta` now records the release whose code is on disk and that is what the check compares; a record from before the field existed says nothing about the bytes, so it materializes again. The legacy migration records it too, so it still re-extracts nothing. `paused` and `draft` did nothing at all. The hook handled `active` and `revoked` and fell through for the rest, and no runtime path read the status, so Pause wrote a property while the plugin kept serving its hooks. Pausing now unregisters the class extender, which is not an uninstall: the files and `PluginMeta` stay, so resuming re-registers the same agent rather than minting a new identity. `installation::resolve` refuses anything but `active`, which stops JS runs, and the startup loader skips extenders whose Installation is not active, so a pause survives a restart. The wasip2 install test asserted the old behaviour, that pausing and resuming leaves the extender registered. It now asserts what pausing is for, and the 'must not re-extract' case it was really about moved to a config change, which is where it happens in practice. * Update an installed plugin in place, rather than uninstalling it first The Installation page showed its pinned release read-only and the only way to get newer code onto a drive was Uninstall then install again. That retires the plugin's agent, so everything it owned stopped being its own, and a second install of the same namespace and name is refused while the first one is there. Uploading a zip on the Installation page now publishes it as a private release, reviews it through the dialog the Store and the zip upload already use, and repoints this same resource at it. The subject, the config and the agent stay; only the release, the grants and the version change, and they change in one commit, because the server checks the grants against the new release's manifest and would refuse a release that arrived on its own. The review dialog takes the wording as a prop and starts its config editor from the config in use rather than the release default, since for an update that is what the installer is deciding about. A zip for a different plugin is refused in the page with the names in the message, ahead of the server refusing it on the identifier comparison. * Record the materialized release from the install hook, and cover the encoding Which release's code is on disk is the Installation's business, not the zip extractor's, so `activate` writes it after `install_or_update_plugin` returns rather than passing it down. A failed extraction then leaves no id, and no id means materialize again, which is the safe direction. It also keeps that function at seven arguments, so clippy gains nothing new. `PluginMeta`'s encoding test constructed the struct literally and so did not compile. It now round-trips the id, and a second test pins the compatibility that matters: a record written before the field existed reads back as `None` rather than as some release, and is written out without the key. * Format the touched browser files with oxfmt The browser workspace formats with oxfmt, not prettier, and `pnpm lint` runs `oxfmt --check` in each package. Four files had been written by prettier, which reformatted them whole, so `lib lint` and `data-browser lint` both failed on the format check. Running the repo's own formatter over them puts them back, and the net change against the previous head is once again only the lines this branch actually adds. `pnpm run -r lint` now exits 0 across the workspace. * fix(lib,react): make "destroyed" one fact on the store (#1573) * fix(lib,react): make "destroyed" one fact on the store Deleting a resource left a "Resource with error" row in the sidebar. A `parent=` query keeps answering with a destroyed child for a moment: the answer was computed before the destroy landed, and `Collection.fetchPage` deliberately races the local index against the server's `/query`. Every reader of such an answer that had not itself witnessed the delete put the row back, and that is not cosmetic — rendering the row asks the store for the resource, which re-creates the entry the destroy had removed and then 404s. The knowledge that a subject was deleted only ever lived in caches scoped to one instance: `hasPendingDestroy` (an outbox entry, dropped the moment the server acks), `Collection._removedSubjects` and `useChildren`'s `removedRef`. A reader created after the delete — a second view on the same query, a folder re-expanded — had none of them, which is why the row came back. The store now owns it, as the in-memory half of the tombstone `removeResource` already wrote to the local database: - `Store.isDestroyed` / `clearDestroyed`, backed by an insertion-ordered set capped at 10k subjects. - `Collection` and `useChildren` read it; their own sets are gone. A destroyed subject is never a member, whatever state arrives for it, and is kept out of `_queriedMembers` so a lifted tombstone can re-admit it. - `applyIncoming` refuses state for a destroyed subject after the ack as well as before it. - `getResourceLoading` answers for one from the tombstone instead of firing a request that can only 404. Only creating a resource again under the same subject lifts a tombstone (`newResource`); a stale answer no longer can. That is a behaviour change: `applyResourceChange` used to treat any matching resource event as a restore, which is the resurrection this fixes. * test(data-browser): give the Markdown schema suite its own timeout fileContentsToTiptapJson awaits a lazily imported collaborative Markdown schema. Whichever of the two tests in that suite runs first pays the import cost, which exceeds vitest's 5s default whenever the test threads are oversubscribed. On the shared runner that is routine: a recent run reported 243s of import time for 86s of wall time, and the suite timed out there twice in a row while passing locally every time. Give the suite the same 60s budget the recovery-enrollment suite already uses for the same reason. * Create a drive's plugin schema terms in one round, not nineteen (#1574) * Create a drive's plugin schema terms in one round, not nineteen Asking for a new plugin on a drive that has no plugin schema yet means waiting for `ensureSchema` to write all nineteen terms of `pluginSchema()` first, and `ensureAll` saved them one after another. Sixteen properties and three classes, each its own awaited round trip, is most of the time between clicking "New plugin" and the page appearing, and it is why `plugins.spec.ts` "a plugin proposes changes, and nothing is written until you approve" times out waiting for the New plugin heading on a cold drive. An earlier pass already batched the orphan lookups; the saves were still sequential, which is the rest of the wait. The terms are independent: they write different resources under the same ontology and none of them reads another's subject, so the order they land in does not matter. Resolve every spec first, collect the ones that have to be created, and send those creates together. The ontology's own list is still assembled in spec order, because that order is what someone reading the ontology sees. * Note the plugin schema batching in the browser changelog * Give the Rust checks container the plugin zip its tests include `server/src/plugins/plugin.rs` and `server/src/handlers/plugin_release_test.rs` `include_bytes!` `browser/e2e/tests/fixtures/test-plugin.zip`, but `rustChecksContainer` mounts only the crate directories and `testdata`, so `/code/browser` does not exist and both test targets fail to compile: error: couldn't read `server/src/handlers/../../../browser/e2e/tests/fixtures/test-plugin.zip` Mount the one file. Nothing else under `browser/` belongs in that container, and mounting the tree would make every front-end edit invalidate the Rust layer. This is the same treatment `testdata/pairing-request.json` already gets a few lines above, for the same reason: a fixture both the Rust and the JavaScript side test against. The two call sites are behind `cfg(all(test, feature = "wasm-plugins"))`, so only the checks container needs it; the build containers are unaffected. * Stop filing every server error in Sentry twice (#1576) * Stop filing every server error in Sentry twice Two things reported the same failure. `sentry_actix` captures a handler's 5xx with the request attached, and `tracing_actix_web` separately logs "Error encountered while processing the incoming HTTP request" at `error!`, which the Sentry tracing layer turned into a second event. The staging floods of September showed it plainly: the issue groups pair up, 6313 events on one side and 6312 on the other, for the same two incidents. On a free-tier quota that is half the budget spent on duplicates. Give the tracing layer an event filter that ignores the `tracing_actix_web` target and delegates everything else to `default_event_filter`, so background work still reports as before. The log line is untouched; only the duplicate report goes away. Also stop the data-browser reporting from a dev server. Vite's hot reload throws while swapping modules ("_s is not a function", "Cannot access X before initialization", all with `@react-refresh` frames), which is not a defect in anything shipped, and those arrived rated above every real bug in the backlog. Setting VITE_SENTRY_ENVIRONMENT still reports from a local build on purpose. * Add changelog entries for the Sentry reporting fixes * Change video link in README Updated video link to a private user image. * Tell people the assistant can build apps and websites (#1577) * Tell people the assistant can build apps and websites Before: the new-resource page listed tables, documents, dashboards and the rest, but nothing on it said apps or websites were possible. Both classes are minted per drive, so they only appear under "Your resource types" once the drive already has one — which never happens to someone who was never told. After: a row of suggestions sits under the AI composer. Clicking one puts a half-written request in the composer with the caret at the end, so the user finishes the sentence instead of sending "build me an app" and being asked what they meant. The row hides once they write their own words, because setting a textarea from code does not go on the browser's undo stack. How: AI_BUILD_SUGGESTIONS in the creation catalog, rendered as a gradient- bordered row in NewRoute under the composer. Unit tests cover the seeds and the hide rule; an e2e test covers clicking, swapping and hiding. * Honour "Enable AI Features" on the new-resource page Before: unchecking Enable AI Features hid the navbar button and the AI chat entry, but left the whole "Build with AI" composer and the Ask AI button in the search field on the page. Submitting either one silently ticked the setting back on, so the only way to get rid of the composer was to not press it. After: the composer, its suggestions and the Ask AI shortcut are gone while the setting is off, and nothing turns it back on behind the user's back. The blank buttons, templates and upload area stay, since none of them are AI features. The no-matches line stops offering the assistant too. How: the section and the Ask AI button render on enableAI, and both setEnableAI(true) calls are dropped now that nothing reaching the assistant is rendered while it is off. An e2e test sets the stored setting and checks what remains. * Make the build suggestions look like the class buttons Before: the suggestions were pills with a sparkle icon and labels written as "An app", "A website". They read as a different kind of control from the blank class buttons right below them, which they are not. After: App, Website, Dashboard, Custom table, on the same button as the blank ones, with a faint rainbow border in place of the grey one. The border comes up to full colour on hover, focus and while a suggestion is in the composer, so it still reads as the assistant's row rather than a second Start blank. How: Suggestion extends BasicChoice and paints its border with a two-layer background-image, so the geometry can never drift from the blank buttons. The sparkle is gone. * Give the build suggestions their class icons Before: the suggestions were the only buttons on the page with a bare label. Next to a row of blank buttons that each carry their class's icon, they read as a different kind of thing. After: each one wears the icon of the class it makes. Dashboard and Custom table take theirs from the class subject, so they are literally the same icon as the blank Dashboard and Table buttons. Website takes the globe already mapped to website-project. App gets the launcher grid, added to the icon map so an app shows the same glyph wherever it appears. How: a suggestion names either a class subject or a drive-minted class shortname, and the row resolves it through getIconForClass like every other class button does. The puzzle piece a new app carries as its placeholder emoji was the obvious choice for App and is deliberately not used: plugins and installations already have it, and an app sits beside both. * Close three gaps the plugin convergence left behind (#1575) A publish refused for claiming the wrong world had already stored the package bytes and cached the release record: the check ran in the handler, after `release::publish_package` had written both. The claim now goes into that function, which asks straight after reading the manifest and before the first write, so a refusal leaves nothing on the node. `expect_world` reads the world from the manifest rather than from the release, because at that point there is no release to read it from. `installation_grants` fell back to the capabilities the manifest declares whenever any step of finding the Installation's approved grants failed. That is the one direction a fallback must not take: the declared set is what the plugin asked for, so a lookup going wrong rewarded asking for more than was approved. Only the absence of an Installation, which is a legacy draft that never had grants, still reads the manifest. An Installation that cannot be read, or whose grants cannot be, now grants nothing and logs why. The class check also went through `Value::to_subjects`, which errors on the scalar `isA` encodings, so an encoding alone could hand over the declared set; `Resource::class_subjects` and `Resource::has_class` are now the one encoding-tolerant reader, shared with `ClassExtender`. An Installation's `release` is declared an atomicURL and held bare `blake3:` ids, which meant every writer set it with datatype validation switched off. Every publish records a `Release` resource and the response carries its subject; the browser's two zip paths were throwing that away and storing the id. `publishZipRelease` returns `subject` and the callers pass it, so the property holds what it is declared to hold and nothing skips validation. `release::resolve` still accepts a bare id, because Installations written before this carry one and have to keep working. One consequence worth knowing: an Installation whose `release` is a URL is resolved with a read check that a bare id skipped. `/releases/<id>` is keyed on the content hash but parented to whichever drive published first, so an agent publishing byte-identical bytes on a second drive gets the first drive's Release resource back and may not be able to read it. They now get a rights error at install time rather than installing from a record they cannot see. That a content-addressed subject is parented to one drive is a deeper wart, left alone here. * Stop a stale auth proof from failing requests that needed no auth (#1578) A browser keeps its authentication proof in the `atomic_session` cookie. `setCookieAuthentication` signed one proof and stored it for a day, and `checkAuthenticationCookie` only asked whether a cookie existed, so nothing ever re-signed it. Since `AUTH_MAX_AGE_MS` landed in 0.41 the server refuses a proof older than five minutes, which made a stale proof the ordinary state of any tab left open, and every request such a tab made was answered 401 -- public ones included. On staging one tab polling the public `GET /server` endpoint produced 4,215 rejections in a row, every one of them carrying the same `signed at` timestamp while the server's clock walked past the window. Two halves: - The browser refreshes before it goes stale. The cookie now lives no longer than the proof it carries, and the freshness check reads the signed timestamp out of the cookie instead of only looking for the cookie's presence, so the request path installs a fresh proof well inside the window. - The server treats an aged-out proof as no proof. An HTTP request, and the headers a socket is opened with, did not ask to be authenticated; a proof that has expired says the caller *was* this agent and is no longer, which is not by itself a reason to refuse. They continue as the public agent and the rights check decides. A signature that does not verify is still refused outright, and so is a stale proof in an `AUTH` frame or a peer handshake, where the caller asked to be authenticated and is owed the answer. * Give the plugin specs a budget that matches what creating a plugin costs (#1579) Several tests in plugins.spec.ts have been failing on develop since 16 September, all at the same point: they click `New plugin` and the page never arrives inside the suite's 10s `expect` budget. It is not a hang. The first plugin on a drive materializes that drive's plugin schema first, and `pluginSchema()` is nineteen terms, each its own resource and its own commit. The browser already sends those writes together, but the server applies commits one at a time: instrumented, all sixteen property saves start within 30ms of each other and finish 130 to 150ms apart, so the cost is additive whatever the client does. Measured against a debug build, the step takes 5.2s on a fresh store and 9.4 to 11.3s once the store holds a handful of drives, which is where this spec runs in a full-suite pass. That is why it passes alone on a laptop and fails in CI. So the wait gets a budget that matches the work, here rather than for the whole suite, and the test gets room for what comes after it. The two places that had the helper's steps copied inline now call the helper, so there is one budget rather than three. This is the budget, not the cost. Nineteen sequential writes to open a page is its own problem and is being looked at separately. * Format the tests that came with the stale-auth-proof fix (#1580) `cargo fmt --all --check` fails on develop's head in three places, all in tests added by #1578: the `sign_message` call in `signed_at`, the `assert_eq!` on `error.error_type`, and the `get_client_agent` binding in `a_cookie_whose_proof_aged_out_is_the_public_agent`. Nothing caught it because #1578's own develop run was concurrency-held behind an earlier push and then cancelled before it ran, so the format gate never executed against that commit. The cost is larger than one red check. The gate runs first and dies about twenty seconds in, before any test starts, so develop and every branch that merges develop produce no test information at all until this lands. Run 4180 on a feature branch already failed this way. This is `cargo fmt --all` output and nothing else: line wrapping plus one trailing comma, no behaviour change. * Put back 184 UI strings that render blank in a production build (#1583) Before: a good deal of the app's text was simply not there once built. The page transition toggle in Settings was an unlabelled checkbox. So were both plugin visibility toggles on the Integrations page. The Installation review had no "Release" tab to click, the new-resource page's assistant suggestions had no names, and 180-odd other strings were missing the same way. After: they render. Verified by rebuilding the frontend and running the specs that were failing on the missing text. Why it happened: `wuchale` compiles the .po catalogs into the bundle at build time and does not extract during a production build, so any string added without re-running extraction has no entry, and a message with no entry renders as nothing. The catalogs had drifted 184 strings behind the source. Two things here. The catalogs are regenerated, which also drops 40 msgids for text that no longer exists, most of them fragments of sentences that were since rewritten as one. And `wuchale.config.js` now points the extractor at the integration screens as well. That second part matters: Connect GitHub, Google Calendar, Notion and the other connect screens are React components inside the pinned `devonian` package, aliased in as `@localthought/atomic-integrations`. The adapter's default `src/**` never reaches them, so a regeneration without that line would have taken 22 strings those screens still use out of the catalog and blanked them, which is how this was found. Three e2e specs were failing on the blank text, and pass with it back: `settings.spec.ts:7`, `integration-store-install.spec.ts:14` and `integration-visibility.spec.ts:119`, along with both `new-resource-catalog` failures. `settings.spec.ts` also asked for the toggle by its old name; #1566 turned that setting into an opt-in and renamed it, so it now matches on the words the setting's own search keywords use. This is worth a guard of its own: nothing today notices when the catalog falls behind, and the symptom is invisible to everything except a person looking at the screen. * Answer the integration spec's catalog mock in the shape the endpoint uses The "integration categories default off" spec stubbed /plugin-catalog with `{ metadata: { ... }, verification }`, which is the publish payload, not the catalog one. `plugin_release::catalog` answers with one flat object per Listing: name, description, publisher, domains, standards, release, releaseId. Reading `entry.domains` off the nested object gave undefined, and spreading it in the store's search filter threw "domains is not iterable" during render, so the error boundary replaced the whole integrations page. The spec then failed on the first thing it looked for after the reload, `[data-integration="mt940"]`, which made it read like a bundled-integration or a preference-persistence bug rather than a bad mock. * Fix twenty-five e2e failures, and two real bugs hiding among them (#1582) * Fix seventeen e2e failures that were never the code under test Before: develop's e2e reported 38 failures. Seventeen of them had nothing to do with the features they name; they were the test setup asking for things the pipeline does not provide. After: all seventeen pass. Verified locally against a built bundle and a real server, the same shape CI runs. Three separate causes. Eleven specs loaded app modules by source path. Seven website specs plus installation-recovery ran `await import('/src/chunks/Website/websiteModel.ts')` (and `'/@fs' + resolve(...)` for `browser/lib/src/ws-v2.ts`) inside `page.evaluate`. Vite's dev server resolves those; the bundle atomic-server embeds does not, so every one of them died on "Failed to fetch dynamically imported module". They were added on 15 September against a dev server and have never passed in CI; `pnpm typecheck` in browser/e2e has been reporting them as TS2307 the whole time. An E2E build now hangs those modules on `window.atomicE2E` through the new `helpers/e2eModules.ts`, attached behind `isE2E()` so a release build neither exposes the registry nor pulls the chunks into its entry graph. Their types are written out in `browser/e2e/global.d.ts`: the e2e project cannot import data-browser's own types without dragging the whole app graph across its rootDir. Website publishing was never switched on for the test server. Without `ATOMIC_WEBSITE_ORIGIN` the server answers every hosting call with "Website hosting is disabled", which surfaces two different ways: `website-versions` and `website.spec.ts` throw out of `hostingRequest`, and `website-inline-rte` and `website-errors` fail on a click, because the resulting toast renders over the preview iframe and intercepts pointer events. `http://sites.localhost:9883` satisfies the separate-domain rule the server enforces, and is set for both the Dagger e2e service and the local script. The apps specs hit the wall #1579 just fixed for plugins. `createApp` calls `ensureSchema(pluginSchema())`, the same nineteen sequential writes, against the suite's 10s budget. Same fix in the same shape: one `newApp` helper with a budget that matches the work, replacing five copies of the same three lines. One more, found on the way: `newResource` looked for a class button by name anywhere under `main`, and #1577 added an assistant suggestion row carrying some of the same words. Two buttons named "Dashboard" is a strict mode violation, which is what `dashboard.spec.ts:287` was failing on. The search is now scoped to the page's sections, where the class buttons live. Not fixed here, and still failing: website-preview-stability, which is a real bug rather than plumbing, since opening "AI edit" tears down the website preview the test asserts must survive. Also integration-workspace, which does not merely import a source path but fetches the module's text and regexes another path out of it. * Fix six more e2e failures, two of them real crashes The integration cluster on develop was three separate things wearing the same clothes. **The mock proxy was at an address the browser cannot reach.** It runs inside the atomic-server container, so from the server's process it is on 127.0.0.1:19090 -- and the e2e build handed that same value to the browser, which runs in the playwright container, where nothing answers there. Every page that lists integrations then carried a second `role="alert"` reading "TypeError: Failed to fetch", so the specs asserting on an alert read the wrong one or failed strict mode. Pointing the bundle at `http://atomic.localhost:19090` reaches the same container through the host mapping chromium is already given, and the server keeps its own loopback URL, which is right for it. Nothing server-side reads a configured proxy origin, and the origin a LocalThought connection stores is used only by the browser, so this value is the browser's alone. The per-test forwarding route in the Pets spec was a workaround for exactly this and goes with it. **Opening an AI chat could take the page down.** `useEditor` replaces the tiptap editor when its dependencies change and destroys the old one, which nulls its `commandManager` while its last state stays readable. So `editor.isEmpty` still answers and `editor.commands` throws. A prefill arriving in that same breath -- which is what "New automation" and "AI edit" do -- left the user at the error boundary with "Cannot read properties of null (reading 'commands')". Both effects now check `isDestroyed`, which also covers an editor whose view has not mounted yet. The prefill is not marked as applied when it is skipped, so the new editor still receives it. **A release is reviewed before it is used.** The store card offers "Open", which raises the installation review, and the draft is one of the choices there. The spec still clicked a "Create draft" button on the card itself. Two specs were reaching into the app by source path, which only a vite dev server can answer: one imported `githubInstaller.ts`, read the served text and regexed the provider bundle's path out of it, the other imported `runScript.ts`. Both now take the modules from the `window.atomicE2E` registry, which gains the installer and `@tomic/lib`. Last, `integration-visibility` was sending `/plugin-catalog` the shape it had before the catalog was flattened. `IntegrationStore` spreads `entry.domains`, so an entry without one took the Integrations page to its error boundary rather than failing an assertion. Verified against a locally built bundle and a real server: all five integration specs and `plugins.spec.ts:121` pass, and Notion and Clockify pass as soon as the proxy is reachable and fail exactly as CI does when it is not. * Put a plugin's release in the commit that creates its Installation Installing a plugin has been refused since the Installation class started requiring `release`: the server answers the commit with "Property .../properties/release missing. Is required in class Installation", the outbox drops it as terminal, and nothing is installed. Both paths are affected, the integration store and a zip upload, since both go through `installRelease`. `store.newResource` signs the genesis commit from the propvals it is given, right there at creation, and holds it on the resource until `save()` moves it into the outbox. A property set between those two calls lands in a later commit, so the genesis commit the server validates is missing it. `release` was set that way. So it moves into the propvals, where the rest of the Installation's fields already are, and the separate `set` goes. Verified against a locally built bundle and a real server: the install now gets through the review dialog and past the commit. The spec still fails here, on a fetch of a class hosted at atomicdata.dev that this sandbox has no route to, which CI does. * Note the plugin install and AI chat fixes in the browser changelog * Check the e2e port before wiping the store, and offer to run the mock proxy Two ways the local e2e server told you the product was broken when it was not. The store was wiped before the port was checked. With an older server still listening, the script printed "Something is already listening" and exited, leaving that process serving a store which had just been deleted underneath it. Runs against that look ordinary and mean nothing: new-resource-catalog lost its search box entirely, which reads like a UI regression and is not one. The port is now checked first, nothing is wiped when it fails, and the message says how to find the stale process. It also warns against `pkill -f atomic-server`, which matches the caller's own command line and kills the shell that ran it, so a restart chained after it silently never happens. The integration specs need an integration proxy, which CI runs beside the server and this script did not run at all. Without it Notion, Clockify and GitHub raise "TypeError: Failed to fetch", and that second role="alert" breaks their own strict-mode alert assertions, so four specs fail for a reason that has nothing to do with them. plugins.spec.ts:1270, red on develop across two CI runs and recorded as a product bug in the automation toggle, is green as soon as the browser can reach a proxy. `--mock-proxy` runs it. It waits for /catalog to answer rather than assuming a started process is a working one, since a proxy that failed to bind produces exactly the silent wrong results this is meant to prevent, and it warns when the built bundle does not mention the port, because VITE_INTEGRATION_PROXY_URL is read at build time and nothing the script does afterwards can rescue a bundle built without it. It gives the server no proxy environment of its own. Passing what .dagger gives its e2e service took plugins.spec.ts from two failures to four, adding :109 and :514, neither of which touches an integration provider. Those variables belong to a container this is not. * Give the plugin sandbox round trips a budget they can actually meet Three of the plugin specs waited on a real trip through a real sandbox with the suite's 10s action budget. The browser posts to the server, the server starts a sandbox, and the plugin's discover phase answers out of it; none of that fits 10s on a loaded box. Measured here, each of the three passes on its own and times out on exactly that step when the suite runs it beside another, which is what CI does on every shard. So the waits are widened where the sandbox is, not for the suite: - GitHub install. When the URL the test waits for appears, the install is still running: the button reads "Connecting…" and is disabled, so "Connections" does not exist yet. The click succeeds at 45s and the test takes 48s, against a 60s per-test default with nothing left over, so the test gets 120s too. This one now passes in the full file. - Both Clockify workspace discoveries. Same shape as the wait `newApp` documents in apps.spec.ts: the budget was never achievable on this hardware, and each assertion is about what came back, not about how fast it came. This does not make the two Clockify specs green. It moves them off the discovery step they never got past and on to later assertions — "skips repeats" at the end of the apply flow, and the transport-error path behind Preview import — which are about behaviour and want their own look. * Let the Clockify apply and its aborted run finish before asserting Both Clockify specs waited ten seconds on something that has to reach the server and come back. The apply writes the plugin source and then pins the release over the network, and the review button only unmounts once both have landed; the transport-error path has to get the aborted run's refusal back before the page can report it. Neither is stuck, which is what made them read as behaviour failures. The review button's count sits at 1 for the whole wait and then the assertion gives up, and the page Playwright captures afterwards has no button and no error alert on it: by then the apply had finished. Run on its own, with the stock ten second budget, the apply spec passes in 53s. It is contention with the other worker, not latency. So both waits get 45s, and the file goes from three failures to one under the same load that produced them. * Let the Notion connection fail properly before reading the alert Pressing Enter starts the installation rather than just validating it: the connection is created against the server, and only when the secret write comes back refused does the page say so. The spec read the alert ten seconds later. That is a round trip, and it does not fit ten seconds while another worker is running, which is every shard in CI. On its own the spec passes in 21s. With this the whole file passes under the load that was failing it. * New video * Revise README for better package descriptions * Update README.md for clarity and formatting Corrected formatting and removed incomplete sentences in the README. * Show a shared chatroom's messages to the person it was shared with (#1581) Before: accepting an invite to a chatroom on someone else's drive opened the thread, with its title, its place in the sidebar and its composer, and said "No messages yet". The messages were there and the guest was allowed to read them; they were simply never fetched. After: the guest sees the conversation. The message list comes from `useChatMessages`, which builds a collection query for children of the chatroom with class Message. `CollectionBuilder` defaults a query's drive scope to `store.getDrive()`, the drive the viewer currently has selected, and the query index is keyed by drive. For the owner that is the same drive the chatroom lives on, so it works; for a guest arriving through an invite it is their own drive, so the query asked the wrong drive and got zero rows back every time. Instrumented against a local reproduction of `chatroom @smoke`: for the invited agent the server's chatroom endpoint, which scopes from the chatroom's own subject, found the one message, while the collection query the view actually uses returned `total_items=0` for the same agent in the same run. With the scope taken from the thread's subject it returns 1 and the spec passes. How: `useChatMessages` passes the thread's subject as the collection's drive, leaving the default in place while the subject is still the `unknown-subject` placeholder. The hook also backs the comments panel, which had the same problem on any resource shared from another drive. * Build the local e2e server the way CI builds it The script pointed at `target/debug/atomic-server` and told you to build with a bare `cargo build -p atomic-server`. CI does neither: `.dagger` builds the e2e server with `--profile e2e`, which the workspace Cargo.toml defines as release brought down to opt-level 2 with debug assertions and overflow checks kept on. The comment on that profile says what it is for, "an optimisation level where commit round-trips stop dominating". opt-level 0 is not a slower version of the same thing. With two Playwright workers the app stops being able to boot: specs across new-resource-catalog fail on a page whose entire snapshot is `img "AtomicServer"`, one of them after a 30s waitForURL, and the same file at --workers=1 passes 7 of 7 every time. Read as test output those look like behaviour bugs, or like CI shard contention. They are neither, and a timing number taken there measures the binary rather than the suite. A green on the slow build still means something, since a spec that passes at opt-level 0 passes at opt-level 2. A red means nothing at all. That asymmetry is the whole reason this matters: the failures are the ones you are reading. So the binary is resolved from target/e2e when it exists, ATOMIC_E2E_BINARY still wins, and target/debug stays as the fallback. All three build hints name the profile, and the missing-binary hint now says when to keep ATOMICSERVER_SKIP_JS_BUILD and when to drop it, since building with it set against no dist is the other way to get a server that serves the wrong frontend. * Do not read a missing lsof as a free port The port check asked `lsof` and took a non-zero exit as "nothing is listening". A container without lsof exits non-zero for a different reason, `2>&1` swallows the "command not found", and the script walks on into the `--fresh` wipe. That is exactly the failure 28f643a was written to prevent: a stale server keeps serving a store that has just been deleted underneath it, and the failures that produces read as product bugs. So the tools are established separately from the answer, and the question is now the one that matters, whether anything is ANSWERING on that port, asked with curl. `--noproxy` because an HTTPS_PROXY in the environment would otherwise send a localhost probe through it. lsof still runs when it is there, since a process can hold the socket without answering, and either answer means busy. With neither tool, the script refuses only when `--fresh` would delete something. Without a wipe an unnoticed stale server can do no worse than make the new one fail to bind, which announces itself. This is not hypothetical, and not only about this script: the same shape of mistake nearly stopped a working rig today, when `ss` and `netstat` were both absent from a container, the pipeline had empty input, and the fallback said nothing was listening. Both servers were fine. * Script the README loop video instead of hand-cutting it (#1586) * fix(demo): keep the simulated collaborators visibly moving Yusuf's canvas cursor was a random walk that clamped its position at the edges without reflecting its velocity, so once the drift pointed outward the cursor sat pinned to a border for tens of seconds. It also drew the second creature's face from nowhere: the strokes appeared while his cursor was somewhere…
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Resumes
schema-in-code#1207/did:ad:frozen(#1208) on currentdevelop, then iterates toward a cleaner public API, more tests, and a faster freeze/resolve path. Related draft: #1209.Clean API
defineSchema/@tomic/lib/schemabuildSchemaLock/verifySchemaLock/isSchemaLockstore.useSchema(schema | lock, { publish? })registerFrozenSchema/loadSchemaLockuseSchemaregisterSchemacreateSchemaPointer/freezeStructureHappy path stays: define → use handles →
save()auto-publishes.Speed
defineSchema)Cache-Control: public, immutable, max-age=31536000onGET /frozenTests
freeze/schema/schema-lock/frozen-resolve/ vectors)frozen*lib tests green (cargo test -p atomic_lib --lib frozen --features db)Still open
registerSchemaonce imports are frozen-nativeFROZEN_REQUESTsync (Phase D)See
planning/did-frozen-schema-iteration.md.