Skip to content

Repair expired cursors from artifacts on buckets that evict current values - #21

Merged
jaredLunde merged 4 commits into
mainfrom
jared/correctness
Oct 1, 2026
Merged

jaredLunde merged 4 commits into
mainfrom
jared/correctness

Conversation

@jaredLunde

@jaredLunde jaredLunde commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Problem

When a watcher's cursor expired, the resync deleted every in-scope key NATS no longer listed. On a bucket whose retention evicts current values (max_age, per-message TTLs, discard: old under a limit), "not in NATS" no longer means "deleted":

  • valid keys that merely aged out were deleted from the fold, and
  • writes made while the watcher was offline that also aged out were never seen.

Reproduced against a live nats-server for both discard: old and max_age. A watcher that stayed live kept the aged-out key while a restarted one deleted it, so the fleet diverged. The formal model missed this because it defined "converged" as "matches NATS" and never evicted a current value.

Fix

  • KvWatcher::retention() reports, from the live stream config, whether retention evicts current values and the first retained revision.

  • ExpiryRepair (Relist / Restore / Auto / None) replaces the reader argument of watch_applied. Option<Arc<dyn KvReader>> still converts into it, so existing call sites compile unchanged (Some(reader) means Relist).

  • Relist (the key-listing diff) is now refused on evicting buckets: the watch fails instead of deleting valid keys.

  • Restore replaces the in-scope fold with the newest published artifact and resumes from its cursor. The artifact must be ahead of the fold, cover the watch's scope, and have its cursor still inside retention (new shared kernel protocol::restore_allowed); otherwise the watch fails with the reason logged at error. It's applied as a chunked diff through parse/apply/store, with the cursor moving only on the final commit, so a crash mid-restore re-runs it and the consumer's domain state sees exactly the changes. ArtifactRestore implements the source over an ArtifactTransport.

  • Auto picks by bucket retention at the moment of expiry.

  • Floor guard: a mid-watch trip now ends the watch with CursorExpired and is repaired in process through the same path, after the deliveries still buffered in the channel are folded.

  • Pre-existing bug fixed along the way: a transient store failure on the flush right before a repair left updates re-queued and invisible to the diff, so a re-queued put could resurrect a key deleted during the gap. This affected the old resync too; repairs now retry the flush until the store holds everything.

  • A cursor-less start on a bucket that has already evicted current values seeds from the artifact instead of an incomplete re-list.

  • Second pre-existing bug, found by the new step-by-step model: a fold with data but no cursor (a torn first checkpoint, or a populated store started without one) was re-listed blind on start, and a re-list never removes keys it doesn't deliver. Such a start is now repaired like an expiry at revision 0. The key listing is also trusted on an evicting bucket that has never evicted anything (first retained revision ≤ 1).

  • Third issue, found by the strengthened fault-injection harness: a restore could make the consumer's in-memory state transiently wrong. A NATS cursor C means every retained message ≤ C is applied, not "the truth at C". An artifact exported while its exporter was catching up therefore lacks, or holds older values for, keys whose latest write is after C. Restoring from one, or over a store holding data ahead of its cursor after a torn write, briefly deleted live keys or moved keys back to older revisions. The restore now takes an artifact entry only when it's newer, and deletes only keys the bucket doesn't list live. ExpiryRepair::Restore now carries the reader that does the listing.

Rollout notes

  • Artifact manifest schema 2 adds the exporting watch's key scope. Older builds refuse schema-2 manifests, so upgrade importing nodes before exporting nodes.
  • Artifacts with no recorded scope (schema 1, or exported via export_to directly) are refused for automatic restore. Trigger an export round right after rollout.
  • Callers passing Some(reader) on an evicting bucket go from silent data loss to a failed watch on expiry. Switch them to ExpiryRepair::Auto in the same deploy.
  • Run exports well inside max_age (or the discard: old turnover), or an expired node has nothing fresh to restore from.
  • First deploy onto a bucket that has already evicted values, with no artifact yet: start one node with ExpiryRepair::None to seed, then export from it. The error message says so.
  • Raw watch_all_from callers now get CursorExpired (not WatchError) when the floor guard trips.

Known limitation

Shutdown isn't observed while a restore is in progress. This is safe, since a kill mid-restore re-runs it, but a long restore on a large on-disk fold can outlast a deploy's grace period. Follow-up: check for shutdown between chunks and make the diff scan cancellable.

Verification

Third commit (e6a5a74):

  • tests/repair_dst.rs now runs on every backend (in-RAM append log, fjall, RocksDB). It covers the normal data path (resume with churn, fresh full re-list), multi-prefix scopes, and artifacts exported mid catch-up. It also checks that apply never sees a key's revision go backward or a live key deleted, and that every mid-run export resumes to the truth. That's ~32k schedules on the append log and every single fault on fjall and RocksDB; the full on-disk sweep is an ignored deep tier, ~11 min. Reverting any of nine repair steps fails it.
  • Live floor-guard trip → restore, in tests/eviction.rs: on a real discard: old bucket, a stalled, floor-guarded watcher is overrun by a 20 MB burst. The server stops pushing to a consumer that doesn't answer flow control, so retention evicts what it never sent. Released, the watcher trips mid-stream, restores in process, and ends identical to the exporter. Passed 5/5 runs, and reverting the trip's routing fails it.
  • tests/model_repair.rs states the cursor invariant precisely: while a cursor is resumable, value + retained tail = truth. It adds the no-regression and no-phantom-delete properties; 11 mutations are caught.
  • Docs: with a store, resume from the store's cursor, not on_applied's. Across a transient store failure, on_applied runs ahead of what the store made durable.

Second commit (34c5cb4), on top of the tests below:

Layer What it checks Teeth
repair::tests (exhaustive, real code) Scope-coverage soundness over every small prefix set and key; check_restore against its spec; restore_diff over every pair of 3-key folds × 4 scopes × every crash point mid-restore Each check compares against an independent spec
tests/repair_dst.rs (fault injection, real watch_applied) ~19k schedules: a transient failure or a crash (before / after / torn) at every store write, 9 scenarios, both delivery orders, batch sizes 1 and 100. After each schedule: the fold equals every write minus every real delete, in-memory state equals the fold, and the cursor never moves backward Reverting any of 7 repair steps fails it
tests/model_repair.rs (Stateright) The repair as separate steps, interleaved with deliveries, writes, eviction, publishes, transient failures and crashes 9 mutations, each caught
tests/model_fleet.rs (Stateright) Exporters that themselves expire, restore, start with no cursor, and publish their actual folds: every artifact is the truth at its cursor. This discharges the "exporters are correct" assumption 2 fleet-poisoning mutations caught
tests/model.rs The post-restore window re-check carries safety (removing it is caught); the freshness check is an early refusal (removing it is not caught, as expected) —
Live NATS Per-message TTL, discard: old with max_msgs, and max_age added to a live bucket all count as evicting current values —

Deep tiers (ignored by default, all pass in release): evicting model.rs rev ≤ 4 (124M unique states), fleet rev ≤ 6, repair steps rev ≤ 5 with 2 crashes and 3 transient failures.

What's still an assumption: the key-listing repair (buckets that keep current values) assumes no admin purge of a delete marker during the milliseconds between taking its listing and starting its re-list. model_repair drops that assumption and the checker reaches the trace; the restore path doesn't depend on it. All model results are exhaustive within the stated bounds, not unbounded proofs.

First commit:

  • tests/eviction.rs (new, live NATS + local object store, both eviction kinds): restore keeps the aged-out key, recovers the gap write, applies the real delete, and ends identical to the exporter; relist is refused with the fold untouched; a stale artifact fails the watch.
  • Unit tests for every restore refusal, scope handling, the floor-guard drain, cursor monotonicity, fresh-start seeding, scope stamping on export, and the re-queued-flush bug. The drain and relist-refusal tests were confirmed to fail with those fixes removed.
  • tests/model.rs: convergence is now judged against every write minus every real delete, with an AgeOut step and a gap write. The shipped repair holds on both bucket kinds, and the key-listing resync on an evicting bucket is kept as a mutation the checker catches (with a lost-write counterexample). A deep evicting run (rev ≤ 4, ~124M unique states) passes; it's marked ignored by default.
  • tests/model_live_watch.rs: same axis for the floor-guard repair, plus mutations for relist-on-evicting and a restore without restore_allowed. Both are caught.
  • Both CI configurations (--features fjall,rocksdb,transport and none) pass, and clippy is clean with -D warnings. The main model takes ~85s in debug.

🤖 Generated with Claude Code

https://claude.ai/code/session_0152kQKDdP8XRhoeqisJYpWr

jaredLunde and others added 4 commits October 1, 2026 13:19
…alues

When a watcher's cursor expired, the resync deleted every in-scope key NATS
no longer listed. On a bucket whose retention evicts current values
(max_age, per-message TTLs, discard:old under a limit), "not in NATS" no
longer means "deleted": the resync dropped valid keys that had merely aged
out, and writes made during the gap that also aged out were never seen.
Reproduced against a live nats-server for both discard:old and max_age.

- KvWatcher::retention() reports, from the live stream config, whether
  retention evicts current values and the first retained revision.
- watch_applied takes an ExpiryRepair (Relist / Restore / Auto / None);
  the old Option<reader> argument still compiles and means Relist.
  Relist (the key-listing diff) is refused on evicting buckets.
- Restore replaces the in-scope fold with the newest published artifact,
  applied as a chunked diff through parse/apply/store, then resumes from
  the artifact's cursor. The artifact must be ahead of the fold, cover the
  watch scope, and still be inside retention (protocol::restore_allowed);
  otherwise the watch fails with the reason logged at error.
  ArtifactRestore implements the source over an ArtifactTransport.
- Artifacts exported through watch_applied record their key scope
  (manifest schema 2; schema 1 still read).
- The live floor guard now ends the watch with CursorExpired and is
  repaired in process through the same path, after draining the deliveries
  still buffered in the channel.
- Repairs retry the pre-diff flush until the store holds every delivered
  update: a re-queued put was invisible to the diff and could resurrect a
  key deleted during the gap (also affected the old resync).
- A cursor-less start on a bucket that already evicted current values
  seeds from the artifact.

Models: tests/model.rs now judges convergence against every write minus
every real delete, adds AgeOut and a gap write, and keeps the key-listing
resync on an evicting bucket as a mutation the checker catches;
tests/model_live_watch.rs gets the same axis plus an unguarded-restore
mutation. Live tests in tests/eviction.rs.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0152kQKDdP8XRhoeqisJYpWr
… data

The step-by-step repair model found a second pre-existing bug: a fold with
data but no cursor (a torn first checkpoint, or a populated store started
without one) was re-listed blind on start. A re-list never removes what it
doesn't deliver, so a key deleted with its marker evicted stayed in the fold
forever. Such a start is now repaired like an expiry at revision 0. The
listing is also trusted on an evicting bucket that has never evicted
anything (first retained revision <= 1), which that repair needs and which
is exact.

Verification added:
- repair::tests: exhaustive checks of the real scope coverage (soundness
  over every small prefix set and key), check_restore against its spec, and
  restore_diff/materialize over every pair of 3-key folds x 4 scopes x
  every crash point mid-restore (re-run converges); manifest schema matrix.
- tests/repair_dst.rs: deterministic fault injection over the real
  watch_applied. A simulated NATS log and a fault store (transient failure;
  crash before, after, or torn) at every store-apply call: ~19k schedules
  across 9 scenarios, both delivery orders, batch sizes 1 and 100. Checks
  convergence to every write minus every real delete, domain state == fold,
  and a monotone cursor. Reverting any of seven repair steps fails it.
- tests/model_repair.rs: the repair as steps, interleaved with deliveries,
  writes, eviction, publishes, transient failures, and crashes; nine
  mutations, each caught. Makes axiom 6's relist window explicit (dropping
  it is reachable; the restore path doesn't need it).
- tests/model_fleet.rs: exporters that themselves expire, restore, start
  cursor-less, and publish their actual folds. Every published artifact is
  the truth at its cursor, discharging the "exporters are correct" axiom.
- tests/model.rs: the post-restore window re-check is load-bearing (caught
  when removed); the restore guard is fail-fast there (not caught).
- Live NATS: per-message TTLs, discard:old under max_msgs, and max_age
  added to a live bucket all classify as evicting current values.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0152kQKDdP8XRhoeqisJYpWr
…kend

The strengthened fault-injection harness found a third issue: a restore
could make the consumer's domain state transiently wrong. A NATS cursor C
means every RETAINED message <= C is applied, not "the truth at C": an
artifact exported while its exporter was catching up lacks, or holds older
values for, keys whose latest write is after C (the resume delivers them).
Restoring from one — or restoring over a fold holding data ahead of its
cursor after a torn write — deleted live keys and moved keys to older
revisions until the resume caught up.

The restore now takes an artifact entry only when it is newer than the
local one, and deletes a key the artifact lacks only when the bucket does
not list it live (ExpiryRepair::Restore now carries the reader that lists
it). In both skipped cases the key's latest write is after C, so the
resume delivers it: no phantom deletes or regressions, even transiently.

Verification:
- tests/repair_dst.rs is generic over the backend (append log, fjall,
  RocksDB) and covers the normal data path (resume with churn, fresh full
  re-list), multi-prefix scopes, and artifacts exported mid catch-up. It
  checks that apply never sees a revision go backward or a live key
  deleted, and that every mid-run export resumes to the truth. ~32k
  schedules on the append log; every single fault on fjall and RocksDB
  (deep tier for the full sweep). Reverting any of nine repair steps
  fails it.
- tests/eviction.rs: a live floor-guard trip mid-watch on a real
  discard:old bucket (a stalled consumer overrun by a 20 MB burst) is
  repaired in process and ends identical to the exporter; reverting the
  trip's routing fails it.
- tests/model_repair.rs states the cursor invariant precisely (while a
  cursor is resumable, value + retained tail = truth) and adds the
  domain no-regression / no-phantom-delete properties; 11 mutations caught.
- repair::tests: the exhaustive diff check now spans every live listing.
- Docs: with a store, resume from the store's cursor, not on_applied's
  (it runs ahead across a transient store failure).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0152kQKDdP8XRhoeqisJYpWr
dl.min.io now answers 410 Gone for every community server binary (only
the commercial AIStor build is still served), so CI failed installing
tools before running anything. Build the same pinned release from source
with `go install` (~45 s, no replace directives in its go.mod). The four
real-S3 tests (conditional-write pointer swap included) pass against it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0152kQKDdP8XRhoeqisJYpWr
@jaredLunde
jaredLunde merged commit 79ad13d into main Oct 1, 2026
1 check passed
@jaredLunde
jaredLunde deleted the jared/correctness branch October 1, 2026 23:14
jaredLunde added a commit that referenced this pull request Oct 1, 2026
…med (#22)

slipstream is a bounded log: folds are the replicas of record. The store layer still assumed the older cache model, which every cursor-expiry bug in #21 traced back to. delete_with_version tombstones now reach watchers and folds as deletes (agreeing with get/scan/keys and the repair listings); Restore/Auto without a store is refused, and an incomplete cursor-less re-list warns; the store invariants, durability rationale, and runbooks are rewritten for replica-of-record (recover a fold by importing an artifact, never by re-listing NATS). Also pins, live, the supersession-only expiry fail-stop on evicting buckets (safe, unavailable until the next export round; export more often on small hot max_age buckets).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
jaredLunde added a commit that referenced this pull request Oct 2, 2026
Cursor-expiry repair for bounded logs (#21) and the replica-of-record
cleanup (#22). Breaking: ExpiryRepair replaces the reader argument (the old
Option<reader> still compiles), Relist is refused on buckets that evict
current values, Restore/Auto require a store, artifact manifests move to
schema 2 (upgrade importers before exporters), delete_with_version
tombstones reach watchers as deletes, and the floor guard reports
CursorExpired.

Also moves the lockfile off yanked fjall/lsm-tree 3.1.4 to 3.1.10 (and
chacha20, spin), so the suite runs on what dependents resolve, and
refreshes the README install snippets to 0.8.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0152kQKDdP8XRhoeqisJYpWr
jaredLunde added a commit that referenced this pull request Oct 2, 2026
Cursor-expiry repair for bounded logs (#21) and the replica-of-record
cleanup (#22). Breaking: ExpiryRepair replaces the reader argument (the old
Option<reader> still compiles), Relist is refused on buckets that evict
current values, Restore/Auto require a store, artifact manifests move to
schema 2 (upgrade importers before exporters), delete_with_version
tombstones reach watchers as deletes, and the floor guard reports
CursorExpired.

Also moves the lockfile off yanked fjall/lsm-tree 3.1.4 to 3.1.10 (and
chacha20, spin), so the suite runs on what dependents resolve, and
refreshes the README install snippets to 0.8.


Claude-Session: https://claude.ai/code/session_0152kQKDdP8XRhoeqisJYpWr

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant