Repository navigation
Repair expired cursors from artifacts on buckets that evict current values - #21
Merged
Merged
Conversation
…alues When a watcher's cursor expired, the resync deleted every in-scope key NATS no longer listed. On a bucket whose retention evicts current values (max_age, per-message TTLs, discard:old under a limit), "not in NATS" no longer means "deleted": the resync dropped valid keys that had merely aged out, and writes made during the gap that also aged out were never seen. Reproduced against a live nats-server for both discard:old and max_age. - KvWatcher::retention() reports, from the live stream config, whether retention evicts current values and the first retained revision. - watch_applied takes an ExpiryRepair (Relist / Restore / Auto / None); the old Option<reader> argument still compiles and means Relist. Relist (the key-listing diff) is refused on evicting buckets. - Restore replaces the in-scope fold with the newest published artifact, applied as a chunked diff through parse/apply/store, then resumes from the artifact's cursor. The artifact must be ahead of the fold, cover the watch scope, and still be inside retention (protocol::restore_allowed); otherwise the watch fails with the reason logged at error. ArtifactRestore implements the source over an ArtifactTransport. - Artifacts exported through watch_applied record their key scope (manifest schema 2; schema 1 still read). - The live floor guard now ends the watch with CursorExpired and is repaired in process through the same path, after draining the deliveries still buffered in the channel. - Repairs retry the pre-diff flush until the store holds every delivered update: a re-queued put was invisible to the diff and could resurrect a key deleted during the gap (also affected the old resync). - A cursor-less start on a bucket that already evicted current values seeds from the artifact. Models: tests/model.rs now judges convergence against every write minus every real delete, adds AgeOut and a gap write, and keeps the key-listing resync on an evicting bucket as a mutation the checker catches; tests/model_live_watch.rs gets the same axis plus an unguarded-restore mutation. Live tests in tests/eviction.rs. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0152kQKDdP8XRhoeqisJYpWr
… data The step-by-step repair model found a second pre-existing bug: a fold with data but no cursor (a torn first checkpoint, or a populated store started without one) was re-listed blind on start. A re-list never removes what it doesn't deliver, so a key deleted with its marker evicted stayed in the fold forever. Such a start is now repaired like an expiry at revision 0. The listing is also trusted on an evicting bucket that has never evicted anything (first retained revision <= 1), which that repair needs and which is exact. Verification added: - repair::tests: exhaustive checks of the real scope coverage (soundness over every small prefix set and key), check_restore against its spec, and restore_diff/materialize over every pair of 3-key folds x 4 scopes x every crash point mid-restore (re-run converges); manifest schema matrix. - tests/repair_dst.rs: deterministic fault injection over the real watch_applied. A simulated NATS log and a fault store (transient failure; crash before, after, or torn) at every store-apply call: ~19k schedules across 9 scenarios, both delivery orders, batch sizes 1 and 100. Checks convergence to every write minus every real delete, domain state == fold, and a monotone cursor. Reverting any of seven repair steps fails it. - tests/model_repair.rs: the repair as steps, interleaved with deliveries, writes, eviction, publishes, transient failures, and crashes; nine mutations, each caught. Makes axiom 6's relist window explicit (dropping it is reachable; the restore path doesn't need it). - tests/model_fleet.rs: exporters that themselves expire, restore, start cursor-less, and publish their actual folds. Every published artifact is the truth at its cursor, discharging the "exporters are correct" axiom. - tests/model.rs: the post-restore window re-check is load-bearing (caught when removed); the restore guard is fail-fast there (not caught). - Live NATS: per-message TTLs, discard:old under max_msgs, and max_age added to a live bucket all classify as evicting current values. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0152kQKDdP8XRhoeqisJYpWr
…kend The strengthened fault-injection harness found a third issue: a restore could make the consumer's domain state transiently wrong. A NATS cursor C means every RETAINED message <= C is applied, not "the truth at C": an artifact exported while its exporter was catching up lacks, or holds older values for, keys whose latest write is after C (the resume delivers them). Restoring from one — or restoring over a fold holding data ahead of its cursor after a torn write — deleted live keys and moved keys to older revisions until the resume caught up. The restore now takes an artifact entry only when it is newer than the local one, and deletes a key the artifact lacks only when the bucket does not list it live (ExpiryRepair::Restore now carries the reader that lists it). In both skipped cases the key's latest write is after C, so the resume delivers it: no phantom deletes or regressions, even transiently. Verification: - tests/repair_dst.rs is generic over the backend (append log, fjall, RocksDB) and covers the normal data path (resume with churn, fresh full re-list), multi-prefix scopes, and artifacts exported mid catch-up. It checks that apply never sees a revision go backward or a live key deleted, and that every mid-run export resumes to the truth. ~32k schedules on the append log; every single fault on fjall and RocksDB (deep tier for the full sweep). Reverting any of nine repair steps fails it. - tests/eviction.rs: a live floor-guard trip mid-watch on a real discard:old bucket (a stalled consumer overrun by a 20 MB burst) is repaired in process and ends identical to the exporter; reverting the trip's routing fails it. - tests/model_repair.rs states the cursor invariant precisely (while a cursor is resumable, value + retained tail = truth) and adds the domain no-regression / no-phantom-delete properties; 11 mutations caught. - repair::tests: the exhaustive diff check now spans every live listing. - Docs: with a store, resume from the store's cursor, not on_applied's (it runs ahead across a transient store failure). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0152kQKDdP8XRhoeqisJYpWr
dl.min.io now answers 410 Gone for every community server binary (only the commercial AIStor build is still served), so CI failed installing tools before running anything. Build the same pinned release from source with `go install` (~45 s, no replace directives in its go.mod). The four real-S3 tests (conditional-write pointer swap included) pass against it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0152kQKDdP8XRhoeqisJYpWr
jaredLunde
added a commit
that referenced
this pull request
Oct 1, 2026
…med (#22) slipstream is a bounded log: folds are the replicas of record. The store layer still assumed the older cache model, which every cursor-expiry bug in #21 traced back to. delete_with_version tombstones now reach watchers and folds as deletes (agreeing with get/scan/keys and the repair listings); Restore/Auto without a store is refused, and an incomplete cursor-less re-list warns; the store invariants, durability rationale, and runbooks are rewritten for replica-of-record (recover a fold by importing an artifact, never by re-listing NATS). Also pins, live, the supersession-only expiry fail-stop on evicting buckets (safe, unavailable until the next export round; export more often on small hot max_age buckets). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This was referenced Oct 2, 2026
Merged
jaredLunde
added a commit
that referenced
this pull request
Oct 2, 2026
Cursor-expiry repair for bounded logs (#21) and the replica-of-record cleanup (#22). Breaking: ExpiryRepair replaces the reader argument (the old Option<reader> still compiles), Relist is refused on buckets that evict current values, Restore/Auto require a store, artifact manifests move to schema 2 (upgrade importers before exporters), delete_with_version tombstones reach watchers as deletes, and the floor guard reports CursorExpired. Also moves the lockfile off yanked fjall/lsm-tree 3.1.4 to 3.1.10 (and chacha20, spin), so the suite runs on what dependents resolve, and refreshes the README install snippets to 0.8. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0152kQKDdP8XRhoeqisJYpWr
jaredLunde
added a commit
that referenced
this pull request
Oct 2, 2026
Cursor-expiry repair for bounded logs (#21) and the replica-of-record cleanup (#22). Breaking: ExpiryRepair replaces the reader argument (the old Option<reader> still compiles), Relist is refused on buckets that evict current values, Restore/Auto require a store, artifact manifests move to schema 2 (upgrade importers before exporters), delete_with_version tombstones reach watchers as deletes, and the floor guard reports CursorExpired. Also moves the lockfile off yanked fjall/lsm-tree 3.1.4 to 3.1.10 (and chacha20, spin), so the suite runs on what dependents resolve, and refreshes the README install snippets to 0.8. Claude-Session: https://claude.ai/code/session_0152kQKDdP8XRhoeqisJYpWr Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
When a watcher's cursor expired, the resync deleted every in-scope key NATS no longer listed. On a bucket whose retention evicts current values (
max_age, per-message TTLs,discard: oldunder a limit), "not in NATS" no longer means "deleted":Reproduced against a live nats-server for both
discard: oldandmax_age. A watcher that stayed live kept the aged-out key while a restarted one deleted it, so the fleet diverged. The formal model missed this because it defined "converged" as "matches NATS" and never evicted a current value.Fix
KvWatcher::retention()reports, from the live stream config, whether retention evicts current values and the first retained revision.ExpiryRepair(Relist/Restore/Auto/None) replaces thereaderargument ofwatch_applied.Option<Arc<dyn KvReader>>still converts into it, so existing call sites compile unchanged (Some(reader)meansRelist).Relist(the key-listing diff) is now refused on evicting buckets: the watch fails instead of deleting valid keys.Restorereplaces the in-scope fold with the newest published artifact and resumes from its cursor. The artifact must be ahead of the fold, cover the watch's scope, and have its cursor still inside retention (new shared kernelprotocol::restore_allowed); otherwise the watch fails with the reason logged aterror. It's applied as a chunked diff throughparse/apply/store, with the cursor moving only on the final commit, so a crash mid-restore re-runs it and the consumer's domain state sees exactly the changes.ArtifactRestoreimplements the source over anArtifactTransport.Autopicks by bucket retention at the moment of expiry.Floor guard: a mid-watch trip now ends the watch with
CursorExpiredand is repaired in process through the same path, after the deliveries still buffered in the channel are folded.Pre-existing bug fixed along the way: a transient store failure on the flush right before a repair left updates re-queued and invisible to the diff, so a re-queued put could resurrect a key deleted during the gap. This affected the old resync too; repairs now retry the flush until the store holds everything.
A cursor-less start on a bucket that has already evicted current values seeds from the artifact instead of an incomplete re-list.
Second pre-existing bug, found by the new step-by-step model: a fold with data but no cursor (a torn first checkpoint, or a populated store started without one) was re-listed blind on start, and a re-list never removes keys it doesn't deliver. Such a start is now repaired like an expiry at revision 0. The key listing is also trusted on an evicting bucket that has never evicted anything (first retained revision ≤ 1).
Third issue, found by the strengthened fault-injection harness: a restore could make the consumer's in-memory state transiently wrong. A NATS cursor C means every retained message ≤ C is applied, not "the truth at C". An artifact exported while its exporter was catching up therefore lacks, or holds older values for, keys whose latest write is after C. Restoring from one, or over a store holding data ahead of its cursor after a torn write, briefly deleted live keys or moved keys back to older revisions. The restore now takes an artifact entry only when it's newer, and deletes only keys the bucket doesn't list live.
ExpiryRepair::Restorenow carries thereaderthat does the listing.Rollout notes
export_todirectly) are refused for automatic restore. Trigger an export round right after rollout.Some(reader)on an evicting bucket go from silent data loss to a failed watch on expiry. Switch them toExpiryRepair::Autoin the same deploy.max_age(or thediscard: oldturnover), or an expired node has nothing fresh to restore from.ExpiryRepair::Noneto seed, then export from it. The error message says so.watch_all_fromcallers now getCursorExpired(notWatchError) when the floor guard trips.Known limitation
Shutdown isn't observed while a restore is in progress. This is safe, since a kill mid-restore re-runs it, but a long restore on a large on-disk fold can outlast a deploy's grace period. Follow-up: check for shutdown between chunks and make the diff scan cancellable.
Verification
Third commit (
e6a5a74):tests/repair_dst.rsnow runs on every backend (in-RAM append log, fjall, RocksDB). It covers the normal data path (resume with churn, fresh full re-list), multi-prefix scopes, and artifacts exported mid catch-up. It also checks thatapplynever sees a key's revision go backward or a live key deleted, and that every mid-run export resumes to the truth. That's ~32k schedules on the append log and every single fault on fjall and RocksDB; the full on-disk sweep is an ignored deep tier, ~11 min. Reverting any of nine repair steps fails it.tests/eviction.rs: on a realdiscard: oldbucket, a stalled, floor-guarded watcher is overrun by a 20 MB burst. The server stops pushing to a consumer that doesn't answer flow control, so retention evicts what it never sent. Released, the watcher trips mid-stream, restores in process, and ends identical to the exporter. Passed 5/5 runs, and reverting the trip's routing fails it.tests/model_repair.rsstates the cursor invariant precisely: while a cursor is resumable, value + retained tail = truth. It adds the no-regression and no-phantom-delete properties; 11 mutations are caught.on_applied's. Across a transient store failure,on_appliedruns ahead of what the store made durable.Second commit (
34c5cb4), on top of the tests below:repair::tests(exhaustive, real code)check_restoreagainst its spec;restore_diffover every pair of 3-key folds × 4 scopes × every crash point mid-restoretests/repair_dst.rs(fault injection, realwatch_applied)tests/model_repair.rs(Stateright)tests/model_fleet.rs(Stateright)tests/model.rsdiscard: oldwithmax_msgs, andmax_ageadded to a live bucket all count as evicting current valuesDeep tiers (ignored by default, all pass in release): evicting
model.rsrev ≤ 4 (124M unique states), fleet rev ≤ 6, repair steps rev ≤ 5 with 2 crashes and 3 transient failures.What's still an assumption: the key-listing repair (buckets that keep current values) assumes no admin purge of a delete marker during the milliseconds between taking its listing and starting its re-list.
model_repairdrops that assumption and the checker reaches the trace; the restore path doesn't depend on it. All model results are exhaustive within the stated bounds, not unbounded proofs.First commit:
tests/eviction.rs(new, live NATS + local object store, both eviction kinds): restore keeps the aged-out key, recovers the gap write, applies the real delete, and ends identical to the exporter; relist is refused with the fold untouched; a stale artifact fails the watch.tests/model.rs: convergence is now judged against every write minus every real delete, with anAgeOutstep and a gap write. The shipped repair holds on both bucket kinds, and the key-listing resync on an evicting bucket is kept as a mutation the checker catches (with a lost-write counterexample). A deep evicting run (rev ≤ 4, ~124M unique states) passes; it's marked ignored by default.tests/model_live_watch.rs: same axis for the floor-guard repair, plus mutations for relist-on-evicting and a restore withoutrestore_allowed. Both are caught.--features fjall,rocksdb,transportand none) pass, and clippy is clean with-D warnings. The main model takes ~85s in debug.🤖 Generated with Claude Code
https://claude.ai/code/session_0152kQKDdP8XRhoeqisJYpWr