Skip to content

Add full-instance backup capture and offline restore - #1432

Closed
michielbdejong wants to merge 2 commits into
developfrom
codex/instance-backup
Closed

michielbdejong wants to merge 2 commits into
developfrom
codex/instance-backup

Conversation

@michielbdejong

@michielbdejong michielbdejong commented Sep 11, 2026 •

Copy link
Copy Markdown
Contributor

A running node can now create a complete local instance checkpoint without exiting. Capture drains admitted operations, gates writes and incoming sync application, copies every redb table from one consistent read transaction plus associated data/config files, then resumes normal operation before ZIP compression. Restore verifies the archive and database into a new directory and leaves it offline until explicitly activated.

Capture, manifest/checksum verification, ZIP creation and offline restore live in atomic_lib::backup, behind an optional native backup feature. CheckpointOptions provides explicit data/config/output roots and build metadata. The server is an adapter for authenticated loopback control, CLI/status and server-specific cache checks; native callers use the same API without atomic-server or Actix. This follows the runtime boundary in #1416 without depending on that PR landing first. The v1 archive layout and CLI commands are preserved.

This checkpoint uses the existing Db maintenance and sync application paths. It is distinct from the desktop NFS graph projection and encrypted drive vault: those edit/merge drive data, while this restores all persisted instance state and configuration. The docs compare these mechanisms and record that desktop integration must stop/flush VFS staging before capture and enforce the offline marker before reconnecting copied identities. No desktop backup UI or staging integration is claimed.

Validation

Current refactor, macOS:

  • sh lib/tests/check-instance-checkpoint.sh: passes the core-only server/Actix dependency gate, 2 standalone tests and 8 core backup tests. Covers originless capture/verify/restore/reopen with identity, phase/pause ordering, overlapping roots, Loro history, blobs/envelopes, storage barriers, failure recovery, archive rejection and paused sync imports.
  • cargo test --locked -p atomic-server --no-default-features --lib: 69 passed.
  • cargo test --locked -p atomic-server --no-default-features --test it instance_backup: real server process, WS replication, backup CLI, later replicated edit, offline restore and startup refusal passed.
  • Changed Rust files pass rustfmt; shell syntax and git diff --check pass.
  • Server checks used ATOMICSERVER_SKIP_JS_BUILD=true with a reused HTML asset. No frontend changes or desktop interaction tests. Existing compiler warnings remain. Dagger now runs the standalone checkpoint gate after workspace tests; the full Dagger pipeline was not run locally.

Operational limits

Local persisted checkpoint; no proof of remote replication freshness or eternal per-edit history. V1 targets native redb on macOS/Linux, requires separate data/config/cache roots, rejects symlinks/special files, and does not install a nightly job or prune archives. Large-dataset pause/ZIP64 benchmarks, physical disk-full/power-loss injection, dedicated Iroh disconnect-during-capture testing and desktop staging/restore acceptance remain documented gaps. No production data was used.

@joepio

joepio commented Sep 11, 2026

Copy link
Copy Markdown
Member

Some doubts about this implementation:

@michielbdejong

michielbdejong commented Sep 14, 2026 •

Copy link
Copy Markdown
Contributor Author

Good questions. The first one is now fixed and the second one is not.

One nice way to do backups would be to actually reuse the "per-doc Loro inside a Merkle tree of resources and blobs" representation that is already used for server sync, collaboration, undo and for vault backup. That could mean full per-op rollback, which could be really powerful if atomic server becomes a more versatile AI-powered System of Record.

You are right to flag that the 'database file zip' this PR uses is in some sense parallel to what already exists in the code base. It would add a very custom format in which to represent the state of a server. This would not be a good format for archaeology - you would need a compatible version of atomic-server to rehydrate it.

You're right, another option is reusing the filesystem view, just zip that up instead of zipping up the database folder. That would be much easier to restore.

But: We would first need to continue the implementation of the filesystem view. Right now files work but resources are just links into atomic-desktop, so creating a backup from a zip of the filesystem view would not actually store your data.

It's not unreasonable to have a distinction between incremental backups and snapshot backups for different operational goals. So yeah I guess to do this properly would still be quite some work. I'll leave this PR open and think about it some more.

@joepio
joepio force-pushed the codex/instance-backup branch from a933ec2 to c9604a2 Compare September 17, 2026 22:11
@michielbdejong

Copy link
Copy Markdown
Contributor Author

I see the target this branch is starting to drift, let's leave it for now and address backups more properly if/when at some point we rethink the storage-and-sync architecture

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants