Add full-instance backup capture and offline restore - #1432
michielbdejong wants to merge 2 commits into
Conversation
|
Some doubts about this implementation:
|
|
Good questions. The first one is now fixed and the second one is not. One nice way to do backups would be to actually reuse the "per-doc Loro inside a Merkle tree of resources and blobs" representation that is already used for server sync, collaboration, undo and for vault backup. That could mean full per-op rollback, which could be really powerful if atomic server becomes a more versatile AI-powered System of Record. You are right to flag that the 'database file zip' this PR uses is in some sense parallel to what already exists in the code base. It would add a very custom format in which to represent the state of a server. This would not be a good format for archaeology - you would need a compatible version of atomic-server to rehydrate it. You're right, another option is reusing the filesystem view, just zip that up instead of zipping up the database folder. That would be much easier to restore. But: We would first need to continue the implementation of the filesystem view. Right now files work but resources are just links into atomic-desktop, so creating a backup from a zip of the filesystem view would not actually store your data. It's not unreasonable to have a distinction between incremental backups and snapshot backups for different operational goals. So yeah I guess to do this properly would still be quite some work. I'll leave this PR open and think about it some more. |
a933ec2 to
c9604a2
Compare
|
I see the target this branch is starting to drift, let's leave it for now and address backups more properly if/when at some point we rethink the storage-and-sync architecture |
A running node can now create a complete local instance checkpoint without exiting. Capture drains admitted operations, gates writes and incoming sync application, copies every redb table from one consistent read transaction plus associated data/config files, then resumes normal operation before ZIP compression. Restore verifies the archive and database into a new directory and leaves it offline until explicitly activated.
Capture, manifest/checksum verification, ZIP creation and offline restore live in
atomic_lib::backup, behind an optional nativebackupfeature.CheckpointOptionsprovides explicit data/config/output roots and build metadata. The server is an adapter for authenticated loopback control, CLI/status and server-specific cache checks; native callers use the same API without atomic-server or Actix. This follows the runtime boundary in #1416 without depending on that PR landing first. The v1 archive layout and CLI commands are preserved.This checkpoint uses the existing
Dbmaintenance and sync application paths. It is distinct from the desktop NFS graph projection and encrypted drive vault: those edit/merge drive data, while this restores all persisted instance state and configuration. The docs compare these mechanisms and record that desktop integration must stop/flush VFS staging before capture and enforce the offline marker before reconnecting copied identities. No desktop backup UI or staging integration is claimed.Validation
Current refactor, macOS:
sh lib/tests/check-instance-checkpoint.sh: passes the core-only server/Actix dependency gate, 2 standalone tests and 8 core backup tests. Covers originless capture/verify/restore/reopen with identity, phase/pause ordering, overlapping roots, Loro history, blobs/envelopes, storage barriers, failure recovery, archive rejection and paused sync imports.cargo test --locked -p atomic-server --no-default-features --lib: 69 passed.cargo test --locked -p atomic-server --no-default-features --test it instance_backup: real server process, WS replication, backup CLI, later replicated edit, offline restore and startup refusal passed.git diff --checkpass.ATOMICSERVER_SKIP_JS_BUILD=truewith a reused HTML asset. No frontend changes or desktop interaction tests. Existing compiler warnings remain. Dagger now runs the standalone checkpoint gate after workspace tests; the full Dagger pipeline was not run locally.Operational limits
Local persisted checkpoint; no proof of remote replication freshness or eternal per-edit history. V1 targets native redb on macOS/Linux, requires separate data/config/cache roots, rejects symlinks/special files, and does not install a nightly job or prune archives. Large-dataset pause/ZIP64 benchmarks, physical disk-full/power-loss injection, dedicated Iroh disconnect-during-capture testing and desktop staging/restore acceptance remain documented gaps. No production data was used.