Current nemesis.clj supports :partition, :kill (SIGKILL), :pause — all process/network-level faults. None simulate real power loss: a killed process still leaves the OS page cache intact, so any WAL data written-but-not-yet-fsync'd survives via normal replay on restart. This is fundamentally different from real power loss, where un-fsynced, page-cache-resident writes can be dropped.
GUARANTEES.md already documents this gap explicitly: "Simultaneous crash of all nodes (e.g. full power loss) is not covered; d-engine uses a MemFirst persistence strategy with periodic batch flushing, so unflushed entries could be lost in that scenario." This statement is currently asserted but never verified against a real fault injection.
Jepsen's own author has built LazyFS specifically for this: a filesystem layer that lets Jepsen drop any writes not yet confirmed via fsync(), used in prior Jepsen analyses (e.g. NATS 2.12.1) to simulate "crash with partial amnesia." No LazyFS integration currently exists in this repo.
Expectation: integrate LazyFS (or an equivalent fsync-aware write-dropping layer) into the nemesis fault set, so we can (a) confirm whether an acknowledged-but-not-yet-durable write is actually lost when a majority of nodes experience this simultaneously, and (b) quantify how wide that risk window actually is in practice — closing the gap between what GUARANTEES.md currently asserts and what's actually been tested.
Refers: deventlab/d-engine#442
Current
nemesis.cljsupports:partition,:kill(SIGKILL),:pause— all process/network-level faults. None simulate real power loss: a killed process still leaves the OS page cache intact, so any WAL data written-but-not-yet-fsync'd survives via normal replay on restart. This is fundamentally different from real power loss, where un-fsynced, page-cache-resident writes can be dropped.GUARANTEES.mdalready documents this gap explicitly: "Simultaneous crash of all nodes (e.g. full power loss) is not covered; d-engine uses aMemFirstpersistence strategy with periodic batch flushing, so unflushed entries could be lost in that scenario." This statement is currently asserted but never verified against a real fault injection.Jepsen's own author has built LazyFS specifically for this: a filesystem layer that lets Jepsen drop any writes not yet confirmed via
fsync(), used in prior Jepsen analyses (e.g. NATS 2.12.1) to simulate "crash with partial amnesia." No LazyFS integration currently exists in this repo.Expectation: integrate LazyFS (or an equivalent fsync-aware write-dropping layer) into the nemesis fault set, so we can (a) confirm whether an acknowledged-but-not-yet-durable write is actually lost when a majority of nodes experience this simultaneously, and (b) quantify how wide that risk window actually is in practice — closing the gap between what
GUARANTEES.mdcurrently asserts and what's actually been tested.Refers: deventlab/d-engine#442