Skip to content

feat: Add LazyFS-based power-loss fault injection to Jepsen nemesis (current faults are process-level only) #25

Description

@JoshuaChi

Current nemesis.clj supports :partition, :kill (SIGKILL), :pause — all process/network-level faults. None simulate real power loss: a killed process still leaves the OS page cache intact, so any WAL data written-but-not-yet-fsync'd survives via normal replay on restart. This is fundamentally different from real power loss, where un-fsynced, page-cache-resident writes can be dropped.

GUARANTEES.md already documents this gap explicitly: "Simultaneous crash of all nodes (e.g. full power loss) is not covered; d-engine uses a MemFirst persistence strategy with periodic batch flushing, so unflushed entries could be lost in that scenario." This statement is currently asserted but never verified against a real fault injection.

Jepsen's own author has built LazyFS specifically for this: a filesystem layer that lets Jepsen drop any writes not yet confirmed via fsync(), used in prior Jepsen analyses (e.g. NATS 2.12.1) to simulate "crash with partial amnesia." No LazyFS integration currently exists in this repo.

Expectation: integrate LazyFS (or an equivalent fsync-aware write-dropping layer) into the nemesis fault set, so we can (a) confirm whether an acknowledged-but-not-yet-durable write is actually lost when a majority of nodes experience this simultaneously, and (b) quantify how wide that risk window actually is in practice — closing the gap between what GUARANTEES.md currently asserts and what's actually been tested.

Refers: deventlab/d-engine#442

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions