Skip to content

Snapshot replication failure handling lacks clarity #449

Description

@JoshuaChi

Problem

When a peer's snapshot transfer encounters backpressure (bounded send channel full) or timeout, the current code path does not have an explicit strategy for:

  1. Retry discipline — how many attempts before giving up? Exponential backoff or fixed delay?
  2. State coherence — if snapshot transfer is interrupted mid-stream, what prevents the follower's state machine from diverging from its log?
  3. Downgrade semantics — should repeated snapshot failure trigger any state transition (e.g., back to Probe)? If yes, does it risk log-state-machine inconsistency?

Impact

  • No observability into why a peer gets stuck in Snapshot state
  • Risk of silent state machine corruption if failure handling is ad-hoc
  • Interacts with the try_send() improvement for AppendEntries (if AppendEntries fails and transitions to Probe, but Snapshot is still in flight on the same peer, behaviors are unclear)

Acceptance

Ticket is resolved when:

  • Snapshot failure modes (backpressure, timeout, stream broken) have explicit, documented handling paths
  • State machine coherence guarantees are spelled out
  • Test coverage exists for each failure mode

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    component:raft-snapshotSnapshot creation, transfer, installation, and compaction triggers.

    Type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions