Problem
When a peer's snapshot transfer encounters backpressure (bounded send channel full) or timeout, the current code path does not have an explicit strategy for:
- Retry discipline — how many attempts before giving up? Exponential backoff or fixed delay?
- State coherence — if snapshot transfer is interrupted mid-stream, what prevents the follower's state machine from diverging from its log?
- Downgrade semantics — should repeated snapshot failure trigger any state transition (e.g., back to Probe)? If yes, does it risk log-state-machine inconsistency?
Impact
- No observability into why a peer gets stuck in
Snapshot state
- Risk of silent state machine corruption if failure handling is ad-hoc
- Interacts with the
try_send() improvement for AppendEntries (if AppendEntries fails and transitions to Probe, but Snapshot is still in flight on the same peer, behaviors are unclear)
Acceptance
Ticket is resolved when:
- Snapshot failure modes (backpressure, timeout, stream broken) have explicit, documented handling paths
- State machine coherence guarantees are spelled out
- Test coverage exists for each failure mode
Problem
When a peer's snapshot transfer encounters backpressure (bounded send channel full) or timeout, the current code path does not have an explicit strategy for:
Impact
Snapshotstatetry_send()improvement for AppendEntries (if AppendEntries fails and transitions to Probe, but Snapshot is still in flight on the same peer, behaviors are unclear)Acceptance
Ticket is resolved when: