Skip to content

Leader produces replication requests regardless of peer reachability #450

Description

@JoshuaChi

Problem

  • Whether a peer gets a new request is decided per proposed batch, not by the
    peer's reachability.
  • While a peer is unreachable, one request per batch (mostly empty) is queued in an
    unbounded per-peer queue and never dropped; the Raft layer is not told.
  • Empty requests share that queue with data requests, with no rate limit or
    deduplication. For keepalive only the most recent one matters.

Observed (3-node local bench, load starts right after leader election)

  • ~1.5s unreachable window -> ~17k queued requests, almost all empty.
  • On recovery the backlog is flushed at once; the follower stops reading (pending
    limit), the leader's send buffer fills, the stream is torn down and reopened
    repeatedly, in-flight responses are lost.
  • The peer ends at ~2% applied within the observation window.

Open questions

  • Empty requests also carry the commit index and serve as the lease send-time
    anchor and the read quorum ack. Which duties need per-batch cadence?

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    raft-clusterCluster-wide operations and coordination issues

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions