Skip to content

fix(read) #381: LinearizableRead must not be served without quorum in multi-voter cluster - #383

Merged
JoshuaChi merged 2 commits into
mainfrom
fix/381-linear-bug
May 18, 2026
Merged

JoshuaChi merged 2 commits into
mainfrom
fix/381-linear-bug

Conversation

@JoshuaChi

@JoshuaChi JoshuaChi commented May 18, 2026 •

Copy link
Copy Markdown
Contributor

Root cause

In Phase 3 of execute_and_process_raft_rpc, a multi-voter leader served
LinearizableRead immediately when last_applied >= read_index, with zero
quorum confirmation. A leader isolated in a minority partition would answer
stale reads — Raft §8 violation.

Changes

leader_state.rs — Phase 3
Add single_voter && guard: single-voter self IS the quorum (safe to serve
immediately); multi-voter reads always queue into pending_reads so they
cannot be served without a network round-trip.

leader_state.rs — handle_append_result (Path A drain)
After quorum confirmation, drain pending_reads entries whose read_index <=
last_applied. Required because pure-read batches never advance commit_index,
so handle_apply_completed (Path B) would never fire for them.


What Does This PR Do?

Fixes a linearizability violation where a partitioned multi-voter leader could
serve stale reads by bypassing quorum confirmation. After this fix, multi-voter
LinearizableRead is always deferred until either a quorum ACK (Path A) or SM
apply (Path B) confirms the leader is still current.

Type:

  • Bug Fix (with test)

Why Is This Needed?

Bug: A multi-voter leader in a minority partition would answer
LinearizableRead immediately from local state, bypassing the Raft §8 requirement
that leadership must be confirmed via a quorum heartbeat before serving reads.
Jepsen register workload with FAULTS=partition reproduces this as a
Knossos linearizability violation.

Fix: Gate Phase 3 fast-path on single_voter. Multi-voter reads enter
pending_reads and are drained only after quorum is confirmed in
handle_append_result.


Checklist

Required:

  • make test passes
  • Added tests for new code
  • Commits squashed to 1-2 logical units

Testing

How tested:

  • Unit tests (deterministic):

    • pending_reads_test::test_multi_voter_fast_path_linear_read_is_queued (T1) — Phase 3 no longer fast-paths multi-voter; fails without fix
    • pending_reads_test::test_pending_reads_drained_by_quorum_ack (T2) — Path A quorum ACK drains the queue; fails without fix
    • pending_reads_test::test_pending_reads_slow_path_drained_by_apply_completed (T3) — Path B regression guard
    • pending_reads_test::test_pending_reads_expired_by_tick (T4) — partition scenario: reads expire via tick()
    • pending_reads_test::test_pending_reads_cleared_on_stepdown (T5) — drain_read_buffer() sends Unavailable on step-down
    • client_read_test::test_linearizable_read_served_without_quorum_in_minority_partition — end-to-end regression through flush_cmd_buffers; fails without fix
  • Integration tests: make run-workload WORKLOAD=register FAULTS=partition TIME_LIMIT=120 (probabilistic; unit tests above are the deterministic guarantee)

For bug fixes:

  • Added test that fails without this fix (T1, T2, and client_read regression)

Does This Follow d-engine's Principles?

  • Solves a real problem for most users (not just my edge case)
  • Keeps implementation simple
  • Doesn't bloat the API surface

Reviewer Notes

Core logic change is 2 surgical edits (~20 lines). The rest is tests and comments.
Focus review on:

  1. leader_state.rs:~3450 — Phase 3 single_voter && guard
  2. leader_state.rs:~1885 — Path A drain loop in handle_append_result

Estimated review complexity:

  • Quick (< 100 lines)
  • Medium (< 300 lines) — core change is ~20 lines; 6 new tests

Summary by CodeRabbit

Bug Fixes

  • Fixed an issue (Bug #381) where linearizable read operations could be served prematurely without proper quorum confirmation, ensuring consistent read behavior across cluster states.

Tests

  • Added regression test suite for linearizable read behavior under various quorum and cluster conditions.

Review Change Stack

… multi-voter cluster

## Root cause

In Phase 3 of `execute_and_process_raft_rpc`, a multi-voter leader served
LinearizableRead immediately when `last_applied >= read_index`, with zero
quorum confirmation. A leader isolated in a minority partition would answer
stale reads — Raft §8 violation.

## Changes

**leader_state.rs — Phase 3**
Add `single_voter &&` guard: single-voter self IS the quorum (safe to serve
immediately); multi-voter reads always queue into `pending_reads` so they
cannot be served without a network round-trip.

**leader_state.rs — handle_append_result (Path A drain)**
After quorum confirmation, drain `pending_reads` entries whose read_index <=
last_applied. Required because pure-read batches never advance commit_index,
so handle_apply_completed (Path B) would never fire for them.
@coderabbitai

coderabbitai Bot commented May 18, 2026 •

Copy link
Copy Markdown

Warning

Rate limit exceeded

@JoshuaChi has exceeded the limit for the number of commits that can be reviewed per hour. Please wait 1 minute and 3 seconds before requesting another review.

You’ve run out of usage credits. Purchase more in the billing tab.

⌛ How to resolve this issue?

After the wait time has elapsed, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans have higher rate limits than the trial, open-source and free plans. In all cases, we re-allow further reviews after a brief timeout.

Please see our FAQ for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: e617218d-06a7-450f-ad4f-bc6e30631efa

📥 Commits

Reviewing files that changed from the base of the PR and between 88b8542 and ac026d0.

📒 Files selected for processing (1)
  • d-engine-core/src/raft_role/leader_state_test/client_read_test.rs
📝 Walkthrough

Walkthrough

This PR fixes Bug #381 in Raft linearizable read handling. The leader now correctly queues multi-voter linearizable read requests during Phase 3 instead of unsafely serving them, and drains queued reads when quorum ACK or state machine apply completion occurs.

Changes

Linearizable Read Pending Queue Bug Fix

Layer / File(s) Summary
Core bug fix: Phase-3 read routing and quorum-ACK draining
d-engine-core/src/raft_role/leader_state.rs
In execute_and_process_raft_rpc Phase 3, single-voter leaders serve immediately when last_applied >= read_index, while multi-voter reads are queued; in handle_append_result quorum-ACK path, queued reads with read_index <= last_applied are drained and served to clients.
Regression test for minority partition scenario
d-engine-core/src/raft_role/leader_state_test/client_read_test.rs
New test validates that multi-voter linearizable reads are deferred and not served immediately in minority partitions where followers do not respond, enforcing that quorum confirmation must occur before service.
Comprehensive pending_reads test suite
d-engine-core/src/raft_role/leader_state_test/mod.rs, d-engine-core/src/raft_role/leader_state_test/pending_reads_test.rs
Five integrated tests validate linearizable read queuing and draining: T1 confirms no fast-path for multi-voter, T2 validates quorum-ACK draining, T3 validates state-machine-apply draining, T4 validates timeout eviction, T5 validates leader step-down. Includes test helpers, leader fixture, and state-machine mocks.

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~25 minutes

Possibly related issues

  • deventlab/d-engine#381: The bug being fixed—linearizable reads in multi-voter clusters were served unsafely in minority partitions without quorum confirmation.

Possibly related PRs

  • deventlab/d-engine#229: Both PRs adjust Raft leader linearizable-read handling in leader_state.rs by deferring/serving reads based on state-machine catch-up via last_applied.
  • deventlab/d-engine#230: Both PRs modify the linearizable-read path in LeaderState to ensure read consistency through state-machine apply verification.
  • deventlab/d-engine#288: Both PRs touch Raft leader read-buffering behavior in leader_state.rs, specifically the draining and clearing paths on role transitions.

Poem

🐰 A queue of reads, once served too fast,
Now waits for quorum's word at last—
When ACKs bloom and state machines align,
The pending batch gets its time to shine! ✨

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly and specifically identifies the main bug fix: preventing LinearizableRead from being served without quorum confirmation in multi-voter Raft clusters, directly addressing issue #381.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch fix/381-linear-bug

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🧹 Nitpick comments (1)
d-engine-core/src/raft_role/leader_state.rs (1)

1888-1899: ⚡ Quick win

Extract shared pending-read drain helper to prevent future path drift.

Path A and handle_apply_completed both implement the same pending_reads drain loop. Centralizing this into one helper reduces divergence risk for future fixes in read serving behavior.

♻️ Suggested refactor
+fn drain_pending_reads_up_to(
+    &mut self,
+    last_applied: u64,
+    ctx: &RaftContext<T>,
+) {
+    let to_serve: Vec<u64> =
+        self.pending_reads.range(..=last_applied).map(|(k, _)| *k).collect();
+    for idx in to_serve {
+        if let Some(batch) = self.pending_reads.remove(&idx) {
+            self.execute_pending_reads(batch.requests, ctx);
+        }
+    }
+}
-        let reads_to_serve: Vec<_> =
-            self.pending_reads.range(..=last_index).map(|(k, _)| *k).collect();
-
-        for read_index in reads_to_serve {
-            if let Some(batch) = self.pending_reads.remove(&read_index) {
-                self.execute_pending_reads(batch.requests, ctx);
-            }
-        }
+        self.drain_pending_reads_up_to(last_index, ctx);
-                let last_applied = ctx.state_machine().last_applied().index;
-                let to_serve: Vec<u64> =
-                    self.pending_reads.range(..=last_applied).map(|(k, _)| *k).collect();
-                for idx in to_serve {
-                    if let Some(batch) = self.pending_reads.remove(&idx) {
-                        self.execute_pending_reads(batch.requests, ctx);
-                    }
-                }
+                self.drain_pending_reads_up_to(ctx.state_machine().last_applied().index, ctx);
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@d-engine-core/src/raft_role/leader_state.rs` around lines 1888 - 1899,
Extract the duplicated pending-reads drain loop into a single helper method
(e.g., drain_pending_reads(&mut self, ctx: &mut ContextType)) that iterates
self.pending_reads up to ctx.state_machine().last_applied().index, removes
entries and calls execute_pending_reads(batch.requests, ctx); then replace the
inline loop in the Path A block and the loop inside handle_apply_completed with
calls to this new drain_pending_reads helper so both paths share the same logic
and avoid drift. Ensure the helper has access to self.pending_reads, calls
execute_pending_reads, and uses the same last_applied calculation
(ctx.state_machine().last_applied().index).
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@d-engine-core/src/raft_role/leader_state_test/client_read_test.rs`:
- Around line 2185-2191: The test currently only checks
resp_rx.try_recv().is_err(), which can be true if the channel closed; after
calling flush() assert that the leader_state.pending_reads (or its accessor used
in the test) contains the queued LinearizableRead instance (e.g., matches the
request id or client token used in this test) to prove the read was actually
enqueued awaiting quorum; locate the pending_reads collection in the same test
scope (or via leader_state.pending_reads) and add a direct assert that its
length increased or that it contains the specific read entry before proceeding
to handle_append_result.

---

Nitpick comments:
In `@d-engine-core/src/raft_role/leader_state.rs`:
- Around line 1888-1899: Extract the duplicated pending-reads drain loop into a
single helper method (e.g., drain_pending_reads(&mut self, ctx: &mut
ContextType)) that iterates self.pending_reads up to
ctx.state_machine().last_applied().index, removes entries and calls
execute_pending_reads(batch.requests, ctx); then replace the inline loop in the
Path A block and the loop inside handle_apply_completed with calls to this new
drain_pending_reads helper so both paths share the same logic and avoid drift.
Ensure the helper has access to self.pending_reads, calls execute_pending_reads,
and uses the same last_applied calculation
(ctx.state_machine().last_applied().index).
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: f4df629a-0e37-4aa2-b8ec-5790e951a240

📥 Commits

Reviewing files that changed from the base of the PR and between 8ed1b89 and 88b8542.

📒 Files selected for processing (4)
  • d-engine-core/src/raft_role/leader_state.rs
  • d-engine-core/src/raft_role/leader_state_test/client_read_test.rs
  • d-engine-core/src/raft_role/leader_state_test/mod.rs
  • d-engine-core/src/raft_role/leader_state_test/pending_reads_test.rs

Comment thread d-engine-core/src/raft_role/leader_state_test/client_read_test.rs
@codecov

codecov Bot commented May 18, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 97.10983% with 10 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
...rc/raft_role/leader_state_test/client_read_test.rs 93.67% 5 Missing ⚠️
.../raft_role/leader_state_test/pending_reads_test.rs 98.12% 5 Missing ⚠️

📢 Thoughts on this report? Let us know!

@JoshuaChi
JoshuaChi merged commit e4654a1 into main May 18, 2026
9 checks passed
@JoshuaChi
JoshuaChi deleted the fix/381-linear-bug branch May 18, 2026 11:35
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant