Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
24 commits
Select commit Hold shift + click to select a range
9e2619b
fix #446: fence FsyncCoordinator against truncation, clamp durable/pe…
JoshuaChi Aug 31, 2026
340b518
fix #446: gate commit quorum and follower ACKs on fsync-durable index…
JoshuaChi Sep 2, 2026
d9a5b17
chore #446: bump grpc-go and x/net to resolve Dependabot alerts (#86-91)
JoshuaChi Sep 2, 2026
97eb1c0
fix #446: single-writer durable_index/persisted_index, close truncati…
JoshuaChi Sep 3, 2026
559708a
fix #446: revert synchronous persist in append_entries (leader perf r…
JoshuaChi Sep 6, 2026
81afa8b
fix #446: keep withheld-ACK queue alive across role transitions
JoshuaChi Sep 8, 2026
368121f
fix #446: merge IO wakeup sources, frontier-based persist, replace_ra…
JoshuaChi Sep 10, 2026
642c778
test #446: persist-frontier + fsync-coordinator coverage
JoshuaChi Sep 10, 2026
9f8ac3f
fix #446: order fsync pending mark term-first to unstick durable_inde…
JoshuaChi Sep 10, 2026
33793fa
fix #446: drain queued IOTask::Persist safely, add fsync/batch metric…
JoshuaChi Sep 13, 2026
dc5096e
fix #446: report post-commit term in AppendEntries reply, dedup per-A…
JoshuaChi Sep 16, 2026
2dd0372
perf #446: instrument write-path stages, enable RocksDB manual WAL flush
JoshuaChi Sep 26, 2026
a20f2ef
test #446: fix pre-existing flaky tests found while validating this b…
JoshuaChi Sep 26, 2026
f0d3d62
test #446: add aws benchmark
JoshuaChi Sep 26, 2026
ea39ae8
perf #446: merge log IO into Raft core via RaftLogCore
JoshuaChi Sep 27, 2026
abde9ba
perf #446: fsync on tokio blocking pool, RocksDB bg jobs 4→2
JoshuaChi Sep 27, 2026
2f45114
refactor #446: delete BufferedRaftLog, RaftLogCore is now the sole Ra…
JoshuaChi Sep 27, 2026
4a054dd
fix #446: replication in-flight window and commit correctness
JoshuaChi Oct 1, 2026
1d165e6
fix #446: end stream_append_entries task once inbound closes and pend…
JoshuaChi Oct 1, 2026
50bf03c
test #446: scope snapshot-transfer purge wait to the leader's node id
JoshuaChi Oct 1, 2026
98ff894
fix #446: key pending ACKs by (term, index) so a stale-term entry can…
JoshuaChi Oct 1, 2026
f1805d1
test #446: wait for fsync completion instead of sleeping in concurren…
JoshuaChi Oct 1, 2026
9b9cad7
fix #446: surface fsync/persist failures in FsyncWorker and flush(); …
JoshuaChi Oct 1, 2026
7c3fe5c
fix #446: return from broadcast_vote_requests once a majority has voted
JoshuaChi Oct 1, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 2 additions & 2 deletions .config/nextest.toml
Original file line number Diff line number Diff line change
Expand Up @@ -19,7 +19,7 @@ threads-required = 'num-cpus'
# Serialize wall-clock perf assertions — CPU contention causes false failures, not the
# assertions. Don't #[ignore] them, isolate instead.
[[profile.default.overrides]]
filter = "test(storage_buffered_raft_log::performance_test)"
filter = "test(raft_log_core::performance_test)"
threads-required = 'num-cpus'

# Per-directory, not per-file — a few dirs also hold light tests, throttled as fallout.
Expand Down Expand Up @@ -47,7 +47,7 @@ filter = "test(tmp_db) | test(config_path_env) | test(data_dir_lock) | test(sing
threads-required = 'num-cpus'

[[profile.ci.overrides]]
filter = "test(storage_buffered_raft_log::performance_test)"
filter = "test(raft_log_core::performance_test)"
threads-required = 'num-cpus'

[[profile.ci.overrides]]
Expand Down
4 changes: 3 additions & 1 deletion .dockerignore
Original file line number Diff line number Diff line change
@@ -1,3 +1,5 @@
**/target/
.git/
examples/
examples/*
!examples/three-nodes-standalone
!examples/client-usage-standalone
6 changes: 6 additions & 0 deletions .github/PULL_REQUEST_TEMPLATE.md
Original file line number Diff line number Diff line change
Expand Up @@ -61,6 +61,12 @@

---

## AI Assistance

- [ ] This PR was written in part with the assistance of generative AI. All ideas and architecture decisions are mine; I have fully reviewed all changes.

---

## Reviewer Notes

(Optional: anything reviewers should focus on)
Expand Down
23 changes: 23 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -32,6 +32,18 @@ All notable changes to this project will be documented in this file.
returns immediately, and entries arriving during an in-flight fsync are coalesced into the same physical
disk flush. Storage-level group commit is restored without artificial batching windows.

- **🛑 Client-acknowledged writes could be lost on correlated power loss (#446)**: Raft commit quorum
counted the leader's own log contribution using its in-memory tail (`last_entry_id()`), not its
fsync-confirmed position (`durable_index()`) — a write could reach a majority-looking commit index,
and be acknowledged to the client, before enough replicas had actually synced it to disk. If those
nodes then lost power before their next fsync, the acknowledged write was gone. Fixed: leader quorum
calculation, follower `AppendEntries` ACK timing (a follower now withholds its response until its own
`durable_index` reaches the acknowledged entry), and single-voter clusters (previously exempted from
this class of fix, see #329) all gate on `durable_index`. RPO=0 for acknowledged writes is now a
mandatory invariant. Net effect: write acknowledgment latency now includes fsync time on a quorum of
replicas — see [Throughput Optimization Guide](./d-engine/src/docs/performance/throughput-optimization-guide.md)
for tuning `idle_flush_interval_ms`.

### Changed

- **MSRV raised to Rust 1.89**: The `data_dir` startup lock (prevents two node processes from
Expand Down Expand Up @@ -65,6 +77,17 @@ All notable changes to this project will be documented in this file.

- **`NodeBuilder` is no longer public** — use `EmbeddedEngine::start_custom`/`StandaloneEngine::run_custom` to plug in a custom storage engine or state machine. See [Migration Guide](./MIGRATION_GUIDE.md) for details.

- **⚠️ `[raft] ordered_channel_capacity` renamed to `max_pending_append_responses`** (#446): Follows the
gRPC `AppendEntries` forwarder rewrite (`FuturesUnordered`-based, no longer strict-FIFO) that shipped
alongside the durability fix above. Old field name is silently ignored, not an error — update existing
configs to the new name to keep the setting in effect.

- **⚠️ `[raft.persistence] strategy` removed** (#446): `PersistenceStrategy` was a single-variant enum
(`MemFirst`) left over from #268; its only meaning now lives in whether an entry has reached
`durable_index`, which is no longer a configurable choice. Existing configs setting `strategy =
"MemFirst"` or `"DiskFirst"` are silently ignored, not an error — remove the field, `flush_policy`
is the only persistence knob now.

---

## [v0.2.4] - 2026-05-23
Expand Down
8 changes: 4 additions & 4 deletions Cargo.lock

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

3 changes: 1 addition & 2 deletions Cargo.toml
Original file line number Diff line number Diff line change
Expand Up @@ -69,8 +69,7 @@ opt-level = 3 # All dependency optimization levels
[profile.release]
incremental = true
debug = false
# opt-level = 3
opt-level = 'z' # Optimized for minimum size
opt-level = 3
overflow-checks = true
lto = true # Eliminate redundant code and reduce size
codegen-units = 1 # Reduce parallel code generation units to improve optimization effect
Expand Down
1 change: 1 addition & 0 deletions benches/embedded-bench/Cargo.toml
Original file line number Diff line number Diff line change
Expand Up @@ -18,3 +18,4 @@ futures = "0.3"
tracing-subscriber = { version = "0.3", features = ["env-filter"] }
tracing = { version = "0.1" }
metrics-exporter-prometheus = "0.17.2"
console-subscriber = "0.5"
75 changes: 65 additions & 10 deletions benches/embedded-bench/Makefile
Original file line number Diff line number Diff line change
@@ -1,27 +1,25 @@
# Makefile for embedded-bench
# Provides benchmark commands matching embedded-bench/reports/v0.2.0/report_v0.2.0_final.md

.PHONY: help build clean test-single-write test-high-conc-write test-linearizable-read test-lease-read test-eventual-read test-hot-key all-tests
.PHONY: help build clean clean-log-db test-single-write test-high-conc-write test-linearizable-read test-lease-read test-eventual-read test-hot-key all-tests \
all-ram-tests ramdisk-create clean-ram-log-db ramdisk-release

# ===============================
# Global Variables
# ===============================
BENCH_BIN := ./target/release/embedded-bench
CONFIG_DIR := ./config
LOG_LEVEL ?= warn

# Node selection (default: n1)
NODE ?= n1
CONFIG_PATH := $(CONFIG_DIR)/$(NODE).toml
SINGLE_NODE_CONFIG_PATH := $(CONFIG_DIR)/single_node.toml
DATA_DIR := ./data/$(NODE)
SINGLE_NODE_DATA_DIR := ./data/single-node

# Metrics port per node — same scheme as examples/three-nodes-standalone
# (8081/8082/8083), reused because the two setups never run concurrently.
METRICS_PORT := $(if $(filter n2,$(NODE)),8082,$(if $(filter n3,$(NODE)),8083,8081))

# ===============================
# Global Variables
# ===============================
LOG_LEVEL ?= warn

# Common parameters (matching Standalone tests)
KEY_SIZE := 8
VALUE_SIZE := 256
Expand All @@ -35,6 +33,25 @@ CLIENTS ?= 1000
VERIFY ?= false
VERIFY_FLAG := $(if $(filter true,$(VERIFY)),--verify-write,)

# RAM disk (macOS only, optional) — isolates this node's storage on its own
# independent RAM disk volume, to strip physical-disk latency out of a
# benchmark. Not a general fix for unrelated flakiness — see
# tickets/milestones/v0.2.5/446-perf-batching-measurement-2026-09-13.md.
# One-shot: `make all-ram-tests`. Manual/single-node: add RAMDISK=true to
# any test-* target (needs the other two nodes already running for quorum).
RAMDISK ?= false
RAMDISK_SIZE_MB ?= 1024
RAMDISK_SECTORS := $(shell echo $$(( $(RAMDISK_SIZE_MB) * 2048 )))
RAMDISK_VOLUMES := RAMDisk1 RAMDisk2 RAMDisk3
RAMDISK_INDEX := $(if $(filter n2,$(NODE)),2,$(if $(filter n3,$(NODE)),3,1))

ifeq ($(RAMDISK),true)
DATA_DIR := /Volumes/RAMDisk$(RAMDISK_INDEX)/$(NODE)
SINGLE_NODE_DATA_DIR := /Volumes/RAMDisk1/single-node
else
DATA_DIR := ./data/$(NODE)
SINGLE_NODE_DATA_DIR := ./data/single-node
endif

# On macOS with Homebrew: auto-detect compression lib paths to skip bundled C++
# compilation of RocksDB dependencies, which fails under macOS 26 + Xcode 26
Expand All @@ -58,7 +75,6 @@ ifneq ($(ZSTD_PREFIX),)
BREW_ROCKSDB_ENV += ZSTD_LIB_DIR=$(ZSTD_PREFIX)/lib
endif


help:
@echo "Embedded-bench Makefile - Performance Testing"
@echo ""
Expand All @@ -83,6 +99,14 @@ help:
@echo "Run All:"
@echo " make all-tests Run all benchmark tests"
@echo ""
@echo "RAM Disk Cluster (macOS only, optional — isolates disk latency):"
@echo " make all-ram-tests NODE=nX Create RAM disks, run full --batch suite on this node (like all-tests)"
@echo " Run in 3 terminals with NODE=n1/n2/n3 to form the cluster"
@echo " make ramdisk-create Create RAMDisk1/2/3 if not already mounted"
@echo " make clean-ram-log-db Wipe node data off the RAM disks (keeps them mounted)"
@echo " make ramdisk-release Unmount RAMDisk1/2/3, freeing the memory"
@echo " Add RAMDISK=true to any single test-* target (needs the other 2 nodes already running)"
@echo ""
@echo "Examples:"
@echo " make test-linearizable-read # Run on n1 (default)"
@echo " make test-linearizable-read NODE=n2 # Run on n2"
Expand All @@ -109,6 +133,7 @@ clean-log-db:
rm -rf ./logs/*
rm -rf ./data/*
rm -rf ./snapshots/*

# ============================================
# Write Performance Tests
# ============================================
Expand Down Expand Up @@ -262,4 +287,34 @@ all-tests: build
put
@echo ""
@echo "Compare results with Standalone mode:"
@echo " Standalone report: ../../benches/standalone-bench/reports/v0.2.2/report_v0.2.2.md"

# ============================================
# RAM Disk Cluster (macOS only, optional)
# ============================================
# Isolates each node's storage on its own independent RAM disk volume — use
# this to strip physical-disk latency out of a benchmark (e.g. to study
# fsync/scheduling behavior in isolation), not as a general fix for
# unrelated flakiness. See tickets/milestones/v0.2.5/446-perf-batching-measurement-2026-09-13.md.

# Ensures the RAM disks exist, then runs the full --batch suite for this
# node on its own RAM disk volume — same scope/shape as `make all-tests
# NODE=nX`, just on RAM disk. This is single-node, like all-tests: run it in
# 3 separate terminals with NODE=n1/n2/n3 to form the cluster, e.g.
# make all-ram-tests NODE=n2 CLIENTS=100
all-ram-tests: ramdisk-create
$(MAKE) all-tests RAMDISK=true NODE=$(NODE) CLIENTS=$(CLIENTS)

# Idempotent — mounts any of RAMDisk1/2/3 that aren't already present.
ramdisk-create:
@[ "$$(uname)" = "Darwin" ] || { echo "RAM disk targets need macOS."; exit 1; }
@for v in $(RAMDISK_VOLUMES); do \
[ -d "/Volumes/$$v" ] || diskutil erasevolume HFS+ $$v `hdiutil attach -nomount ram://$(RAMDISK_SECTORS)` >/dev/null; \
done

# Wipes node data off the RAM disks; keeps the volumes mounted.
clean-ram-log-db:
@for v in $(RAMDISK_VOLUMES); do rm -rf /Volumes/$$v/*; done

# Unmounts the RAM disks entirely, releasing the memory back to the OS.
ramdisk-release:
@for v in $(RAMDISK_VOLUMES); do diskutil eject /Volumes/$$v 2>/dev/null || true; done
12 changes: 5 additions & 7 deletions benches/embedded-bench/config/n1.toml
Original file line number Diff line number Diff line change
Expand Up @@ -15,13 +15,6 @@ max_drain = 1024
# Maximum number of commands to accumulate in a single batch during drain operations
max_batch_size = 200

[raft.persistence]
strategy = "MemFirst"
flush_policy = { Batch = { idle_flush_interval_ms = 1000 } }
# Maximum number of log entries to buffer in memory
# when using async persistence strategies (MemFirst/Batched)
max_buffered_entries = 10000

[raft.metrics]
enable_backpressure = false
enable_batch = false
Expand All @@ -31,3 +24,8 @@ enable = false

[storage]
unified_db = false

[raft.replication]
max_inflight_append_requests = 256
append_entries_max_entries_per_replication = 256
replication_send_queue_capacity = 1024
12 changes: 5 additions & 7 deletions benches/embedded-bench/config/n2.toml
Original file line number Diff line number Diff line change
Expand Up @@ -15,13 +15,6 @@ max_drain = 1024
# Maximum number of commands to accumulate in a single batch during drain operations
max_batch_size = 200

[raft.persistence]
strategy = "MemFirst"
flush_policy = { Batch = { idle_flush_interval_ms = 1000 } }
# Maximum number of log entries to buffer in memory
# when using async persistence strategies (MemFirst/Batched)
max_buffered_entries = 10000

[raft.metrics]
enable_backpressure = false
enable_batch = false
Expand All @@ -31,3 +24,8 @@ enable = false

[storage]
unified_db = false

[raft.replication]
max_inflight_append_requests = 256
append_entries_max_entries_per_replication = 256
replication_send_queue_capacity = 1024
12 changes: 5 additions & 7 deletions benches/embedded-bench/config/n3.toml
Original file line number Diff line number Diff line change
Expand Up @@ -15,13 +15,6 @@ max_drain = 1024
# Maximum number of commands to accumulate in a single batch during drain operations
max_batch_size = 200

[raft.persistence]
strategy = "MemFirst"
flush_policy = { Batch = { idle_flush_interval_ms = 1000 } }
# Maximum number of log entries to buffer in memory
# when using async persistence strategies (MemFirst/Batched)
max_buffered_entries = 10000

[raft.metrics]
enable_backpressure = false
enable_batch = false
Expand All @@ -31,3 +24,8 @@ enable = false

[storage]
unified_db = false

[raft.replication]
max_inflight_append_requests = 256
append_entries_max_entries_per_replication = 256
replication_send_queue_capacity = 1024
Loading
Loading