Starting fresh with the latest version? No migration needed - skip this guide and go to Quick Start.
Starting from v0.2.0, the WAL (Write-Ahead Log) format for file-based state machines has changed to support absolute expiration time semantics.
Old Format (pre-v0.2.0):
Entry fields: ..., ttl_secs: u32 (4 bytes, relative TTL)
New Format (v0.2.0+):
Entry fields: ..., expire_at_secs: u64 (8 bytes, absolute expiration time in UNIX seconds)
- Crash Safety: Absolute expiration time ensures TTL correctness across restarts
- Deterministic Semantics: Matches industry-standard lease semantics (absolute expiry)
- No TTL Reset: TTL no longer resets on node restart
- Reading old WAL files will cause deserialization errors
- Node startup will fail if old WAL files are present
Best for: Development, testing, or non-production environments
- Backup your data (optional, if you need to preserve state)
- Stop the node gracefully
- Delete old WAL directory:
rm -rf /path/to/storage/wal/* - Start with v0.2.0
Best for: Production clusters with replication (3+ nodes)
Since d-engine uses Raft consensus, you can perform a rolling upgrade:
- Ensure cluster is healthy (all nodes synchronized)
- For each node:
- Stop the node gracefully (ensure data is persisted)
- Upgrade to v0.2.0
- Clear WAL directory:
rm -rf /path/to/storage/wal/* - Start the node (it will catch up from other nodes)
- Repeat for all nodes one by one
The cluster will remain available during the upgrade (assuming you have 3+ nodes).
Best for: Large WAL files or single-node setups
- On old version (pre-v0.2.0):
- Trigger a snapshot to persist current state
- Wait for snapshot to complete
- Verify snapshot file exists:
/path/to/storage/snapshots/
- Upgrade to v0.2.0
- Clear WAL:
rm -rf /path/to/storage/wal/* - Start node - it will restore from the snapshot
After upgrading, verify:
# Check node starts without errors
journalctl -u d-engine -f
# Verify TTL entries expire correctly
# (create a key with TTL and wait for expiration)
# Check logs for WAL-related errors
grep "WAL" /var/log/d-engine.log| Aspect | Old (pre-v0.2.0) | New (v0.2.0+) |
|---|---|---|
| TTL Storage | Relative (seconds from now) | Absolute (UNIX timestamp) |
| After Restart | TTL resets 🔄 | TTL preserved ✅ |
| WAL Replay | All entries loaded | Expired entries skipped ✅ |
| Expiration Semantics | Relative TTL | Absolute timestamp ✅ |
| Crash Safe | ❌ No | ✅ Yes |
- Documentation: See examples/ for updated usage patterns
- Issues: Report migration issues on GitHub Issues
- Discussion: Ask questions in GitHub Discussions
| Version | WAL Format | Wire Protocol | Migration Required |
|---|---|---|---|
| v0.1.x | Relative TTL | Compatible | - |
| v0.2.0–v0.2.2 | Absolute expiration | Compatible | ✅ Yes (clear WAL from v0.1.x) |
| v0.2.3 | Same as v0.2.0+ | Incompatible | ✅ Yes (protobuf enum changes + API changes) |
| v0.2.4 | Same as v0.2.0+ | Compatible (additive) | ✅ Yes (delete snapshot/ — format changed to CF export) |
| v0.2.5 | Same as v0.2.0+ | Compatible (additive) | lease.enabled; remove entry_term() from custom SM; use start_node instead of start_custom_with_config) |
v0.2.3 introduces breaking wire protocol changes due to protobuf enum value shifts to comply with buf lint standards.
Wire Protocol Incompatibility:
- v0.2.3 nodes CANNOT communicate with v0.2.2 or earlier nodes
- No rolling upgrade possible - all cluster nodes must upgrade simultaneously
- Client SDKs must be upgraded to v0.2.3 to connect to upgraded clusters
| Role | Old Value | New Value | New Constant |
|---|---|---|---|
| - | - | 0 | NODE_ROLE_UNSPECIFIED |
| Follower | 0 | 1 | NODE_ROLE_FOLLOWER |
| Candidate | 1 | 2 | NODE_ROLE_CANDIDATE |
| Leader | 2 | 3 | NODE_ROLE_LEADER |
| Learner | 3 | 4 | NODE_ROLE_LEARNER |
| Status | Old Value | New Value | New Constant |
|---|---|---|---|
| - | - | 0 | NODE_STATUS_UNSPECIFIED |
| Promotable | 0 | 1 | NODE_STATUS_PROMOTABLE |
| ReadOnly | 1 | 2 | NODE_STATUS_READ_ONLY |
| Active | 2 | 3 | NODE_STATUS_ACTIVE |
| Error | Old Value | New Value | New Constant |
|---|---|---|---|
| - | - | 0 | ERROR_CODE_UNSPECIFIED |
| NotLeader | 1 | 1 | ERROR_CODE_NOT_LEADER |
| ... (others unchanged) |
- Enum Prefixes: All enum values now have proper prefixes (
NODE_ROLE_*,NODE_STATUS_*,ERROR_CODE_*) - Field Naming: All fields now use snake_case naming (
leader_id,prev_log_index,last_log_index, etc.)
Update any TOML configuration files that reference enum values:
Old (v0.2.2):
[cluster]
node_id = 1
role = 0 # Follower
status = 0 # PromotableNew (v0.2.3):
[cluster]
node_id = 1
role = 1 # NODE_ROLE_FOLLOWER
status = 1 # NODE_STATUS_PROMOTABLESince the wire protocol is incompatible, you must upgrade all nodes at once:
- Schedule maintenance window (cluster will be unavailable during upgrade)
- Stop all nodes in the cluster
- Upgrade binaries to v0.2.3 on all nodes
- Update configuration files (see Step 1)
- Start all nodes simultaneously
- Verify cluster health (check logs, run health checks)
For Production Clusters:
If you require high availability during upgrade:
- Set up a parallel v0.2.3 cluster (new hardware/instances)
- Migrate data to the new cluster (application-level migration)
- Switch traffic to new cluster
- Decommission old cluster
All client applications must upgrade their d-engine SDK to v0.2.3:
Cargo.toml:
[dependencies]
d-engine = { version = "0.2.3", features = ["client"] }Rebuild and redeploy all client applications before connecting to upgraded cluster.
After upgrade, verify:
# Check all nodes started successfully
journalctl -u d-engine -f
# Verify cluster health
curl http://localhost:8080/health
# Test basic operations
d-engine-cli put test-key test-value
d-engine-cli get test-keyv0.2.3 introduces breaking API changes to unify client interfaces and improve developer experience.
Old (v0.2.2):
use d_engine::client::KvClient;
use d_engine::client::KvError;
async fn example(client: impl KvClient) -> Result<(), KvError> {
// ...
}New (v0.2.3):
use d_engine::client::ClientApi;
use d_engine::client::ClientApiError;
async fn example(client: impl ClientApi) -> Result<(), ClientApiError> {
// ...
}Migration Steps:
- Replace
KvClientwithClientApiin trait bounds - Replace
KvErrorwithClientApiErrorin error handling - Update imports:
use d_engine::client::{ClientApi, ClientApiError};
Old (v0.2.2): Default = MemFirst (write to memory, async flush to disk)
New (v0.2.3): Default = DiskFirst (Raft protocol compliance)
Migration:
If you want to restore v0.2.2 behavior, add to config:
[raft.persistence]
persistence_strategy = "MemFirst"MemFirst trades durability for performance. Only use in scenarios where data loss is acceptable.
- CompareAndSwap (CAS): New atomic operation added
- Drain-based batching: Performance improvements (no API changes)
- Client::refresh(): New method for leader rediscovery
Last Updated: February 2026
If you implemented a custom StateMachine, update apply_chunk:
// Old (v0.2.3)
fn apply_chunk(&mut self, entries: Vec<Entry>) -> Result<(), ...>
// New (v0.2.4) — slice of ApplyEntry (decoded key/value/TTL)
fn apply_chunk(&mut self, entries: &[ApplyEntry]) -> Result<(), ...>ApplyEntry carries decoded fields directly — no proto parsing needed in your impl.
v0.2.4 removes internal implementation details that were accidentally exposed as pub.
All removed items were internal — they were never part of the documented public API.
This method returned client_id (not a node ID), which was semantically incorrect.
// Old (v0.2.3) — broken semantics, removed
let id = client.node_id();
// New (v0.2.4) — use EmbeddedEngine instead
let id = engine.node_id();get_linearizable(), get_lease(), and get_eventual() are now only available via
the ClientApi trait. The return type is unified to Option<Bytes> (was Option<ClientResult>).
// Old (v0.2.3)
let result = client.get_linearizable("key").await?;
let value = result.map(|r| r.value); // extra unwrap needed
// New (v0.2.4) — add trait import, get Bytes directly
use d_engine_client::ClientApi;
let value = client.get_linearizable("key").await?; // Option<Bytes>These were internal lifecycle methods. Use EmbeddedEngine::wait_ready() instead.
// Old (v0.2.3)
node.set_rpc_ready(true);
let ready = node.is_rpc_ready();
// New (v0.2.4) — use the engine-level API
engine.wait_ready(Duration::from_secs(5)).await?;// Old (v0.2.3)
let node_id = node.node_config.cluster.node_id;
// New (v0.2.4)
let node_id = node.node_id();Use the documented constructors instead.
// Old (v0.2.3)
NodeBuilder::init(config, shutdown_rx)
// New (v0.2.4)
NodeBuilder::new(None, shutdown_rx).node_config(config)
// or
NodeBuilder::from_cluster_config(cluster_config, shutdown_rx)This type was defined but never used. Remove any references to it.
These are internal connection pool types. Use Client and ClientBuilder instead.
v0.2.4 changes the internal snapshot format from RocksDB checkpoint (v0.2.3) to CF export (v0.2.4). Existing v0.2.3 snapshots cannot be loaded by v0.2.4 nodes.
- Ensure the cluster is healthy — all nodes synchronized before upgrading.
- Delete old snapshots before starting the upgraded node:
rm -rf /path/to/db_root_dir/snapshot/
- Upgrade to v0.2.4 — the node replays from WAL on first start.
- A new snapshot is created automatically once the log-size threshold is reached.
Ensure sufficient disk space for WAL replay. Nodes that have fallen far behind may trigger a snapshot install from the leader instead of local WAL replay.
Watch buffer overflow is no longer silent. When a per-watcher channel fills up, the server now:
- Sends a
CANCELEDsentinel (event_type = WATCH_EVENT_TYPE_CANCELED,error = WATCH_BUFFER_OVERFLOW) - Unregisters the watcher — no further events are delivered on that stream
Previously, events were silently dropped and the watcher was silently unregistered with no client notification.
Non-breaking — clients that handle stream close (None / Status::UNAVAILABLE) continue to work.
The CANCELED event simply provides early notification before the stream closes.
Recommended: update watch consumers to handle CANCELED for clean re-sync:
while let Some(event) = stream.next().await {
match event {
Ok(ev) if ev.event_type == WatchEventType::Canceled as i32 => {
// buffer overflow: re-sync current state then re-register
let current = client.read(ev.key.clone()).await?;
process_snapshot(current);
stream = client.watch(ev.key).await?;
}
Ok(ev) => handle_event(ev),
Err(e) => return Err(e),
}
}See Watch Feature Guide for details.
The entry_term() method has been removed from the StateMachine trait. Term lookup belongs
to the Raft log layer, not the state machine — this was always an architectural layering issue.
StateMachine implementations.
If you implemented a custom StateMachine, delete the entry_term method from your impl block.
Built-in implementations (FileStateMachine, RocksDBStateMachine, SledStateMachine) are
already updated and require no changes.
// Old (v0.2.4) — delete this block
fn entry_term(&self, entry_id: u64) -> Option<u64> {
Some(1)
}Compilation will fail with "method not found in trait" if you forget to remove it.
apply_snapshot_from_file now returns Result<SnapshotApplyResult> instead of Result<()>, so
the caller knows whether the install actually ran or was a no-op instead of every outcome looking
identical.
pub enum SnapshotApplyResult {
Applied { last_included: LogId }, // install actually ran
IgnoredStale { current: LogId }, // incoming.index < current — no-op, not an error
IgnoredDuplicate { current: LogId }, // incoming == current, same term — no-op, not an error
}StateMachine implementations.
// Old (v0.2.4)
async fn apply_snapshot_from_file(
&self,
metadata: &SnapshotMetadata,
path: PathBuf,
) -> Result<()> {
// ... install logic ...
Ok(())
}
// New (v0.2.5)
async fn apply_snapshot_from_file(
&self,
metadata: &SnapshotMetadata,
path: PathBuf,
) -> Result<SnapshotApplyResult> {
// ... same install logic ...
Ok(SnapshotApplyResult::Applied { last_included: new_last_included })
}Return IgnoredStale for a snapshot strictly older than what's already applied, and
IgnoredDuplicate only when the index and term both match what's already applied — that's
Raft's own idempotency check on the install call, not a failure. An equal index with a
different term is not a duplicate: it means two different leadership terms produced
conflicting content at the same boundary, which the public error contract classifies as fatal —
return Err(SnapshotError::BoundaryConflict { .. }) for that case. Reserve Err generally for a
genuine failure (checksum mismatch, I/O error, corrupted data, or this boundary conflict).
Built-in implementations (FileStateMachine, RocksDBStateMachine) are already updated.
Compilation fails with "incompatible type for trait" until a custom implementation is updated.
The learner's eager startup snapshot pull is removed. A new learner no longer calls
Transport::request_snapshot_from_leader to fetch a snapshot on startup; it relies on the leader's
PUSH replication loop instead (AppendEntries, or InstallSnapshot when its log falls behind the
purge boundary).
Transport and StateMachineHandler implementations.
Two trait methods are removed:
Transport::request_snapshot_from_leader(...)— delete this method from any customTransport.StateMachineHandler::apply_snapshot_stream_from_leader(...)— delete this method from any customStateMachineHandler.
Compilation fails with "method ... is not a member of trait" until the custom implementation is updated. The built-in implementations are already updated.
The public API now exposes start_node() directly instead of routing through a thin wrapper.
// Old (v0.2.4) — removed
EmbeddedEngine::start_custom_with_config(storage, sm, config).await?;
// New (v0.2.5)
EmbeddedEngine::start_node(config, storage, sm).await?; // embedded
StandaloneEngine::start_node(config, storage, sm, shutdown_rx).await?; // standalonestart_node calls config.validate() internally — callers may pass validated or unvalidated configs.
Snapshot last_included was previously computed by subtracting retained_log_entries
from last_applied, introducing a label/data mismatch. Fixed: last_included == last_applied
always.
Transparent — no action required. Nodes that installed snapshots with the old label may have double-applied entries; upgrading to v0.2.5 prevents this from recurring. Existing snapshots are compatible.
The enabled flag under [raft.state_machine.lease] has been removed. TTL expiration is now always active — no opt-in required.
In v0.2.4, omitting enabled = true caused a fatal crash on put_with_ttl. This bug is fixed in v0.2.5.
If your config contains enabled = true or enabled = false, remove the line:
# Old (v0.2.4) — remove this line
[raft.state_machine.lease]
enabled = true # ← delete
# New (v0.2.5) — TTL always active, no flag needed
[raft.state_machine.lease]
cleanup_interval_ms = 1000Impact: None if you do not touch the config — unrecognised fields are ignored. The only behavioral change is that TTL expiration is now unconditionally enabled.
This section is optional — both fields have defaults and existing configs work without changes.
[raft.read_actor]
# mpsc channel buffer for Eventual/LeaseRead fast path.
# Rule of thumb: ≥ 2× peak concurrent readers. Default: 512.
channel_capacity = 512
# Max reads drained per wakeup. Default: 100.
max_drain = 100No migration action required unless you want to tune read concurrency.
data_dir (formerly db_root_dir) is no longer a config file field. Pass it as an explicit argument to the engine constructor instead:
// Before (v0.2.4)
let engine = EmbeddedEngine::start_with("config.toml").await?;
// config.toml had [cluster] data_dir = "./db"
// After (v0.2.5)
let engine = EmbeddedEngine::start_with("./db", "config.toml").await?;All engine constructors (start, start_with, start_custom, start_node, run, run_with, run_custom) now require data_dir as the first argument. Remove [cluster] data_dir / [cluster] db_root_dir from all config files — the entry is silently ignored if left in place, but it's misleading.
[cluster] log_dir is removed (it was validated but never consumed).
snapshots_dir is no longer configurable — always data_dir/snapshots.
Only affects code that plugged in a custom storage engine or state machine directly via NodeBuilder. If you used EmbeddedEngine or StandaloneEngine, no change needed.
// Before (v0.2.4)
let node = NodeBuilder::new(data_dir, None, shutdown_rx)
.storage_engine(storage_engine)
.state_machine(state_machine)
.start()
.await?;
node.run().await?;
// After (v0.2.5)
EmbeddedEngine::start_custom(data_dir, storage_engine, state_machine, None).await?;
// or, for a standalone gRPC server:
StandaloneEngine::run_custom(data_dir, storage_engine, state_machine, shutdown_rx, None).await?;start_custom/run_custom cover the same ground — explicit data_dir, custom StorageEngine/StateMachine, optional config file.
Last Updated: August 2026