Skip to content

Latest commit

 

History

History
726 lines (485 loc) · 22.6 KB

File metadata and controls

726 lines (485 loc) · 22.6 KB

Migration Guide for d-engine

🎯 For New Users

Starting fresh with the latest version? No migration needed - skip this guide and go to Quick Start.


🚨 For v0.1.x Users: WAL Format Change

What Changed

Starting from v0.2.0, the WAL (Write-Ahead Log) format for file-based state machines has changed to support absolute expiration time semantics.

Old Format (pre-v0.2.0):

Entry fields: ..., ttl_secs: u32 (4 bytes, relative TTL)

New Format (v0.2.0+):

Entry fields: ..., expire_at_secs: u64 (8 bytes, absolute expiration time in UNIX seconds)

Why This Change?

  • Crash Safety: Absolute expiration time ensures TTL correctness across restarts
  • Deterministic Semantics: Matches industry-standard lease semantics (absolute expiry)
  • No TTL Reset: TTL no longer resets on node restart

Impact

⚠️ WAL files from pre-v0.2.0 are NOT compatible with v0.2.0+

  • Reading old WAL files will cause deserialization errors
  • Node startup will fail if old WAL files are present

Migration Strategies

Option 1: Clean Start (Recommended for Development)

Best for: Development, testing, or non-production environments

  1. Backup your data (optional, if you need to preserve state)
  2. Stop the node gracefully
  3. Delete old WAL directory:
    rm -rf /path/to/storage/wal/*
  4. Start with v0.2.0

⚠️ Warning: This will lose all uncommitted/unreplicated data in the WAL.


Option 2: Rolling Upgrade (Production Cluster)

Best for: Production clusters with replication (3+ nodes)

Since d-engine uses Raft consensus, you can perform a rolling upgrade:

  1. Ensure cluster is healthy (all nodes synchronized)
  2. For each node:
    • Stop the node gracefully (ensure data is persisted)
    • Upgrade to v0.2.0
    • Clear WAL directory: rm -rf /path/to/storage/wal/*
    • Start the node (it will catch up from other nodes)
  3. Repeat for all nodes one by one

The cluster will remain available during the upgrade (assuming you have 3+ nodes).


Option 3: Snapshot-based Migration

Best for: Large WAL files or single-node setups

  1. On old version (pre-v0.2.0):
    • Trigger a snapshot to persist current state
    • Wait for snapshot to complete
    • Verify snapshot file exists: /path/to/storage/snapshots/
  2. Upgrade to v0.2.0
  3. Clear WAL: rm -rf /path/to/storage/wal/*
  4. Start node - it will restore from the snapshot

Verification After Migration

After upgrading, verify:

# Check node starts without errors
journalctl -u d-engine -f

# Verify TTL entries expire correctly
# (create a key with TTL and wait for expiration)

# Check logs for WAL-related errors
grep "WAL" /var/log/d-engine.log

TTL Behavior Changes

Aspect Old (pre-v0.2.0) New (v0.2.0+)
TTL Storage Relative (seconds from now) Absolute (UNIX timestamp)
After Restart TTL resets 🔄 TTL preserved ✅
WAL Replay All entries loaded Expired entries skipped ✅
Expiration Semantics Relative TTL Absolute timestamp ✅
Crash Safe ❌ No ✅ Yes

Need Help?


Timeline

Version WAL Format Wire Protocol Migration Required
v0.1.x Relative TTL Compatible -
v0.2.0–v0.2.2 Absolute expiration Compatible ✅ Yes (clear WAL from v0.1.x)
v0.2.3 Same as v0.2.0+ Incompatible ✅ Yes (protobuf enum changes + API changes)
v0.2.4 Same as v0.2.0+ Compatible (additive) ✅ Yes (delete snapshot/ — format changed to CF export)
v0.2.5 Same as v0.2.0+ Compatible (additive) ⚠️ Minor (remove lease.enabled; remove entry_term() from custom SM; use start_node instead of start_custom_with_config)

🚨 For v0.2.2 Users: Protobuf Enum Breaking Changes in v0.2.3

What Changed

v0.2.3 introduces breaking wire protocol changes due to protobuf enum value shifts to comply with buf lint standards.

⚠️ Critical Impact

Wire Protocol Incompatibility:

  • v0.2.3 nodes CANNOT communicate with v0.2.2 or earlier nodes
  • No rolling upgrade possible - all cluster nodes must upgrade simultaneously
  • Client SDKs must be upgraded to v0.2.3 to connect to upgraded clusters

Enum Value Changes

NodeRole Enum

Role Old Value New Value New Constant
- - 0 NODE_ROLE_UNSPECIFIED
Follower 0 1 NODE_ROLE_FOLLOWER
Candidate 1 2 NODE_ROLE_CANDIDATE
Leader 2 3 NODE_ROLE_LEADER
Learner 3 4 NODE_ROLE_LEARNER

NodeStatus Enum

Status Old Value New Value New Constant
- - 0 NODE_STATUS_UNSPECIFIED
Promotable 0 1 NODE_STATUS_PROMOTABLE
ReadOnly 1 2 NODE_STATUS_READ_ONLY
Active 2 3 NODE_STATUS_ACTIVE

ErrorCode Enum

Error Old Value New Value New Constant
- - 0 ERROR_CODE_UNSPECIFIED
NotLeader 1 1 ERROR_CODE_NOT_LEADER
... (others unchanged)

Additional Protobuf Changes

  • Enum Prefixes: All enum values now have proper prefixes (NODE_ROLE_*, NODE_STATUS_*, ERROR_CODE_*)
  • Field Naming: All fields now use snake_case naming (leader_id, prev_log_index, last_log_index, etc.)

Migration Steps for v0.2.2 → v0.2.3 Protobuf Changes

Step 1: Update Configuration Files

Update any TOML configuration files that reference enum values:

Old (v0.2.2):

[cluster]
node_id = 1
role = 0        # Follower
status = 0      # Promotable

New (v0.2.3):

[cluster]
node_id = 1
role = 1        # NODE_ROLE_FOLLOWER
status = 1      # NODE_STATUS_PROMOTABLE

Step 2: Upgrade All Cluster Nodes Simultaneously

⚠️ No Rolling Upgrade Possible

Since the wire protocol is incompatible, you must upgrade all nodes at once:

  1. Schedule maintenance window (cluster will be unavailable during upgrade)
  2. Stop all nodes in the cluster
  3. Upgrade binaries to v0.2.3 on all nodes
  4. Update configuration files (see Step 1)
  5. Start all nodes simultaneously
  6. Verify cluster health (check logs, run health checks)

For Production Clusters:

If you require high availability during upgrade:

  1. Set up a parallel v0.2.3 cluster (new hardware/instances)
  2. Migrate data to the new cluster (application-level migration)
  3. Switch traffic to new cluster
  4. Decommission old cluster

Step 3: Upgrade Client SDKs

All client applications must upgrade their d-engine SDK to v0.2.3:

Cargo.toml:

[dependencies]
d-engine = { version = "0.2.3", features = ["client"] }

Rebuild and redeploy all client applications before connecting to upgraded cluster.

Step 4: Verification

After upgrade, verify:

# Check all nodes started successfully
journalctl -u d-engine -f

# Verify cluster health
curl http://localhost:8080/health

# Test basic operations
d-engine-cli put test-key test-value
d-engine-cli get test-key

🚨 For v0.2.2 Users: API Changes in v0.2.3

What Changed

v0.2.3 introduces breaking API changes to unify client interfaces and improve developer experience.

Breaking Changes

1. Unified Client API Trait

Old (v0.2.2):

use d_engine::client::KvClient;
use d_engine::client::KvError;

async fn example(client: impl KvClient) -> Result<(), KvError> {
    // ...
}

New (v0.2.3):

use d_engine::client::ClientApi;
use d_engine::client::ClientApiError;

async fn example(client: impl ClientApi) -> Result<(), ClientApiError> {
    // ...
}

Migration Steps:

  • Replace KvClient with ClientApi in trait bounds
  • Replace KvError with ClientApiError in error handling
  • Update imports: use d_engine::client::{ClientApi, ClientApiError};

2. Default Persistence Strategy

Old (v0.2.2): Default = MemFirst (write to memory, async flush to disk)

New (v0.2.3): Default = DiskFirst (Raft protocol compliance)

Migration:

If you want to restore v0.2.2 behavior, add to config:

[raft.persistence]
persistence_strategy = "MemFirst"

⚠️ Warning: MemFirst trades durability for performance. Only use in scenarios where data loss is acceptable.


Non-Breaking Changes

  • CompareAndSwap (CAS): New atomic operation added
  • Drain-based batching: Performance improvements (no API changes)
  • Client::refresh(): New method for leader rediscovery

Last Updated: February 2026


🚨 For v0.2.3 Users: StateMachine::apply_chunk Signature Change (#388)

If you implemented a custom StateMachine, update apply_chunk:

// Old (v0.2.3)
fn apply_chunk(&mut self, entries: Vec<Entry>) -> Result<(), ...>

// New (v0.2.4) — slice of ApplyEntry (decoded key/value/TTL)
fn apply_chunk(&mut self, entries: &[ApplyEntry]) -> Result<(), ...>

ApplyEntry carries decoded fields directly — no proto parsing needed in your impl.


🚨 For v0.2.3 Users: API Surface Changes in v0.2.4 (#326)

What Changed

v0.2.4 removes internal implementation details that were accidentally exposed as pub. All removed items were internal — they were never part of the documented public API.

Breaking Changes

1. EmbeddedClient::node_id() removed

This method returned client_id (not a node ID), which was semantically incorrect.

// Old (v0.2.3) — broken semantics, removed
let id = client.node_id();

// New (v0.2.4) — use EmbeddedEngine instead
let id = engine.node_id();

2. GrpcClient convenience methods now require ClientApi trait in scope

get_linearizable(), get_lease(), and get_eventual() are now only available via the ClientApi trait. The return type is unified to Option<Bytes> (was Option<ClientResult>).

// Old (v0.2.3)
let result = client.get_linearizable("key").await?;
let value = result.map(|r| r.value); // extra unwrap needed

// New (v0.2.4) — add trait import, get Bytes directly
use d_engine_client::ClientApi;
let value = client.get_linearizable("key").await?; // Option<Bytes>

3. Node::set_rpc_ready(), is_rpc_ready(), ready_notifier() removed from public API

These were internal lifecycle methods. Use EmbeddedEngine::wait_ready() instead.

// Old (v0.2.3)
node.set_rpc_ready(true);
let ready = node.is_rpc_ready();

// New (v0.2.4) — use the engine-level API
engine.wait_ready(Duration::from_secs(5)).await?;

4. Node::node_config field is no longer public

// Old (v0.2.3)
let node_id = node.node_config.cluster.node_id;

// New (v0.2.4)
let node_id = node.node_id();

5. NodeBuilder::init() is no longer public

Use the documented constructors instead.

// Old (v0.2.3)
NodeBuilder::init(config, shutdown_rx)

// New (v0.2.4)
NodeBuilder::new(None, shutdown_rx).node_config(config)
// or
NodeBuilder::from_cluster_config(cluster_config, shutdown_rx)

6. QuorumStatus removed

This type was defined but never used. Remove any references to it.

7. ClientInner and ConnectionPool no longer public

These are internal connection pool types. Use Client and ClientBuilder instead.



🚨 For v0.2.3 Users: Snapshot Format Change in v0.2.4

What Changed

v0.2.4 changes the internal snapshot format from RocksDB checkpoint (v0.2.3) to CF export (v0.2.4). Existing v0.2.3 snapshots cannot be loaded by v0.2.4 nodes.

Migration Steps

  1. Ensure the cluster is healthy — all nodes synchronized before upgrading.
  2. Delete old snapshots before starting the upgraded node:
    rm -rf /path/to/db_root_dir/snapshot/
  3. Upgrade to v0.2.4 — the node replays from WAL on first start.
  4. A new snapshot is created automatically once the log-size threshold is reached.

Ensure sufficient disk space for WAL replay. Nodes that have fallen far behind may trigger a snapshot install from the leader instead of local WAL replay.


For v0.2.3 Users: Watch Behavior Change in v0.2.4 (#294)

What Changed

Watch buffer overflow is no longer silent. When a per-watcher channel fills up, the server now:

  1. Sends a CANCELED sentinel (event_type = WATCH_EVENT_TYPE_CANCELED, error = WATCH_BUFFER_OVERFLOW)
  2. Unregisters the watcher — no further events are delivered on that stream

Previously, events were silently dropped and the watcher was silently unregistered with no client notification.

Impact

Non-breaking — clients that handle stream close (None / Status::UNAVAILABLE) continue to work. The CANCELED event simply provides early notification before the stream closes.

Recommended: update watch consumers to handle CANCELED for clean re-sync:

while let Some(event) = stream.next().await {
    match event {
        Ok(ev) if ev.event_type == WatchEventType::Canceled as i32 => {
            // buffer overflow: re-sync current state then re-register
            let current = client.read(ev.key.clone()).await?;
            process_snapshot(current);
            stream = client.watch(ev.key).await?;
        }
        Ok(ev) => handle_event(ev),
        Err(e) => return Err(e),
    }
}

See Watch Feature Guide for details.


For v0.2.4 Users: StateMachine::entry_term() Removed in v0.2.5 (#418)

What Changed

The entry_term() method has been removed from the StateMachine trait. Term lookup belongs to the Raft log layer, not the state machine — this was always an architectural layering issue.

Impact

⚠️ Breaking for custom StateMachine implementations.

If you implemented a custom StateMachine, delete the entry_term method from your impl block. Built-in implementations (FileStateMachine, RocksDBStateMachine, SledStateMachine) are already updated and require no changes.

// Old (v0.2.4) — delete this block
fn entry_term(&self, entry_id: u64) -> Option<u64> {
    Some(1)
}

Compilation will fail with "method not found in trait" if you forget to remove it.


For v0.2.4 Users: StateMachine::apply_snapshot_from_file Return Type Changed in v0.2.5 (#436)

What Changed

apply_snapshot_from_file now returns Result<SnapshotApplyResult> instead of Result<()>, so the caller knows whether the install actually ran or was a no-op instead of every outcome looking identical.

pub enum SnapshotApplyResult {
    Applied { last_included: LogId },    // install actually ran
    IgnoredStale { current: LogId },     // incoming.index < current — no-op, not an error
    IgnoredDuplicate { current: LogId }, // incoming == current, same term — no-op, not an error
}

Impact

⚠️ Breaking for custom StateMachine implementations.

// Old (v0.2.4)
async fn apply_snapshot_from_file(
    &self,
    metadata: &SnapshotMetadata,
    path: PathBuf,
) -> Result<()> {
    // ... install logic ...
    Ok(())
}

// New (v0.2.5)
async fn apply_snapshot_from_file(
    &self,
    metadata: &SnapshotMetadata,
    path: PathBuf,
) -> Result<SnapshotApplyResult> {
    // ... same install logic ...
    Ok(SnapshotApplyResult::Applied { last_included: new_last_included })
}

Return IgnoredStale for a snapshot strictly older than what's already applied, and IgnoredDuplicate only when the index and term both match what's already applied — that's Raft's own idempotency check on the install call, not a failure. An equal index with a different term is not a duplicate: it means two different leadership terms produced conflicting content at the same boundary, which the public error contract classifies as fatal — return Err(SnapshotError::BoundaryConflict { .. }) for that case. Reserve Err generally for a genuine failure (checksum mismatch, I/O error, corrupted data, or this boundary conflict).

Built-in implementations (FileStateMachine, RocksDBStateMachine) are already updated. Compilation fails with "incompatible type for trait" until a custom implementation is updated.


For v0.2.4 Users: Learner Initial-Snapshot PULL Path Removed in v0.2.5 (#436)

What Changed

The learner's eager startup snapshot pull is removed. A new learner no longer calls Transport::request_snapshot_from_leader to fetch a snapshot on startup; it relies on the leader's PUSH replication loop instead (AppendEntries, or InstallSnapshot when its log falls behind the purge boundary).

Impact

⚠️ Breaking for custom Transport and StateMachineHandler implementations.

Two trait methods are removed:

  • Transport::request_snapshot_from_leader(...) — delete this method from any custom Transport.
  • StateMachineHandler::apply_snapshot_stream_from_leader(...) — delete this method from any custom StateMachineHandler.

Compilation fails with "method ... is not a member of trait" until the custom implementation is updated. The built-in implementations are already updated.


For v0.2.4 Users: start_custom_with_config → start_node() (#415)

What Changed

The public API now exposes start_node() directly instead of routing through a thin wrapper.

// Old (v0.2.4) — removed
EmbeddedEngine::start_custom_with_config(storage, sm, config).await?;

// New (v0.2.5)
EmbeddedEngine::start_node(config, storage, sm).await?;           // embedded
StandaloneEngine::start_node(config, storage, sm, shutdown_rx).await?;  // standalone

start_node calls config.validate() internally — callers may pass validated or unvalidated configs.


For v0.2.4 Users: Snapshot Label Fix (#418)

What Changed

Snapshot last_included was previously computed by subtracting retained_log_entries from last_applied, introducing a label/data mismatch. Fixed: last_included == last_applied always.

Impact

Transparent — no action required. Nodes that installed snapshots with the old label may have double-applied entries; upgrading to v0.2.5 prevents this from recurring. Existing snapshots are compatible.


For v0.2.4 Users: TTL Config Change in v0.2.5 (#398)

What Changed

The enabled flag under [raft.state_machine.lease] has been removed. TTL expiration is now always active — no opt-in required.

In v0.2.4, omitting enabled = true caused a fatal crash on put_with_ttl. This bug is fixed in v0.2.5.

Migration

If your config contains enabled = true or enabled = false, remove the line:

# Old (v0.2.4) — remove this line
[raft.state_machine.lease]
enabled = true   # ← delete

# New (v0.2.5) — TTL always active, no flag needed
[raft.state_machine.lease]
cleanup_interval_ms = 1000

Impact: None if you do not touch the config — unrecognised fields are ignored. The only behavioral change is that TTL expiration is now unconditionally enabled.


For v0.2.4 Users: New [raft.read_actor] Config Section in v0.2.5 (#392)

This section is optional — both fields have defaults and existing configs work without changes.

[raft.read_actor]
# mpsc channel buffer for Eventual/LeaseRead fast path.
# Rule of thumb: ≥ 2× peak concurrent readers. Default: 512.
channel_capacity = 512

# Max reads drained per wakeup. Default: 100.
max_drain = 100

No migration action required unless you want to tune read concurrency.


For v0.2.4 Users: data_dir moved out of config in v0.2.5

data_dir (formerly db_root_dir) is no longer a config file field. Pass it as an explicit argument to the engine constructor instead:

// Before (v0.2.4)
let engine = EmbeddedEngine::start_with("config.toml").await?;
// config.toml had [cluster] data_dir = "./db"

// After (v0.2.5)
let engine = EmbeddedEngine::start_with("./db", "config.toml").await?;

All engine constructors (start, start_with, start_custom, start_node, run, run_with, run_custom) now require data_dir as the first argument. Remove [cluster] data_dir / [cluster] db_root_dir from all config files — the entry is silently ignored if left in place, but it's misleading.

[cluster] log_dir is removed (it was validated but never consumed).

snapshots_dir is no longer configurable — always data_dir/snapshots.


For v0.2.4 Users: NodeBuilder is no longer public in v0.2.5

Only affects code that plugged in a custom storage engine or state machine directly via NodeBuilder. If you used EmbeddedEngine or StandaloneEngine, no change needed.

// Before (v0.2.4)
let node = NodeBuilder::new(data_dir, None, shutdown_rx)
    .storage_engine(storage_engine)
    .state_machine(state_machine)
    .start()
    .await?;
node.run().await?;

// After (v0.2.5)
EmbeddedEngine::start_custom(data_dir, storage_engine, state_machine, None).await?;
// or, for a standalone gRPC server:
StandaloneEngine::run_custom(data_dir, storage_engine, state_machine, shutdown_rx, None).await?;

start_custom/run_custom cover the same ground — explicit data_dir, custom StorageEngine/StateMachine, optional config file.


Last Updated: August 2026