Near-data BPE tokenization over NVMe-over-Fabrics
Move tokenization to the storage node—not the dataset to the compute node.
NDT-BPE is a research prototype that offloads byte-pair encoding (BPE) tokenization to an NVMe-oF storage server. The compute node sends metadata-only custom NVMe commands; the storage target reads input blocks, invokes the tokenizer runtime through shared memory, and writes token IDs back to the SSD locally.
The current data path uses Data Length=0 for opcode 0xD4. In a 509 MB single-shard validation, measured NVMe/TCP traffic fell from 2.03 GB in the previous implementation to an average of 0.84 MB across three runs. This result demonstrates payload elimination during tokenization; it does not by itself establish an end-to-end speedup.
- Metadata-only compute-to-storage commands over NVMe/TCP
- Storage-local input reads and output writes through SPDK bdev
- Multi-slot tokenizer runtime with double-buffered shared memory and
eventfd - Arrow IPC text-buffer indexing and reusable NDT-native staging
- Python bindings backed by
io_uringNVMe passthrough - One-command patching and builds for all three components
flowchart LR
subgraph C[Compute node]
APP[Python / scheduler]
BIND[ndt_compute]
URING[io_uring]
APP --> BIND --> URING
end
subgraph S[Storage node]
NVMF[Patched SPDK NVMe-oF target]
SHM[Shared-memory slots]
BPE[bpe_process runtime]
SSD[(NVMe SSD)]
NVMF -->|local read| SSD
NVMF <--> SHM <--> BPE
NVMF -->|local write| SSD
end
URING -->|opcode 0xD4: LBA, NLB, slot, output LBA| NVMF
NVMF -->|completion only| URING
NDT-BPE consists of three independently built pieces:
| Component | Location | Responsibility |
|---|---|---|
| Compute client | compute/ |
Arrow metadata, FIEMAP/LBA scheduling, io_uring, Python API |
| SPDK target | storage/spdk/ + scripts/patches/ |
Custom command handling and storage-local bdev I/O |
| BPE runtime | storage/runtime/ |
Shared-memory request processing and tokenization |
.
├── compute/ # Compute-side C++ library and Python binding
├── storage/
│ ├── runtime/ # bpe_process, worker runtime, monitor
│ ├── spdk/ # Pinned upstream SPDK submodule
│ └── setting.sh # Compute-side NVMe-oF connect/mount helper
├── scripts/
│ ├── patch.sh # Patch and build all three components
│ └── patches/
│ └── spdk-ndt-bpe.patch # Reproducible SPDK modifications
├── .gitmodules
└── LICENSE
NDT-BPE currently targets Linux on x86-64. A two-node setup is recommended.
- Ubuntu 22.04 or a comparable Linux distribution
- A compute node and a storage node connected over TCP
- NVMe device dedicated to the SPDK target
- Python 3.12 with
venv - GCC/G++ with C++17 support
- CMake, Make, pkg-config, and Git
- Kernel NVMe/TCP and
io_uringsupport - Huge pages and SPDK build dependencies
- Root privileges for device binding, NVMe-oF target startup, and mounting
Warning
SPDK takes exclusive ownership of the configured PCI device. Verify PCI and BDEV carefully. Never bind a system or mounted data device to SPDK.
git clone --recurse-submodules git@github.com:doogunwo/NDT-BPE.git
cd NDT-BPEOn the build machine:
sudo ./storage/spdk/scripts/pkgdep.shInstall python3-venv, nvme-cli, and tmux if they are not already available.
./scripts/patch.shThis command:
- initializes the pinned SPDK, liburing, and tokenizers-cpp submodules;
- applies
scripts/patches/spdk-ndt-bpe.patchidempotently; - creates
compute/venvand builds the Python binding; - builds
bpe_process,bpe_process_main2, andbpe_monitor; - builds the patched
nvmf_tgt.
Useful variants:
./scripts/patch.sh --apply-only
./scripts/patch.sh --skip-spdk-build
JOBS=8 ./scripts/patch.shBuild outputs:
compute/venv/ Python environment
storage/runtime/bin/bpe_process_main2 Storage runtime
storage/runtime/bin/bpe_monitor Runtime monitor
storage/spdk/build/bin/nvmf_tgt Patched NVMe-oF target
Create .env in the repository root. The file is ignored by Git.
TARGET_IP=192.168.1.160
STORAGE_IP=192.168.1.160
TARGET_PORT=4420
NQN=nqn.2025-01.io.spdk:cnode1
MNT=/mnt/nvme
DEV=/dev/nvme2n1
PCI=0000:03:00.0
BDEV=NVMe0Adapt every value to your environment. Device node numbers such as /dev/nvme2n1 may change after reconnecting.
cd storage/runtime
sudo ./bin/bpe_process_main2 \
--mode=arrow \
--exec-mode=process \
--workers=16The runtime creates per-slot shared-memory buffers and waits for the SPDK target's requests. For a background tmux session, use:
./storage/runtime/start_bpe_runtime.sh ndt-runtime arrow process 16Start the target:
sudo ./storage/spdk/build/bin/nvmf_tgt -m 0x1In another terminal, bind the intended NVMe device and configure the TCP transport, subsystem, namespace, and listener with storage/spdk/scripts/rpc.py. Use the values from .env; SPDK's standard NVMe-oF target documentation describes the corresponding RPC commands.
sudo modprobe nvme_tcp nvme_fabrics
./storage/setting.shstorage/setting.sh discovers the target, connects the NQN, detects the SPDK namespace, and mounts it at MNT.
source compute/venv/bin/activate
python - <<'PY'
import ndt_compute as ndt
result = ndt.tokenize_to_nvme(
dev_path="/dev/ng2n1",
input_path="/mnt/nvme/data-00000-of-00080.arrow",
output_path="/mnt/nvme/data-00000-of-00080.bin",
slots=16,
queue_depth=64,
)
print(result)
PYUse the generic character device (/dev/ngXnY) associated with the connected SPDK namespace. The block device (/dev/nvmeXnY) is used for mounting; the generic device is used for NVMe passthrough commands.
Arrow input is converted once into a reusable NDT-native staged representation. Dataset ingestion and repeated tokenization are intentionally treated as separate phases:
- Ingestion: parse Arrow metadata and create aligned, framed storage chunks.
- Tokenization: send only LBA/length/slot/output metadata and completion traffic.
Do not include first-time ingestion or output-pool creation in a steady-state tokenization measurement unless the experiment explicitly evaluates end-to-end ingestion.
Single Arrow shard, 509,476,864 processed bytes, 3,887 commands, one runtime slot, three repetitions:
| Metric | Previous path | Metadata-only path |
|---|---|---|
| Mean NVMe/TCP traffic | 2,034,442,416 B | 841,040 B |
| Traffic reduction | — | 99.9587% |
| Mean elapsed time | 157.753 s | 139.159 s |
| Output validation | Non-zero tokens | Non-zero tokens |
The remaining traffic is command, completion, and NVMe/TCP protocol overhead. A one-slot Ray baseline completed in 74.657 seconds in the same pilot, so performance claims require the multi-shard, multi-slot experiments rather than this transfer validation alone.
The SPDK patch does not apply
The patch targets the SPDK commit pinned by this repository. Check for local submodule changes:
git -C storage/spdk status --short
git -C storage/spdk rev-parse HEADRun ./scripts/patch.sh --apply-only after restoring the expected clean submodule checkout.
The generic NVMe device changed
List the namespace and its generic device after every reconnect:
sudo nvme list
ls -l /dev/ng*n* /dev/nvme*n*The runtime cannot open shared memory or eventfds
Start the BPE runtime before submitting custom commands and verify that all workers are alive. Stale IPC objects can be removed with:
make -C storage/runtime clean_shmNDT-BPE is an active research prototype, not a production storage system. Interfaces, command layouts, staging formats, and deployment scripts may change. Current work focuses on multi-shard scaling, data-movement attribution, scheduling, and end-to-end performance evaluation.
Licensed under the Apache License 2.0. The SPDK submodule and third-party dependencies retain their respective licenses.