rand-fast is a Linux performance diagnostic tool. It measures scheduler
latency, CPU usage, disk I/O latency, network retransmissions, off-CPU wait
time, and memory pressure for a single process and its threads, using Aya and
eBPF.
- Linux 5.8+ with BTF (validated on 6.12.105 host and 7.2.0-rc6 QEMU guest; the used tracepoints have been available since 2.6)
- Kernel built with
CONFIG_BPF,CONFIG_BPF_SYSCALL,CONFIG_BPF_EVENTS,CONFIG_DEBUG_INFO_BTF - Tracepoints under
/sys/kernel/debug/tracing/events(or/sys/kernel/tracing) - Rust nightly with the
rust-srccomponent bpf-linker(0.9.x, e.g.LLVM_SYS_191_PREFIX=/usr/lib/llvm-19 cargo install bpf-linker --version 0.9.14)- Permission to load eBPF programs and open perf events:
root, orCAP_BPF+CAP_PERFMON(kernel 5.8+), orCAP_SYS_ADMINon older kernels - QEMU 7.2+ with KVM (
/dev/kvm) for verifier and privileged smoke tests without hostsudo(useskernel-serverbpf-nextbzImageand~/.bpf_selftests/root.img)
The build script compiles the eBPF object with bpfel-unknown-none and embeds it in the userspace binary.
cargo build --releaseThis produces:
target/release/fast— the diagnostic CLItarget/release/sched-workload— scheduler workload generatortarget/release/fast-workload— workload generators for io, net, lock, and memory
All subcommands take --pid <TGID> and --duration (for example 10s or
500ms). Threads are discovered through /proc/<pid>/task and re-synced
while the collection runs; the report is printed when the collection ends or
the process exits.
| Subcommand | Measures | State |
|---|---|---|
sched |
runnable → running scheduler latency | measured |
cpu |
CPU usage and on-CPU hot stacks | measured, symbolized top stacks |
io |
block I/O latency | measured, per-device latency with op split |
net |
TCP retransmissions | stub (RTT/endpoint fields zeroed) |
off-cpu |
off-CPU wait time | measured, stacks not symbolized |
memory |
PSI, page faults, swap | measured from /proc |
diagnose |
ranked likely causes | heuristic /proc signals only |
daemon |
flight recorder | stub (synthetic ring data) |
Measures the time a thread spends runnable before it starts running, using
the sched_wakeup and sched_switch tracepoints.
sudo ./target/release/fast sched --pid 1234 --duration 10sSample output (target under 8 hog workers contending):
PID: sched-workload (18342)
Duration: 5s
Samples: 4871
Lost events: 0
Scheduler latency
p50 7 µs
p95 2.3 ms
p99 4.0 ms
max 5.2 ms
CPU latency (running CPU)
cpu 0 samples 610 p50 6 µs p95 1.9 ms p99 3.7 ms max 4.8 ms
Slow events
> 1ms 297
> 10ms 0
> 50ms 0
Verification: start ./target/release/sched-workload target --duration 30s --period 1ms, run fast sched against its PID (idle p95 in the low µs
range), then repeat while ./target/release/sched-workload hog --duration 30s --workers 8 runs (p95 and slow-event counts rise clearly).
Samples on-CPU stacks of the target threads at --frequency (default 99 Hz)
by attaching a BPF program to a perf cpu-clock event on every online CPU,
and reports process CPU usage from /proc tick deltas. Top stacks are
symbolized: kernel frames via the running kernel's kallsyms (full addresses
require root), user frames via blazesym against the target process' live
/proc/<pid> state. The sample count scales with sampling rate × CPU time:
an idle target yields ~0 samples, a busy target ≈ frequency × duration ×
busy-core count, regardless of how many times the threads wake up.
sudo ./target/release/fast cpu --pid 1234 --duration 10s --frequency 199Sample output (lock-hog fixture, 16 workers):
PID: fast-workload (326)
Duration: 5s
Samples: 3489
Lost: 0
CPU usage: 97.7% (over 5.0s wall, 4019 ticks total)
On-CPU samples (hot stacks)
stack (952, 144) samples 302 (8.7%)
0 [k] native_queued_spin_lock_slowpath+0x15
1 [k] _raw_spin_lock+0x29
2 [k] futex_wake+0x172
3 [k] do_futex+0xd8
4 [k] __x64_sys_futex+0x136
5 [k] do_syscall_64+0xaa
6 [k] entry_SYSCALL_64_after_hwframe+0x76
7 syscall+0x1d (libc.so.6)
stack (-14, 2560) samples 196 (5.6%)
0 std::sys::backtrace::__rust_begin_short_backtrace::<...>+0x90 (fast-workload)
stack (3252, 144) samples 190 (5.4%)
0 [k] native_queued_spin_lock_slowpath+0x13e
...
Per-CPU samples
cpu 0 samples 492
Correlation
CPU saturation likely contributes to scheduler latency (CPU 97.7% with 3489 samples)
Kernel frames carry the [k] prefix and are resolved through kallsyms;
user frames are resolved against the target's live maps (libc frames stay
raw on systems without libc symbol tables). Kernel stacks are stored only
for samples that interrupt the target inside the kernel; user-context
samples report user frames alone.
Verification: fast cpu against ./target/release/sched-workload target --duration 30s collects 0 samples and near-zero usage; against
./target/release/sched-workload hog --duration 30s --workers 8 the sample
count approaches frequency × duration × busy cores (3943 samples at 99 Hz
× 8 cores × 5 s) and the futex wait path shows up symbolized under
fast-workload lock-hog. Rate scaling: 3928 samples at 99 Hz vs 15566 at
396 Hz against the same hog (3.96x, expected ~4x).
Measures block I/O latency from block_rq_issue / block_rq_complete, and
reads /proc/<pid>/io counters for byte deltas. The tracepoint payloads are
parsed (offsets verified against the guest format files the init script
dumps): each event carries the device (major:minor), start sector, size in
sectors, and the operation (read/write) read from the rwbs field.
Requests are paired per request, not per issuing thread: the pending map is
keyed by (device, start sector), so completions that run in IRQ/softirq
context still attribute to the issuing thread and completion rate tracks
issue rate.
--threshold (default 10ms) controls the slow-I/O counter.
sudo ./target/release/fast io --pid 1234 --duration 10s --threshold 10msSample output (io-hog, sequential O_DIRECT reads on the QEMU scratch disk):
PID: fast-workload (204)
Duration: 5s
Samples: 151281
Lost events: 0
Slow > 10ms: 1
rchar: 634273792 bytes, wchar: 0 bytes
I/O latency
samples 151281
p50 47 µs
p95 65 µs
p99 87 µs
max 536.3 ms
ops: read 151281 (590.9 MiB)
Per-device latency
dev vda (254:0) samples 151281 sectors 1210248
ops: read 151281 (590.9 MiB)
p50 47 µs p95 65 µs p99 87 µs max 536.3 ms
Slow I/O > 10ms (top 1 of 1)
latency device op sectors bytes tid
536.3 ms vda (254:0) read 8 4.0 KiB 207
The device label resolves through /sys/dev/block (name plus
major:minor); without a sysfs entry it falls back to major:minor.
Verification: fast io against sleep 60 collects no samples; against
./target/release/fast-workload io-hog --duration 30s --workers 2 --path <file-on-real-block-fs> the report names the device holding the file, the
rchar delta and per-device byte totals grow together, and the slow-I/O
table lists the slowest requests above --threshold. Per-request pairing:
284120 completions matched 290619 reads in the smoke run (97.8%; the old
per-TID matching paired only a fraction of a percent because completions
run in IRQ context). The default path lives under /tmp, which is often
tmpfs — reads there never reach the block_rq_* tracepoints, so only the
rchar delta moves; point --path at a file on a real block filesystem
(or the QEMU scratch disk) to collect latency samples.
Counts TCP retransmissions of target threads from
tcp_retransmit_skb. RTT and endpoint fields are stubs (see limitations).
sudo ./target/release/fast net --pid 1234 --duration 10sSample output (idle process on a quiet network):
PID: sleep (18600)
Duration: 5s
Retransmissions: 0
Lost events: 0
Endpoints
No TCP samples collected (no retransmissions or RTT >0). Test with: python3 -m http.server 8000 & curl http://127.0.0.1:8000/
Local fixture: start a local server and generate traffic from the target process.
Note: connection vs transfer vs retransmission delays are distinguished by RTT (transfer) and retrans flag.
Verification: fast net against ./target/release/fast-workload net-hog --duration 30s --workers 4 exercises loopback request/response traffic.
Loopback rarely retransmits, so the counter usually stays 0; to produce real
retransmissions, add artificial loss (requires root): tc qdisc add dev lo root netem loss 1% and clean up with tc qdisc del dev lo root.
Measures off-CPU wait by pairing sched_switch (a target thread switches
out in a sleepable state) with sched_wakeup (it becomes runnable again),
with kernel stack ids for hot wait stacks.
sudo ./target/release/fast off-cpu --pid 1234 --duration 10sSample output (lock-hog, 16 contending threads on 8 CPUs):
PID: fast-workload (18700)
Duration: 5s
Samples: 557348
Lost: 0
Off-CPU wait
p50 4 µs
p95 43 µs
p99 200 µs
max 2.7 ms
Hot wait stacks
stack 210 samples 294700
Verification: fast off-cpu against ./target/release/sched-workload target --duration 30s collects ~1000 timer-sleep waits per second (p50 ≈ the sleep
period); against ./target/release/fast-workload lock-hog --duration 30s --workers 16 the sample count explodes into hundreds of thousands of short
futex waits. Use more workers than CPUs so waiters actually park instead of
spinning.
Polls PSI (/proc/pressure/memory), per-process page faults
(/proc/<pid>/stat), and swap usage (/proc/meminfo) every 200 ms.
sudo ./target/release/fast memory --pid 1234 --duration 10sSample output (mem-hog, 64 MiB rounds):
PID: fast-workload (18800)
Duration: 5s
Memory pressure
PSI some avg10: 0.0% max 0.2%
PSI full avg10: 0.0%
Page faults: minflt 12264 (2452.8/s) majflt 0 (0.0/s)
Swap used: 0 KB
Note: system vs process distinguished; PSI is system-level, faults are per-process.
Verification: fast memory against sleep 60 shows a near-zero fault rate;
against ./target/release/fast-workload mem-hog --duration 30s the minor
fault rate climbs into the thousands per second. Requires CONFIG_PSI for
the PSI lines.
Runs a heuristic ranking of likely causes from /proc signals (loadavg, I/O
counters, TCP retransmission indicator, voluntary context switches, PSI). It
does not run the eBPF collectors.
sudo ./target/release/fast diagnose --pid 1234 --duration 10sSample output:
Diagnosing PID: sched-workload (18900) for 10s
Ranked causes (deterministic):
1. CPU contention 100.0% evidence: loadavg 8.42, estimated CPU pressure 100%
2. Scheduler latency 20.0% evidence: scheduler p95 estimated from run-queue (requires eBPF for precise)
3. Disk I/O 5.0% evidence: rchar 12345 bytes
4. Network 5.0% evidence: TCP retrans indicator 0
5. Lock contention 4.5% evidence: voluntary_ctxt_switches 1500
6. Memory pressure 0.9% evidence: PSI memory avg10 0.3%
Evidence preserved per signal; confidence is tested and deterministic.
For competing bottlenecks, synthetic fixtures (hog, dd, iperf, futex, stress --vm) validate ranking.
Note: full diagnose would run 'fast sched/cpu/io/net/off-cpu/memory' collectors in parallel for 10s
Verification: start ./target/release/sched-workload hog --duration 30s --workers 8 and run diagnose against it; CPU contention ranks first. (The
exact confidence values depend on machine load; the ranking is
deterministic for the same signals.)
Prototype flight recorder: keeps a rolling 60 s ring of scheduler/CPU/IO summary values and writes an incident JSON file when the ring detects the configured scheduler p95 trigger. Collection data is currently synthetic and the trigger path is unreachable yet (see limitations); today an incident is preserved when the observed process exits.
sudo ./target/release/fast daemon --pid 1234 --duration 30s --trigger 10ms --output ./incidentsSample output (observed process exits mid-run, preserving the window):
Flight recorder for sched-workload (18900)
Budget: CPU <2%, mem <10MB, rolling 60s
Trigger: scheduler p95 > 10ms
Output: ./incidents
Process exited, preserving final window
Incident preserved to ./incidents/incident-18900.json
Flight recorder stopped after 6s
Incidents preserved: 0
Resource budget: CPU <2%, mem <10MB, ring 600 entries, 10485760 bytes budget
Restart behavior: ring persists to ./incidents and reloads on start
Verification: start ./target/release/sched-workload hog --duration 5s and
run the daemon against it with a longer duration; when the hog exits, the
daemon preserves the current window as an incident file under --output and
stops. Automatic trigger-fired preservation is not reachable yet (ring data
is synthetic, see limitations).
Honest state of each area, as of this version:
cpu: kernel stacks are stored only for samples that interrupt the target inside the kernel (user-context samples carry user frames alone); libc frames stay raw on systems without libc symbol tables. The usage percentage is derived from/proctick deltas, not from the eBPF samples.io: requests are keyed by (device, start sector). Two outstanding requests on the exact same key keep only the first issue (BPF_NOEXIST), empty flush requests all share sector 0, and a completion from a non-target process for the same key would consume the target's pending entry and be misattributed to it — the request pointer that would remove these races is not reachable from tracepoint programs. Payload offsets are verified against the 7.2.x tracepoint format files; older kernel series (e.g. 5.x, where the fields sit at different offsets) would need re-verification.net: onlytcp_retransmit_skbis tracked; RTT and address/port fields are zeroed, so no endpoint or RTT table exists yet.off-cpu: waits pair switch-out with the next wakeup, so the final wake-to-run dispatch is counted as scheduler latency instead; stack ids capture the waking context, not the sleeping frame, wait reasons are not classified, and stack ids are not symbolized.memory: pure/procpolling (PSI, stat, meminfo). TheMEMORY_EVENTSeBPF map is emitted but not consumed by userspace.diagnose: heuristic ranking from/proconly; no eBPF collection, and confidences are rough estimates, not measured percentages of the slowdown.daemon: ring entries are synthetic (no real collection yet), incident files contain the ring window only, and the printed "restart behavior" is not implemented.
The privileged smoke matrix runs every eBPF-backed subcommand idle vs under its matching load fixture and records verifier results and deltas:
# Inside the QEMU guest (root), with the release binaries in the working dir:
tools/qemu-smoke.sh # writes raw reports to /tmp/fast-smokeThe script exits non-zero if any verifier load or collection fails.
It also runs on a host where you hold the required capabilities.
tools/qemu-guest-init.sh boots a minimal initramfs-only guest (busybox,
the release binaries, a loopback link, and an optional ext4 scratch disk for
the io-hog fixture) and runs the matrix automatically:
truncate -s 256M /tmp/fast-io.img && mkfs.ext4 -q -F /tmp/fast-io.img
gcc -static -O2 -o /tmp/mount2 tools/qemu-guest-mount.c
# build the initramfs around tools/qemu-guest-init.sh, then:
qemu-system-x86_64 -enable-kvm -m 2048 -smp 8 \
-kernel vmlinuz -initrd initramfs.img.gz \
-append "console=ttyS0 rdinit=/init loglevel=3 panic=-1" \
-nographic -no-reboot -monitor none \
-drive file=/tmp/fast-io.img,format=raw,if=virtioRecorded results (QEMU KVM guest, 8 vCPUs, kernel 7.2.3-arch1-3 with BTF,
2026-09-08, v1.0 code; the sched 7.2.0-rc6 row keeps the earlier record):
| Command | Verifier | Idle | Load |
|---|---|---|---|
sched |
pass | p50 6µs p95 19µs p99 29µs max 131µs, slow>1ms 0 | p50 3µs p95 5µs p99 1.1ms max 2.7ms, slow>1ms 46 |
sched (7.2.0-rc6, 2026-08-29) |
pass | p50 7µs p95 9µs p99 11µs max 69µs, slow>1ms 0 | p50 3µs p95 2.3ms p99 4.0ms max 5.2ms, slow>1ms 297 |
cpu |
pass | 0 samples, usage 0.1% | 3943 samples (99 Hz × 8 busy cores × 5 s), usage 99.0%, symbolized top stacks |
io |
pass | 0 samples | 151281 samples on vda (254:0), p50 47µs p99 87µs, 590.9 MiB read |
net |
pass | retransmissions 0 | retransmissions 0 (loopback does not retransmit) |
off-cpu |
pass | 4621 samples | 529503 samples |
v1.0 accuracy checks from the same run (tools/qemu-smoke.sh prints them):
- cpu rate scaling: 3928 samples at 99 Hz → 15566 samples at 396 Hz against the same hog (3.96x, expected ~4x) — sample count tracks frequency × CPU time, not wakeups.
- cpu symbolization: the lock-hog futex wait path appears symbolized
(
[k] futex_wait/do_futexkernel frames plus user frames). - io per-request pairing: 284120 completions for 290619 reads (97.8%); the pre-v1.0 per-TID matching paired roughly 0.03% (67 of ~214k).
All five eBPF-backed programs load and attach in the guest, and the load runs show the expected signal deltas.
cargo test --workspace
cargo clippy --workspace --all-targets -- -D warningsBoth commands also build the eBPF object. CI must keep clippy clean with
-D warnings.
Validated in QEMU (qemu-system-x86_64 -enable-kvm -smp 8 -kernel build/bpf-next/arch/x86/boot/bzImage -drive file=~/.bpf_selftests/root.img) with the guest running as root:
# Inside the guest (e.g. via vmtest rootfs at /root/bpf):
./sched-workload-bin target --duration 20s --period 1ms &
./fast-bin sched --pid <TARGET_PID> --duration 5s # idle
./sched-workload-bin hog --duration 10s --workers 8 & # contention
./fast-bin sched --pid <TARGET_PID> --duration 5sResult (2026-08-29, 7.2.0-rc6):
- eBPF verifier: pass (
sched_wakeup/sched_switchprograms load) - Idle: p50 7µs p95 9µs p99 11µs max 69µs, slow>1ms 0
- Contention (16 hog workers): p50 3µs p95 2.3ms p99 4.0ms max 5.2ms, slow>1ms 297
Contention shows a clear latency increase, satisfying the v0.1 completion condition.
A repeatable host-side helper is still available for manual runs:
# In one shell, start a periodic target.
./target/release/sched-workload target --duration 30s --period 1ms
# In another shell, use the printed target PID while the target is running.
sudo ./target/release/fast sched --pid <TARGET_PID> --duration 10s
# For a contention run, start a second helper in parallel.
./target/release/sched-workload hog --duration 30s --workers 8The current commands measure runnable-to-running scheduler latency, on-CPU activity, block I/O, TCP retransmissions, off-CPU waits, and memory pressure for one process at a time. CPU stack symbolization and per-request I/O attribution landed in v1.0; connection-level RTT, automatic diagnosis from real collectors, and long-running recording are planned for later versions.