Skip to content

Add workload, a generator for insert/delete benchmark streams - #114

Merged
TomTonic merged 3 commits into
mainfrom
feat/110-workload-generator
Sep 27, 2026
Merged

TomTonic merged 3 commits into
mainfrom
feat/110-workload-generator

Conversation

@TomTonic

@TomTonic TomTonic commented Sep 27, 2026 •

Copy link
Copy Markdown
Owner

Closes #110.

A new subpackage, github.com/TomTonic/rtcompare/workload, that generates mutation streams for benchmarks of mutable data structures. It is independent of #112 and #113: it branches off main and merges cleanly onto that stack (checked with git merge-tree).

API

type Op struct { ID uint32; Kind Kind }          // 8 bytes; Kind = Insert | Delete
type Config struct { Seed uint64; Ratio float64; MaxBurst, LiveTarget int; Victims Policy }

func Build(target int, c Config) ([]Op, error)   // empty → exactly 0..target-1, Ratio·target insertions
func Cycle(target int, c Config) ([]Op, error)   // 0..target-1 → 0..target-1, (Ratio−1)·target transients

type Cursor struct{ /* position */ }
func (c *Cursor) Batch(ops []Op, apply func([]Op)) rtcompare.Batch
func (c *Cursor) Advance(ops []Op, n uint64, apply func([]Op))
func (c *Cursor) Settle(ops []Op, apply func([]Op))

func Check(ops []Op, start, end []uint32) error

Compared with the sketch in the issue:

  • Errors instead of panics. Build and Cycle return error for invalid configurations, per AGENTS.md.
  • Kind instead of Delete bool. It leaves room for the lookup mix listed under extensions and still takes 8 bytes.
  • MaxBurst int instead of Burst func. Simpler, and it covers the "uniform 1..16" case from the issue. A function can be added later if needed.
  • Advance is exported as well, for callers who drive the replay themselves.
  • apply receives contiguous chunks. The loop sits in the caller's code, where the compiler can inline the structure's methods. Chunks are split at the end of the cycle.

How the simulation works

  • Insertion bursts of 1..MaxBurst alternate with deletion bursts.
  • Insertions choose between permanent and transient IDs in proportion to how many of each remain, so they spread evenly over the stream. Transient IDs are fresh (target, target+1, …).
  • The expected length of a deletion burst is the expected number of transient insertions per burst × live/LiveTarget. The equilibrium therefore sits at LiveTarget (default target/8, at most half the transients).
  • Burst lengths are rounded stochastically. With plain rounding, streams with few transients (Ratio 1.01) never deleted until the very end; the test caught that.
  • Victims: Uniform (swap-remove), FIFO (queue with compaction) and LIFO.
  • Seeds go through splitmix64, so seed 0 is deterministic. Without that, NewDPRNG(0) would pick a random state.

Tests

  • Table-driven: targets {1, 10, 1000, 20000} × Ratio {1, 1.01, 1.5, 2, 3} × {Uniform, FIFO, LIFO}. Each stream must pass Check, have the exact insert/delete counts and interleave.
  • Behaviour:
    • determinism, including seed 0;
    • each policy deletes what its name promises;
    • the live count settles at LiveTarget;
    • config errors, and Check errors on hand-built invalid streams;
    • the cursor wraps for arbitrary n (including 0 and 2·len+5), and Settle restores the start state.
  • Fuzz test of generation and replay: 60 s, 14 million executions, no finding.
  • End to end: Compare of map vs Set3 on a cycle, followed by Settle and a size check.
  • Checks: go test ./..., -race and golangci-lint are clean, and coverage is 99.4 %.

Example and one finding

ExampleCycle compares map[uint64]struct{} with Set3 (already in go.mod, so no new dependency) on a cycle over 100,000 elements. Along the way I measured the cost per operation across the phases of a cycle:

  • First pass: 35 to 128 ns per op, because the map grows to the peak size of the cycle.
  • Every later pass: 17 to 20 ns, the same in every phase.

The docs (Cursor, HOWTO) and the example therefore recommend one untimed cycle before measuring.

One more observation: on main, the example's measured series starts at 68/56 ns and settles at 17/12 ns only after about 60 batches. With #112's warm-up it starts straight in the steady state. So that is the #111 effect again, not the workload.

Docs

  • HOWTO: new section "Benchmarking insertions and deletions", covering why add-then-remove misleads, the cursor, settling and priming.
  • README: a Features bullet and an API block.

Not included (listed in the issue as possible extensions)

Lookup mix, Zipf skew, Group for multimaps, the AgeBiased policy and a compact encoding.

Addendum: one call for two questions (7984d77)

Correct use previously required running the first cycle untimed, pairing the cursor with its structure and calling Settle, and building both structures in a balanced order. Leaving out the first pass must not mean ignoring growth altogether, though, or structures with expensive growth would look better than they are. Hence:

  • Compare(target, a, b Structure, Options) (Result, error): each Structure[S] only needs Name, New() S and Apply(S, []Op). It returns two reports that answer two different questions:
    • SteadyState: cost per insert/delete once the structure is in use. It runs on a cycle; the first pass is automatically untimed.
    • Build: cost of one complete build of a fresh structure. Creation, capacity hint, every resize and the garbage all count. Whole builds, because with segments the median would ignore exactly the rare expensive resizes.
    • The build comparison has its own defaults (BuildRepeats 31, BuildValidationRuns 10, about 1,300 builds) and always sets GCBetween. With rtcompare's defaults it would take several minutes at 100,000 elements.
    • The two steady-state structures are built alternately in chunks of 256 ops, so neither is "built last".
    • SkipBuild exists as an escape hatch, and is documented as distorting the result.
  • Replay: bundles the cycle, the apply function and the cursor per structure. The candidate primes automatically in Setup before the first batch, and Settle is a method.
  • Check sizes its model for the peak state.

ExampleCompare (map vs Set3, 100,000 elements) on all three branches merged together:

steady state, per insertion or deletion …:  -39.71% [-40.99%, -37.69%]  resolved
build from empty to 100000 elements, per whole build: -40.46% [-51.43%, -33.52%]  resolved

On main alone, the steady-state part still shows the settling at the start of Collect (drift −50 %). With #112's warm-up it is gone, as shown above.

Tests: go test ./..., -race and fuzzing are green, lint reports 0 issues, and coverage is 98.2 %.

🤖 Generated with Claude Code

TomTonic and others added 3 commits September 27, 2026 19:04
Benchmarking a mutable data structure needs a stream of insertions and
deletions. The obvious "insert, then remove the same element" loop
measures a structure that never changes shape. The honest alternatives
trip over invalid operations, over victim selection inside the timed
loop, and over state that drifts across batches (#110).

- Build(target, Config) goes from empty to exactly the IDs 0 to
  target-1, inserting Ratio*target elements on the way. The transient
  ones are deleted again.
- Cycle(target, Config) starts and ends with 0 to target-1 present. In
  between it inserts and deletes (Ratio-1)*target transient elements in
  bursts, and keeps about LiveTarget of them present.
- Both streams are precomputed by simulation, so they are valid by
  construction, and they are reproducible from a seed. Victims can be
  Uniform, FIFO or LIFO. An Op is 8 bytes.
- Cursor keeps the position in a cycle with the data structure
  instance. Batch returns an rtcompare.Batch that continues where the
  last one stopped and wraps around. Settle restores the start state.
- Check replays a stream against a model and reports the first invalid
  operation.

Tests are table-driven over sizes, ratios and policies, plus a fuzz
test of generation and replay (14M executions without a finding) and
an end-to-end Compare of map against Set3. Coverage is 99.4%. The
example, which is not run by go test, compares map with Set3 on a
100,000-element cycle.

Measuring the phases of that cycle showed that the first pass costs up
to 128 ns per op against 17 to 20 afterwards, because the map grows to
its peak size. The docs therefore recommend one untimed cycle before
measuring.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Using the workload package correctly meant remembering several things.
The first pass of a cycle had to be played untimed. A cursor had to be
kept with its structure and settled afterwards. The two structures had
to be built so that neither came last. And leaving out the first pass
must not mean dropping growth altogether, or structures whose growth
is expensive would look better than they are.

- Compare(target, a, b, Options) takes two Structures, each only New
  and Apply, and answers two questions separately. SteadyState is the
  cost of one insertion or deletion, measured on a cycle after one
  untimed pass. Build is the cost of one whole build of a fresh
  structure, so creation, capacity hints, resizing and garbage all
  count. The build comparison has its own defaults (31 repeats, 10
  validation runs, about 1,300 builds) and always sets GCBetween.
- Both steady-state structures are built in alternating chunks of 256
  operations, so neither is built last.
- Replay bundles a cycle, the apply function and the cursor for one
  structure. Its candidate plays one cycle untimed in Setup before the
  first batch. Settle restores the start state.
- Check sizes its model for the largest state the stream reaches.

The example now compares map with Set3 through Compare. At 100,000
elements the build comparison resolves Set3 as 43% faster to build. On
main, the steady-state comparison still shows the start-of-Collect
settling that #112's warm-up removes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@TomTonic
TomTonic merged commit 4cc3ab2 into main Sep 27, 2026
5 checks passed
@TomTonic
TomTonic deleted the feat/110-workload-generator branch September 27, 2026 19:20
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Feature: generator for realistic insert/delete workloads (valid by construction, replayable, steady-state cycles)

1 participant