Add workload, a generator for insert/delete benchmark streams - #114
Merged
Merged
Conversation
Benchmarking a mutable data structure needs a stream of insertions and deletions. The obvious "insert, then remove the same element" loop measures a structure that never changes shape. The honest alternatives trip over invalid operations, over victim selection inside the timed loop, and over state that drifts across batches (#110). - Build(target, Config) goes from empty to exactly the IDs 0 to target-1, inserting Ratio*target elements on the way. The transient ones are deleted again. - Cycle(target, Config) starts and ends with 0 to target-1 present. In between it inserts and deletes (Ratio-1)*target transient elements in bursts, and keeps about LiveTarget of them present. - Both streams are precomputed by simulation, so they are valid by construction, and they are reproducible from a seed. Victims can be Uniform, FIFO or LIFO. An Op is 8 bytes. - Cursor keeps the position in a cycle with the data structure instance. Batch returns an rtcompare.Batch that continues where the last one stopped and wraps around. Settle restores the start state. - Check replays a stream against a model and reports the first invalid operation. Tests are table-driven over sizes, ratios and policies, plus a fuzz test of generation and replay (14M executions without a finding) and an end-to-end Compare of map against Set3. Coverage is 99.4%. The example, which is not run by go test, compares map with Set3 on a 100,000-element cycle. Measuring the phases of that cycle showed that the first pass costs up to 128 ns per op against 17 to 20 afterwards, because the map grows to its peak size. The docs therefore recommend one untimed cycle before measuring. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Using the workload package correctly meant remembering several things. The first pass of a cycle had to be played untimed. A cursor had to be kept with its structure and settled afterwards. The two structures had to be built so that neither came last. And leaving out the first pass must not mean dropping growth altogether, or structures whose growth is expensive would look better than they are. - Compare(target, a, b, Options) takes two Structures, each only New and Apply, and answers two questions separately. SteadyState is the cost of one insertion or deletion, measured on a cycle after one untimed pass. Build is the cost of one whole build of a fresh structure, so creation, capacity hints, resizing and garbage all count. The build comparison has its own defaults (31 repeats, 10 validation runs, about 1,300 builds) and always sets GCBetween. - Both steady-state structures are built in alternating chunks of 256 operations, so neither is built last. - Replay bundles a cycle, the apply function and the cursor for one structure. Its candidate plays one cycle untimed in Setup before the first batch. Settle restores the start state. - Check sizes its model for the largest state the stream reaches. The example now compares map with Set3 through Compare. At 100,000 elements the build comparison resolves Set3 as 43% faster to build. On main, the steady-state comparison still shows the start-of-Collect settling that #112's warm-up removes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #110.
A new subpackage,
github.com/TomTonic/rtcompare/workload, that generates mutation streams for benchmarks of mutable data structures. It is independent of #112 and #113: it branches offmainand merges cleanly onto that stack (checked withgit merge-tree).API
Compared with the sketch in the issue:
BuildandCyclereturnerrorfor invalid configurations, per AGENTS.md.Kindinstead ofDelete bool. It leaves room for the lookup mix listed under extensions and still takes 8 bytes.MaxBurst intinstead ofBurst func. Simpler, and it covers the "uniform 1..16" case from the issue. A function can be added later if needed.Advanceis exported as well, for callers who drive the replay themselves.applyreceives contiguous chunks. The loop sits in the caller's code, where the compiler can inline the structure's methods. Chunks are split at the end of the cycle.How the simulation works
target, target+1, …).live/LiveTarget. The equilibrium therefore sits atLiveTarget(defaulttarget/8, at most half the transients).Uniform(swap-remove),FIFO(queue with compaction) andLIFO.NewDPRNG(0)would pick a random state.Tests
Check, have the exact insert/delete counts and interleave.LiveTarget;Checkerrors on hand-built invalid streams;n(including 0 and 2·len+5), andSettlerestores the start state.Compareof map vs Set3 on a cycle, followed bySettleand a size check.go test ./...,-raceandgolangci-lintare clean, and coverage is 99.4 %.Example and one finding
ExampleCyclecomparesmap[uint64]struct{}withSet3(already ingo.mod, so no new dependency) on a cycle over 100,000 elements. Along the way I measured the cost per operation across the phases of a cycle:The docs (Cursor, HOWTO) and the example therefore recommend one untimed cycle before measuring.
One more observation: on
main, the example's measured series starts at 68/56 ns and settles at 17/12 ns only after about 60 batches. With #112's warm-up it starts straight in the steady state. So that is the #111 effect again, not the workload.Docs
Not included (listed in the issue as possible extensions)
Lookup mix, Zipf skew,
Groupfor multimaps, the AgeBiased policy and a compact encoding.Addendum: one call for two questions (7984d77)
Correct use previously required running the first cycle untimed, pairing the cursor with its structure and calling Settle, and building both structures in a balanced order. Leaving out the first pass must not mean ignoring growth altogether, though, or structures with expensive growth would look better than they are. Hence:
Compare(target, a, b Structure, Options) (Result, error): eachStructure[S]only needsName,New() SandApply(S, []Op). It returns two reports that answer two different questions:SteadyState: cost per insert/delete once the structure is in use. It runs on a cycle; the first pass is automatically untimed.Build: cost of one complete build of a fresh structure. Creation, capacity hint, every resize and the garbage all count. Whole builds, because with segments the median would ignore exactly the rare expensive resizes.BuildRepeats31,BuildValidationRuns10, about 1,300 builds) and always setsGCBetween. With rtcompare's defaults it would take several minutes at 100,000 elements.SkipBuildexists as an escape hatch, and is documented as distorting the result.Replay: bundles the cycle, the apply function and the cursor per structure. The candidate primes automatically inSetupbefore the first batch, andSettleis a method.Checksizes its model for the peak state.ExampleCompare(map vs Set3, 100,000 elements) on all three branches merged together:On
mainalone, the steady-state part still shows the settling at the start ofCollect(drift −50 %). With #112's warm-up it is gone, as shown above.Tests:
go test ./...,-raceand fuzzing are green, lint reports 0 issues, and coverage is 98.2 %.🤖 Generated with Claude Code