diff --git a/HOWTO.md b/HOWTO.md index f44accd..3ca6f98 100644 --- a/HOWTO.md +++ b/HOWTO.md @@ -133,6 +133,101 @@ cost gets measured along with your code. See [Attenuation](#attenuation-your-number-is-real-but-smaller-than-the-truth) for what this costs you. +## Benchmarking insertions and deletions + +If your candidates are data structures and the question is how fast they +*change* — insert, delete, grow, shrink — the obvious batch is a trap: + +```go +Batch: func(n uint64) { + for i := range n { + m[key(i)] = value // insert + delete(m, key(i)) // and take it out again, so the map stays the same + } +} +``` + +That measures a structure that never changes shape. The element goes into +the same slot it just left; nothing splits, merges, resizes, rehashes or +leaves a tombstone, and those are exactly the costs that separate one data +structure from another in real use. The honest alternatives have their own +traps: picking what to delete inside the timed loop adds work to both +candidates and dilutes the difference; a stream that deletes elements that +aren't there times no-ops; a structure that keeps growing over the run is a +different structure at the last sample than at the first. + +The `workload` package does all of this. You describe each data structure by +how to create an empty one and how to apply an operation to it: + +```go +goMap := workload.Structure[map[uint64]struct{}]{ + Name: "map", + New: func() map[uint64]struct{} { return make(map[uint64]struct{}, 100_000) }, + Apply: func(m map[uint64]struct{}, run []workload.Op) { + for _, op := range run { + if op.Kind == workload.Insert { + m[uint64(op.ID)] = struct{}{} + } else { + delete(m, uint64(op.ID)) + } + } + }, +} +res, err := workload.Compare(100_000, goMap, otherSet, workload.Options{}) +fmt.Println(res) +``` + +The IDs are abstract; map them to your own keys and values, for example +through a precomputed slice of keys. You get **two answers**, because there are +two different questions, and mixing them into one number would get both +wrong: + +- **Steady state:** what does one insertion or deletion cost in a structure + that has been in use for a while? Both structures are filled to 100,000 + elements, then a *cycle* is replayed on them: bursts of insertions and + deletions of 100,000 further, transient elements, about 12,500 of them present + at a time, ending exactly where it started, so it can repeat endlessly. +- **Build:** what does it cost to build such a structure from empty, with the + same kind of back-and-forth along the way? Every sample is one complete + build of a fresh structure. + +The reason for the split: the first pass of a cycle is where a structure grows +to the largest size the cycle reaches. For a Go map that pass cost up to seven +times as much per operation as every pass after it. Timing it along with the +steady state would spread a one-time cost over the measurement, in a +proportion that depends on how long the run was. Dropping it would favour +structures whose growth is expensive. So it is measured where it belongs, in +the build, where creating the structure with or without a capacity hint, every +resize and the garbage all count, and the steady state starts after one +untimed cycle. If the two answers disagree, that is the result: one structure +can be faster to use and slower to build. + +What `workload.Compare` takes care of, so that you don't have to: + +- Every operation is valid by construction: nothing is inserted twice and + nothing is deleted that isn't there. +- The streams are precomputed; the timed loop only reads the next operation. +- Both structures are built in alternating chunks, so neither is the one built + last. +- The first pass of the cycle is untimed, and the position in the cycle stays + with each structure across batches. +- Whole builds take milliseconds, so the build comparison uses its own, smaller + defaults (31 repeats, 10 validation runs, about 1,300 builds in all) and + collects garbage between batches, so that one build's garbage is not + collected in the middle of the next. + +`workload.Config` tunes the streams: `Ratio` (insertions per element at rest, +default 2), burst length, and `Victims`, which chooses what a deletion removes: +`Uniform` (the default, like a general-purpose map), `FIFO` (a queue or a +retention window), or `LIFO` (a stack or undo log). + +For setups `Compare` does not cover, the parts are available on their own: +`workload.Cycle` and `workload.Build` make the streams, `workload.Replay` +replays a cycle on one structure (including the untimed first pass and +`Settle`, which returns the structure to its start state), and +`workload.Check` replays any stream against a model and reports the first +invalid operation. + ## The long version: what `Compare` does, step by step This is the sequence `Compare` runs automatically. Read it if you want to diff --git a/README.md b/README.md index 01e4686..004a881 100644 --- a/README.md +++ b/README.md @@ -24,6 +24,7 @@ Keywords: benchmarking, performance, bootstrap, runtime comparison, statistics, - Detect a trend across a measurement run, which resampling structurally cannot see because it discards the order the samples arrived in. - Resample in blocks when the measurements are correlated enough that treating them as independent would overstate confidence. - Deterministic PRNG for reproducible inputs, and a crypto/rand-backed one where unpredictability is wanted. +- Compare mutable data structures under realistic insert/delete workloads in one call (`workload`), with separate answers for the steady state and for building to size; the streams are valid by construction, reproducible from a seed, and cyclic, so that a batch can replay them endlessly without rebuilding the structure. ## What this cannot tell you @@ -162,6 +163,13 @@ Primitives: - `SampleTime()` / `DiffTimeStamps()` — high-resolution timestamps, and `GetSampleTimePrecision()` for the smallest interval they can resolve here. - `Median` / `QuickMedian` / `Statistics` — small statistics helpers. +Workloads for mutable data structures (`github.com/TomTonic/rtcompare/workload`): + +- `workload.Compare(target, a, b, Options)` — the whole job in one call: each `Structure` says how to create an empty structure and apply operations to it, and you get two reports, one for the steady state (per insertion or deletion, after one untimed cycle) and one for building from empty (per whole build, growth included). +- `workload.Cycle(target, Config)` / `workload.Build(target, Config)` — the streams: a cycle of insertions and deletions that ends where it started, and a build from empty with a realistic history. +- `workload.Replay` — replays a cycle on one structure instance, with the untimed first pass, and `Settle` to return it to its start state. `workload.Cursor` is the bare position, for doing it by hand. +- `workload.Check(ops, start, end)` — replays a stream against a model and reports the first invalid operation. + A note on the threshold of `0.0`: every threshold is evaluated as `delta >= t`, so at zero the question is "at least as fast", not "faster". Quantized timings tie often, and every tie counts towards it. Ask for a threshold above zero if you mean strictly faster. Note on negative `relativeGains`: Negative thresholds are allowed and are diff --git a/go.mod b/go.mod index 740d1df..402711d 100644 --- a/go.mod +++ b/go.mod @@ -1,6 +1,6 @@ module github.com/TomTonic/rtcompare -go 1.26.0 +go 1.26.8 require ( github.com/TomTonic/Set3 v0.4.2 diff --git a/workload/compare.go b/workload/compare.go new file mode 100644 index 0000000..37e3496 --- /dev/null +++ b/workload/compare.go @@ -0,0 +1,229 @@ +package workload + +import ( + "fmt" + "runtime" + "strings" + + "github.com/TomTonic/rtcompare" +) + +// Structure describes one data structure under test for [Compare]: how to +// create an empty one, and how to apply operations to it. +type Structure[S any] struct { + // Name labels the structure in the results. + Name string + + // New returns an empty structure, created the way the program would create + // it. Whether it is given a capacity hint is part of what the build + // comparison measures, so give it one exactly when the program would. + New func() S + + // Apply performs the operations of run on s, in order. Keep the loop over + // run in this function, where the compiler can inline the structure's + // methods. + Apply func(s S, run []Op) +} + +// check rejects a structure that cannot be used. +func (s Structure[S]) check(position string) error { + if s.New == nil || s.Apply == nil { + return fmt.Errorf("workload: structure %s (%q) needs both New and Apply", position, s.Name) + } + return nil +} + +// Options configures [Compare]. The zero value is usable and selects the +// documented defaults. +type Options struct { + // Config configures the two streams, see [Config]. + Config Config + + // SteadyState are the options of the steady-state comparison, passed to + // [rtcompare.Compare] unchanged. + SteadyState rtcompare.CompareOptions + + // Build are the options of the build comparison. Each of its samples is a + // whole build, which takes milliseconds rather than the microseconds of a + // calibrated batch, so a zero Collect.Repeats selects [BuildRepeats] and a + // zero ValidationRuns selects [BuildValidationRuns] rather than + // rtcompare's defaults, and Collect.GCBetween is always set, since every + // build leaves a whole structure of garbage behind that would otherwise be + // collected in the middle of the next candidate's build. + Build rtcompare.CompareOptions + + // SkipBuild omits the build comparison. The steady-state result alone + // ignores what growing to size costs, and so favours structures whose + // growth is expensive; set it only when that cost is known not to matter. + SkipBuild bool +} + +// BuildRepeats and BuildValidationRuns are the build comparison's defaults for +// Repeats and ValidationRuns. They are smaller than rtcompare's, because a +// sample is a whole build: with them a build comparison performs about 1,300 +// builds, some ten seconds for 100,000 elements, where rtcompare's defaults +// would take several minutes. +const ( + BuildRepeats = 31 + BuildValidationRuns = 10 +) + +// Result holds the two answers [Compare] gives, which are answers to two +// different questions and are deliberately not combined into one number. +type Result struct { + // Target is the number of elements at rest, and CycleOps and BuildOps the + // lengths of the two streams. + Target int + CycleOps, BuildOps int + + // SteadyState compares the cost of one insertion or deletion in a + // structure that has been in use for a while, averaged over the cycle. Its + // per-operation figures are nanoseconds per insertion or deletion. + SteadyState rtcompare.Report + + // Build compares the cost of building a structure from empty to Target + // elements, with the transient insertions and deletions a real build sees, + // including creation, every resize and the garbage it leaves. Its + // per-operation figures are nanoseconds per whole build. It is the zero + // Report when Options.SkipBuild was set. + Build rtcompare.Report +} + +// String renders both answers, each under the question it answers. +func (r Result) String() string { + var b strings.Builder + fmt.Fprintf(&b, "steady state, per insertion or deletion in a structure of %d elements (cycle of %d operations, first pass untimed):\n", r.Target, r.CycleOps) + b.WriteString(indent(r.SteadyState.String())) + if r.Build.SamplesA != nil { + fmt.Fprintf(&b, "\n\nbuild from empty to %d elements, per whole build (%d operations):\n", r.Target, r.BuildOps) + b.WriteString(indent(r.Build.String())) + } + return b.String() +} + +func indent(s string) string { + return " " + strings.ReplaceAll(s, "\n", "\n ") +} + +// Compare compares two data structures under a realistic mix of insertions +// and deletions, and answers two questions separately: what an operation costs +// once a structure is in use, and what it costs to build one. +// +// Parameters: target is the number of elements the structures hold at rest, +// at least one; a and b are the two structures, see [Structure]; opt +// configures the streams and both comparisons, see [Options]. +// +// It returns both reports, or an error if a structure lacks New or Apply, if +// the configuration is invalid (see [Build] and [Cycle]), or if a comparison +// fails. +// +// Use it for any comparison of mutable data structures. It exists because +// doing this by hand has several traps, and each of them quietly changes the +// answer: +// +// - The steady-state comparison builds one structure of each kind to target +// elements with a [Build] stream, applied to the two alternately in chunks +// so that neither is built last, and then replays a [Cycle] on each +// through a [Replay]. The first pass of the cycle is untimed: it is where +// the structure grows to the cycle's peak size, a one-time cost that would +// otherwise be spread over the measurement in proportions that depend on +// how long the run was. +// - That growth is not dropped, it is measured where it belongs. The build +// comparison times whole builds, each from a structure fresh from New to +// target elements, so creation, capacity hints, every resize and the +// garbage are included. A structure that cannot be presized and has to +// grow pays for it here, and one that grows cheaply shows it here. +// +// The two answers can disagree, and then that is the result: one structure +// can be faster to use and slower to build. Read both. +// +// The cost is that of two [rtcompare.Compare] calls, the build one dominated +// by the builds it performs, about 1,300 at the defaults; see +// [BuildRepeats]. +// +// res, err := workload.Compare(100_000, +// workload.Structure[map[uint64]struct{}]{Name: "map", New: newMap, Apply: applyToMap}, +// workload.Structure[*set3.Set3[uint64]]{Name: "Set3", New: newSet3, Apply: applyToSet3}, +// workload.Options{}) +// if err != nil { ... } +// fmt.Println(res) +func Compare[SA, SB any](target int, a Structure[SA], b Structure[SB], opt Options) (Result, error) { + if err := a.check("A"); err != nil { + return Result{}, err + } + if err := b.check("B"); err != nil { + return Result{}, err + } + cycle, err := Cycle(target, opt.Config) + if err != nil { + return Result{}, err + } + build, err := Build(target, opt.Config) + if err != nil { + return Result{}, err + } + res := Result{Target: target, CycleOps: len(cycle), BuildOps: len(build)} + + res.SteadyState, err = compareSteadyState(a, b, build, cycle, opt.SteadyState) + if err != nil { + return res, fmt.Errorf("workload: steady-state comparison: %w", err) + } + if !opt.SkipBuild { + res.Build, err = rtcompare.Compare(buildCandidate(a, build), buildCandidate(b, build), buildOptions(opt.Build)) + if err != nil { + return res, fmt.Errorf("workload: build comparison: %w", err) + } + } + return res, nil +} + +// compareSteadyState builds one structure of each kind, alternately, and +// compares the cycle replayed on each. +func compareSteadyState[SA, SB any](a Structure[SA], b Structure[SB], build, cycle []Op, opt rtcompare.CompareOptions) (rtcompare.Report, error) { + sa, sb := a.New(), b.New() + applyA := func(run []Op) { a.Apply(sa, run) } + applyB := func(run []Op) { b.Apply(sb, run) } + buildAlternately(build, applyA, applyB) + ra, rb := NewReplay(cycle, applyA), NewReplay(cycle, applyB) + return rtcompare.Compare(ra.Candidate(a.Name), rb.Candidate(b.Name), opt) +} + +// alternateChunk is how many operations of the build stream go to one +// structure before the other gets its turn. +const alternateChunk = 256 + +// buildAlternately applies the stream to both structures in alternating +// chunks. Built one after the other, the structure built second would be the +// one in the caches and in fresher memory when the measurement starts, which +// has been measured to make an identical structure a few percent faster. +func buildAlternately(ops []Op, applyA, applyB func([]Op)) { + for start := 0; start < len(ops); start += alternateChunk { + run := ops[start:min(start+alternateChunk, len(ops))] + applyA(run) + applyB(run) + } +} + +// buildCandidate times whole builds: each operation of its batch creates a +// structure and applies the entire build stream to it. +func buildCandidate[S any](s Structure[S], build []Op) rtcompare.Candidate { + return rtcompare.Candidate{Name: s.Name, Batch: func(n uint64) { + for range n { + fresh := s.New() + s.Apply(fresh, build) + runtime.KeepAlive(fresh) + } + }} +} + +// buildOptions fills in the build comparison's own defaults. +func buildOptions(opt rtcompare.CompareOptions) rtcompare.CompareOptions { + if opt.Collect.Repeats == 0 { + opt.Collect.Repeats = BuildRepeats + } + if opt.ValidationRuns == 0 { + opt.ValidationRuns = BuildValidationRuns + } + opt.Collect.GCBetween = true + return opt +} diff --git a/workload/compare_test.go b/workload/compare_test.go new file mode 100644 index 0000000..590d41e --- /dev/null +++ b/workload/compare_test.go @@ -0,0 +1,167 @@ +package workload + +import ( + "strings" + "testing" + + "github.com/TomTonic/rtcompare" +) + +// countingSet is a set that counts how many operations it was given, and +// fails the test on an invalid one, for checking what Compare and Replay do +// to the structures they are handed. +type countingSet struct { + t *testing.T + present map[uint32]bool + applied int +} + +func newCountingSet(t *testing.T) *countingSet { + return &countingSet{t: t, present: map[uint32]bool{}} +} + +func (s *countingSet) apply(run []Op) { + for _, op := range run { + if (op.Kind == Insert) == s.present[op.ID] { + s.t.Fatalf("invalid %s of element %d", op.Kind, op.ID) + } + s.present[op.ID] = op.Kind == Insert + if op.Kind == Delete { + delete(s.present, op.ID) + } + } + s.applied += len(run) +} + +// TestReplayPrimesOnceBeforeTheFirstBatch checks the automatic handling of a +// cycle's first pass, in which a structure grows to its peak size once and +// never again. Replay's candidate has to play one full cycle in its Setup, +// outside the measured region, before the first batch and never again; the +// batches have to continue the cycle validly from there, and Settle has to +// bring the structure back to its start state. +func TestReplayPrimesOnceBeforeTheFirstBatch(t *testing.T) { + const target = 200 + ops, err := Cycle(target, Config{Seed: 4}) + if err != nil { + t.Fatalf("unexpected error: %v", err) + } + s := newCountingSet(t) + s.apply(opsInserting(ids(target))) + s.applied = 0 + r := NewReplay(ops, s.apply) + c := r.Candidate("set") + if c.Name != "set" || c.Setup == nil { + t.Fatalf("candidate lacks its name or its priming Setup: %+v", c) + } + c.Setup() + if s.applied != len(ops) { + t.Errorf("priming applied %d operations, want one full cycle of %d", s.applied, len(ops)) + } + c.Batch(7) + c.Setup() + c.Batch(uint64(len(ops))) + if want := 2*len(ops) + 7; s.applied != want { + t.Errorf("after priming and two batches %d operations were applied, want %d", s.applied, want) + } + r.Settle() + if len(s.present) != target { + t.Errorf("%d elements present after Settle, want %d", len(s.present), target) + } +} + +// opsInserting returns a stream that inserts the given IDs. +func opsInserting(xs []uint32) []Op { + ops := make([]Op, len(xs)) + for i, id := range xs { + ops[i] = Op{ID: id, Kind: Insert} + } + return ops +} + +// TestBuildAlternatelyGivesNeitherStructureTheLastWord checks how Compare +// builds the two structures of the steady-state comparison: in alternating +// chunks, both receiving the whole stream in order, so that neither is the one +// built last. Built one after the other, the second has been measured to be a +// few percent faster for that alone. +func TestBuildAlternatelyGivesNeitherStructureTheLastWord(t *testing.T) { + ops := opsInserting(ids(3*alternateChunk + 10)) + var order []string + var gotA, gotB []Op + buildAlternately(ops, + func(run []Op) { order = append(order, "A"); gotA = append(gotA, run...) }, + func(run []Op) { order = append(order, "B"); gotB = append(gotB, run...) }) + if got := strings.Join(order, ""); got != "ABABABAB" { + t.Errorf("chunk order: got %s, want ABABABAB", got) + } + if len(gotA) != len(ops) || len(gotB) != len(ops) || gotA[len(ops)-1] != ops[len(ops)-1] { + t.Errorf("each structure should receive the whole stream in order: %d and %d of %d", len(gotA), len(gotB), len(ops)) + } +} + +// TestBuildCandidateTimesWholeBuilds checks the unit of the build comparison: +// one operation is one complete build of a fresh structure, so that creation, +// resizing and everything else growth costs lands in each sample whole rather +// than in a few batches the median would ignore. +func TestBuildCandidateTimesWholeBuilds(t *testing.T) { + build, err := Build(100, Config{Seed: 2}) + if err != nil { + t.Fatalf("unexpected error: %v", err) + } + var sets []*countingSet + s := Structure[*countingSet]{ + Name: "counting", + New: func() *countingSet { cs := newCountingSet(t); sets = append(sets, cs); return cs }, + Apply: func(cs *countingSet, run []Op) { cs.apply(run) }, + } + buildCandidate(s, build).Batch(3) + if len(sets) != 3 { + t.Fatalf("expected 3 fresh structures, got %d", len(sets)) + } + for i, cs := range sets { + if cs.applied != len(build) || len(cs.present) != 100 { + t.Errorf("build %d: %d operations applied, %d elements present", i, cs.applied, len(cs.present)) + } + } +} + +// TestBuildOptionsDefaults checks the build comparison's own defaults, which +// keep a comparison of whole builds affordable, and that it always collects +// garbage between batches while leaving explicit settings alone. +func TestBuildOptionsDefaults(t *testing.T) { + got := buildOptions(rtcompare.CompareOptions{}) + if got.Collect.Repeats != BuildRepeats || got.ValidationRuns != BuildValidationRuns || !got.Collect.GCBetween { + t.Errorf("defaults not applied: %+v", got) + } + got = buildOptions(rtcompare.CompareOptions{Collect: rtcompare.CollectOptions{Repeats: 51}, ValidationRuns: 20}) + if got.Collect.Repeats != 51 || got.ValidationRuns != 20 || !got.Collect.GCBetween { + t.Errorf("explicit settings not kept: %+v", got) + } +} + +// TestCompareRejectsIncompleteInput checks that Compare fails before +// measuring anything when a structure cannot be used or the streams cannot be +// generated, and names the problem. +func TestCompareRejectsIncompleteInput(t *testing.T) { + ok := Structure[*countingSet]{Name: "ok", New: func() *countingSet { return newCountingSet(t) }, Apply: func(cs *countingSet, run []Op) { cs.apply(run) }} + cases := []struct { + name string + a Structure[*countingSet] + target int + want string + }{ + {"returns error for a structure without New", Structure[*countingSet]{Name: "x", Apply: ok.Apply}, 10, `A ("x") needs both`}, + {"returns error for a structure without Apply", Structure[*countingSet]{Name: "x", New: ok.New}, 10, "needs both"}, + {"returns error for an empty target", ok, 0, "target"}, + } + for _, c := range cases { + t.Run(c.name, func(t *testing.T) { + _, err := Compare(c.target, c.a, ok, Options{}) + if err == nil || !strings.Contains(err.Error(), c.want) { + t.Errorf("got %v, want an error mentioning %q", err, c.want) + } + }) + } + if _, err := Compare(10, ok, Structure[*countingSet]{Name: "y"}, Options{}); err == nil || !strings.Contains(err.Error(), `B ("y")`) { + t.Errorf("got %v, want an error naming structure B", err) + } +} diff --git a/workload/example_test.go b/workload/example_test.go new file mode 100644 index 0000000..36db158 --- /dev/null +++ b/workload/example_test.go @@ -0,0 +1,106 @@ +package workload_test + +import ( + "fmt" + "strings" + "testing" + + set3 "github.com/TomTonic/Set3" + "github.com/TomTonic/rtcompare" + "github.com/TomTonic/rtcompare/workload" +) + +// goMap and set3Set describe the two sets the example compares. Both are +// created with a capacity hint for the elements at rest, as a program that +// knows its size would. +func goMap(target int) workload.Structure[map[uint64]struct{}] { + return workload.Structure[map[uint64]struct{}]{ + Name: "map", + New: func() map[uint64]struct{} { return make(map[uint64]struct{}, target) }, + Apply: func(m map[uint64]struct{}, run []workload.Op) { + for _, op := range run { + if op.Kind == workload.Insert { + m[uint64(op.ID)] = struct{}{} + } else { + delete(m, uint64(op.ID)) + } + } + }, + } +} + +func set3Set(target int) workload.Structure[*set3.Set3[uint64]] { + return workload.Structure[*set3.Set3[uint64]]{ + Name: "Set3", + New: func() *set3.Set3[uint64] { return set3.EmptyWithCapacity[uint64](uint32(target)) }, + Apply: func(s *set3.Set3[uint64], run []workload.Op) { + for _, op := range run { + if op.Kind == workload.Insert { + s.Add(uint64(op.ID)) + } else { + s.Remove(uint64(op.ID)) + } + } + }, + } +} + +// ExampleCompare compares Go's built-in map with Set3 under a realistic mix of +// insertions and deletions over 100,000 elements, and answers two questions: +// what an insertion or deletion costs once a set is in use, and what it costs +// to build one. Everything else, the streams, building both sets, the +// untimed first pass and the cursors, is done by Compare. +func ExampleCompare() { + const target = 100_000 + res, err := workload.Compare(target, goMap(target), set3Set(target), workload.Options{}) + if err != nil { + panic(err) + } + fmt.Println(res) +} + +// TestCompareAnswersBothQuestions runs a reduced version of ExampleCompare, +// which is not run by go test since its output varies. It checks the one-call +// comparison of two data structures end to end: both the steady-state and the +// build comparison have to run on real sets and be reported under their own +// headings, and SkipBuild has to leave the build comparison out. +func TestCompareAnswersBothQuestions(t *testing.T) { + const target = 3000 + quick := rtcompare.CompareOptions{ + Collect: rtcompare.CollectOptions{Repeats: 11, InnerLoops: 500}, + SkipValidation: true, + Resamples: 300, + } + res, err := workload.Compare(target, goMap(target), set3Set(target), workload.Options{ + SteadyState: quick, + Build: rtcompare.CompareOptions{Collect: rtcompare.CollectOptions{Repeats: 11, InnerLoops: 1}, SkipValidation: true, Resamples: 300}, + }) + if err != nil { + t.Fatalf("unexpected error: %v", err) + } + if len(res.SteadyState.SamplesA) != 11 || len(res.Build.SamplesA) != 11 { + t.Errorf("expected 11 samples in each comparison, got %d and %d", len(res.SteadyState.SamplesA), len(res.Build.SamplesA)) + } + if res.CycleOps == 0 || res.BuildOps == 0 || res.Target != target { + t.Errorf("stream lengths not reported: %+v", res) + } + // A whole build of thousands of elements costs far more than one of its + // operations. + if res.Build.NsPerOpA < 100*res.SteadyState.NsPerOpA { + t.Errorf("a build (%.0f ns) should cost far more than one operation (%.1f ns)", res.Build.NsPerOpA, res.SteadyState.NsPerOpA) + } + out := res.String() + for _, want := range []string{"steady state, per insertion or deletion", "build from empty to 3000 elements, per whole build"} { + if !strings.Contains(out, want) { + t.Errorf("output lacks %q:\n%s", want, out) + } + } + + res, err = workload.Compare(target, goMap(target), set3Set(target), workload.Options{SteadyState: quick, SkipBuild: true}) + if err != nil { + t.Fatalf("unexpected error: %v", err) + } + if res.Build.SamplesA != nil || strings.Contains(res.String(), "build from empty") { + t.Errorf("SkipBuild should leave the build comparison out:\n%s", res) + } +} diff --git a/workload/replay.go b/workload/replay.go new file mode 100644 index 0000000..7829b56 --- /dev/null +++ b/workload/replay.go @@ -0,0 +1,218 @@ +package workload + +import ( + "fmt" + "slices" + + "github.com/TomTonic/rtcompare" +) + +// Cursor is a position in a [Cycle], and belongs to the data structure +// instance the cycle is replayed on. +// +// It exists because the position is a property of the structure, not of the +// batch function: after half a cycle, the structure holds the transient +// elements of that half, and only a replay that continues from there stays +// valid. A new batch closure that started again at zero would insert elements +// that are already present. Keep one Cursor per instance, next to it, and use +// it for every comparison the instance takes part in. +// +// The zero value is at the start of the cycle. +type Cursor struct { + pos int +} + +// Position returns the index of the next operation to be applied. +func (c *Cursor) Position() int { return c.pos } + +// Batch returns an [rtcompare.Batch] that applies the next n operations of ops +// on every call, wrapping around at the end of the cycle. +// +// Parameters: ops is a stream from [Cycle], and must be the same on every call +// with this Cursor; apply performs the operations it is handed on the data +// structure. apply receives contiguous runs of the stream, so the loop over +// them sits in the caller's code, where the compiler can inline the data +// structure's methods; a run is split at the end of the cycle, so apply may be +// called more than once per batch. +// +// Use it as a candidate's Batch. One operation is one unit of work: the +// per-operation cost rtcompare reports is the average of the insertions and +// deletions in the stream. +// +// Prefer [Replay], which also plays the first pass untimed; with a bare Cursor +// that is up to the caller, see [Replay] for why it matters. +// +// cur := &workload.Cursor{} +// candidate := rtcompare.Candidate{Name: "map", Batch: cur.Batch(ops, func(run []workload.Op) { +// for _, op := range run { +// if op.Kind == workload.Insert { +// m[uint64(op.ID)] = struct{}{} +// } else { +// delete(m, uint64(op.ID)) +// } +// } +// })} +// +// An empty ops makes a batch that does nothing. +func (c *Cursor) Batch(ops []Op, apply func([]Op)) rtcompare.Batch { + return func(n uint64) { c.Advance(ops, n, apply) } +} + +// Advance applies the next n operations of ops through apply, wrapping around +// at the end of the cycle, and moves the cursor past them. It is what the +// function returned by Batch does, for callers who drive the replay +// themselves. +func (c *Cursor) Advance(ops []Op, n uint64, apply func([]Op)) { + if len(ops) == 0 { + return + } + for n > 0 { + run := ops[c.pos:] + if uint64(len(run)) > n { + run = run[:n] + } + apply(run) + n -= uint64(len(run)) + c.pos += len(run) + if c.pos == len(ops) { + c.pos = 0 + } + } +} + +// Settle applies the rest of the cycle, if the cursor is not at its start, so +// that the structure is back in the state the cycle starts from: exactly the +// IDs 0 to target-1 present. +// +// Call it outside the measured region, for instance in a Candidate's Teardown +// or between comparisons, before anything else is measured on the structure. +func (c *Cursor) Settle(ops []Op, apply func([]Op)) { + if c.pos == 0 || len(ops) == 0 { + return + } + apply(ops[c.pos:]) + c.pos = 0 +} + +// Check replays ops against a model set that starts as start, and reports the +// first operation that would be invalid: an insertion of a present element, a +// deletion of an absent one, or an unknown Kind. It also reports whether the +// model ends as end, ignoring order. It returns nil if the stream is valid. +// +// Use it in a test of anything that produces or transforms a stream. For a +// [Build] stream over target elements, start is empty and end is 0 to +// target-1; for a [Cycle], both are 0 to target-1. +func Check(ops []Op, start, end []uint32) error { + // Sized for the largest the model can get, so that it never grows while + // the stream is replayed. + inserts := 0 + for _, op := range ops { + if op.Kind == Insert { + inserts++ + } + } + present := make(map[uint32]struct{}, len(start)+inserts) + for _, id := range start { + if _, dup := present[id]; dup { + return fmt.Errorf("workload: start lists element %d twice", id) + } + present[id] = struct{}{} + } + for i, op := range ops { + _, has := present[op.ID] + switch { + case op.Kind == Insert && has: + return fmt.Errorf("workload: op %d inserts element %d, which is already present", i, op.ID) + case op.Kind == Insert: + present[op.ID] = struct{}{} + case op.Kind == Delete && !has: + return fmt.Errorf("workload: op %d deletes element %d, which is not present", i, op.ID) + case op.Kind == Delete: + delete(present, op.ID) + default: + return fmt.Errorf("workload: op %d has unknown %s", i, op.Kind) + } + } + return compareEnd(present, end) +} + +// compareEnd reports the first difference between the model's final contents +// and the expected ones. +func compareEnd(present map[uint32]struct{}, end []uint32) error { + want := make(map[uint32]struct{}, len(end)) + for _, id := range end { + want[id] = struct{}{} + if _, ok := present[id]; !ok { + return fmt.Errorf("workload: element %d should be present at the end, but is not", id) + } + } + var extra []uint32 + for id := range present { + if _, ok := want[id]; !ok { + extra = append(extra, id) + } + } + if len(extra) > 0 { + slices.Sort(extra) + return fmt.Errorf("workload: %d elements are present at the end that should not be, the first being %d", len(extra), extra[0]) + } + return nil +} + +// Replay plays a [Cycle] on one data structure instance. It bundles the +// stream, the function that applies operations to the structure, and the +// position in the cycle, so that they cannot be mismatched: create one Replay +// per structure instance, next to the structure. +// +// It also takes care of the first pass. That pass is where the structure grows +// to the largest size the cycle reaches, and it costs far more than every pass +// after it: for a Go map of 100,000 elements at Ratio 2, stretches of the first +// pass cost 35 to 128 ns per operation where later passes cost 17 to 20. That +// growth happens once in a program's life and would otherwise be spread over +// the measurement, in proportions that depend on how long it ran. The +// candidate from [Replay.Candidate] therefore plays one full cycle, untimed, +// before its first measured batch. What the growth costs is a question of its +// own, which [Compare] answers separately with a [Build] stream. +type Replay struct { + ops []Op + apply func([]Op) + cursor Cursor + primed bool +} + +// NewReplay returns a Replay of ops, a stream from [Cycle], applied to one +// structure through apply. apply receives contiguous runs of the stream, so +// that its loop over them sits where the compiler can inline the structure's +// methods; see [Cursor.Batch]. +func NewReplay(ops []Op, apply func([]Op)) *Replay { + return &Replay{ops: ops, apply: apply} +} + +// Prime plays one full cycle, untimed, the first time it is called, and does +// nothing after that. The candidate calls it before its first batch; call it +// yourself only when driving the replay by other means. +func (r *Replay) Prime() { + if r.primed { + return + } + r.primed = true + r.cursor.Advance(r.ops, uint64(len(r.ops)), r.apply) +} + +// Candidate returns an [rtcompare.Candidate] that replays the cycle, continuing +// where the previous batch stopped. Its Setup primes the structure before the +// first batch, outside the measured region, and does nothing after that. +func (r *Replay) Candidate(name string) rtcompare.Candidate { + return rtcompare.Candidate{ + Name: name, + Setup: r.Prime, + Batch: func(n uint64) { r.cursor.Advance(r.ops, n, r.apply) }, + } +} + +// Settle plays the rest of the cycle, so that the structure is back in the +// state the cycle starts from. Call it before measuring anything else on the +// structure. +func (r *Replay) Settle() { + r.cursor.Settle(r.ops, r.apply) +} diff --git a/workload/simulate.go b/workload/simulate.go new file mode 100644 index 0000000..2c6d569 --- /dev/null +++ b/workload/simulate.go @@ -0,0 +1,158 @@ +package workload + +import "github.com/TomTonic/rtcompare" + +// simulation produces a stream by playing it out: it tracks which transient +// elements are present at every point, so that a deletion can only name one +// that is and an insertion only one that is not. That is what makes the +// streams valid by construction. +type simulation struct { + c Config + rng rtcompare.DPRNG + ops []Op + + // permanent holds the elements to be inserted that remain at the end, + // taken from the back; empty for a cycle, whose permanent elements are + // present from the start. + permanent []uint32 + + nextTransient uint32 // the next fresh transient ID + transientsLeft int // transient insertions still to be made + live liveSet + + // meanTransients is the expected number of transient insertions per burst, + // which the deletion bursts are scaled to so that the number of transient + // elements present settles at LiveTarget. + meanTransients float64 +} + +func newSimulation(c Config, target, transients int) *simulation { + return &simulation{ + c: c, + rng: rtcompare.NewDPRNG(mixSeed(c.Seed)), + nextTransient: uint32(target), + transientsLeft: transients, + live: liveSet{policy: c.Victims}, + } +} + +// run alternates bursts of insertions and deletions until every insertion is +// made, then deletes the transient elements still present. +func (s *simulation) run() { + pending := len(s.permanent) + s.transientsLeft + s.ops = make([]Op, 0, pending+s.transientsLeft) + meanBurst := float64(s.c.MaxBurst+1) / 2 + s.meanTransients = meanBurst * float64(s.transientsLeft) / float64(pending) + for len(s.permanent)+s.transientsLeft > 0 { + s.insertBurst() + s.deleteBurst() + } + for s.live.len() > 0 { + for range min(1+int(s.rng.Uint32N(uint32(s.c.MaxBurst))), s.live.len()) { + s.ops = append(s.ops, Op{ID: s.live.remove(&s.rng), Kind: Delete}) + } + } +} + +// insertBurst inserts between 1 and MaxBurst elements. Each one is transient +// or permanent in proportion to how many of each are left, so that both kinds +// are spread evenly over the stream. +func (s *simulation) insertBurst() { + for range 1 + s.rng.Uint32N(uint32(s.c.MaxBurst)) { + left := len(s.permanent) + s.transientsLeft + if left == 0 { + return + } + if int(s.rng.Uint32N(uint32(left))) < s.transientsLeft { + id := s.nextTransient + s.nextTransient++ + s.transientsLeft-- + s.live.add(id) + s.ops = append(s.ops, Op{ID: id, Kind: Insert}) + continue + } + last := len(s.permanent) - 1 + s.ops = append(s.ops, Op{ID: s.permanent[last], Kind: Insert}) + s.permanent = s.permanent[:last] + } +} + +// deleteBurst deletes a number of transient elements whose expectation is the +// expected number of transient insertions per burst, scaled by how far the +// elements present exceed LiveTarget. At LiveTarget insertions and deletions +// balance, so that is where the count settles. +func (s *simulation) deleteBurst() { + n := s.live.len() + if n == 0 { + return + } + mean := s.meanTransients * float64(n) / float64(s.c.LiveTarget) + // Uniform on [0, 2*mean], rounded stochastically so that the expectation + // survives even when it is a small fraction: plain rounding would turn + // every burst below one half into no deletion at all, and a stream with + // few transients would then delete them all at the end, in one block. + x := 2 * mean * s.rng.Float64() + d := int(x) + if s.rng.Float64() < x-float64(d) { + d++ + } + d = min(d, n) + for range d { + s.ops = append(s.ops, Op{ID: s.live.remove(&s.rng), Kind: Delete}) + } +} + +// liveSet holds the transient elements present, in the order of their +// insertion as far as the policy needs it. +type liveSet struct { + policy Policy + ids []uint32 + head int // FIFO only: the oldest element still present +} + +func (l *liveSet) len() int { return len(l.ids) - l.head } + +func (l *liveSet) add(id uint32) { l.ids = append(l.ids, id) } + +// remove takes out the element the policy selects and returns it. +func (l *liveSet) remove(rng *rtcompare.DPRNG) uint32 { + switch l.policy { + case FIFO: + id := l.ids[l.head] + l.head++ + if l.head >= 1024 && 2*l.head >= len(l.ids) { + // Reclaim the consumed front so the slice does not grow without + // bound over a long stream. + l.ids = append(l.ids[:0], l.ids[l.head:]...) + l.head = 0 + } + return id + case LIFO: + last := len(l.ids) - 1 + id := l.ids[last] + l.ids = l.ids[:last] + return id + default: + i := rng.Uint32N(uint32(len(l.ids))) + id := l.ids[i] + last := len(l.ids) - 1 + l.ids[i] = l.ids[last] + l.ids = l.ids[:last] + return id + } +} + +// mixSeed spreads a seed over all 64 bits with one round of splitmix64, so +// that small seeds do not start the generator in a low-entropy state and so +// that zero stays a fixed seed, which rtcompare.NewDPRNG would otherwise take +// as a request for a random one. +func mixSeed(seed uint64) uint64 { + z := seed + 0x9E3779B97F4A7C15 + z = (z ^ (z >> 30)) * 0xBF58476D1CE4E5B9 + z = (z ^ (z >> 27)) * 0x94D049BB133111EB + z ^= z >> 31 + if z == 0 { + return 1 + } + return z +} diff --git a/workload/workload.go b/workload/workload.go new file mode 100644 index 0000000..8552e89 --- /dev/null +++ b/workload/workload.go @@ -0,0 +1,255 @@ +// Package workload generates realistic streams of insertions and deletions for +// benchmarking mutable data structures such as maps, sets, trees, tries, +// indexes, caches and queues with rtcompare. +// +// A hand-written "insert an element, then remove the same element" loop +// measures a structure that never changes shape: the node splits, merges, +// resizes, rehashes and tombstones that dominate real workloads never happen. +// Getting a realistic stream right by hand runs into four traps, and this +// package exists to avoid them: +// +// 1. Invalid operations. A deletion must target an element that is present +// at that point of the stream, and an insertion one that is not, or part of +// the timed work is a no-op. Every stream here is valid by construction, +// and [Check] proves it against a model. +// 2. Overhead in the timed loop. Choosing a victim while timing adds cache +// misses and branches to every candidate and dilutes the difference between +// them. The streams are precomputed; the timed loop only reads the next +// entry, sequentially. +// 3. State drift across batches. rtcompare calls a batch many times, and a +// structure that grows or shrinks without bound is a different structure at +// the last sample than at the first. A [Cycle] ends exactly in its start +// state, so a [Replay] can repeat it endlessly without an untimed rebuild. +// 4. Shared state across comparisons. The position in a cycle belongs to the +// data structure instance, not to the batch closure. A [Replay] carries +// it with the structure, and [Replay.Settle] brings the structure back to +// its start state before anything else is measured on it. +// +// Two more traps sit in the measurement itself. The first pass of a cycle is +// where a structure grows to its peak size, a one-time cost that must neither +// be spread over a steady-state measurement nor be dropped. And whichever of +// two structures is built last starts ahead. +// +// [Compare] does all of this in one call and answers the two questions a +// mutation benchmark has, what an operation costs in steady state and what a +// build costs, separately. [Cycle], [Build], [Replay], [Cursor] and [Check] +// are the parts it is made of, for setups it does not cover. +// +// The generator works on abstract element IDs; the caller maps them to its own +// keys and values. IDs 0 to target-1 are the elements the structure holds at +// rest, and transient IDs follow from target upwards, each inserted once and +// deleted once. +package workload + +import ( + "fmt" + "math" + + "github.com/TomTonic/rtcompare" +) + +// Kind is what an [Op] does. +type Kind uint8 + +const ( + // Insert adds an element that is not present. + Insert Kind = iota + // Delete removes an element that is present. + Delete +) + +// String implements [fmt.Stringer]. +func (k Kind) String() string { + switch k { + case Insert: + return "insert" + case Delete: + return "delete" + default: + return fmt.Sprintf("Kind(%d)", uint8(k)) + } +} + +// Op is one mutation of a stream: an element ID and what to do with it. It +// occupies 8 bytes, so a cycle of 12.7 million operations takes about 100 MB. +type Op struct { + ID uint32 + Kind Kind +} + +// Policy decides which present transient element a deletion removes. +type Policy int + +const ( + // Uniform deletes a transient element drawn uniformly from those present, + // so that some are deleted soon after their insertion and some long after. + // It is the zero value, and what a general-purpose map or set sees. + Uniform Policy = iota + + // FIFO deletes the oldest transient element present, as a queue, a + // time-series retention window or a log compaction does. + FIFO + + // LIFO deletes the youngest transient element present, as a stack or an + // undo log does. + LIFO +) + +// String implements [fmt.Stringer]. +func (p Policy) String() string { + switch p { + case Uniform: + return "Uniform" + case FIFO: + return "FIFO" + case LIFO: + return "LIFO" + default: + return fmt.Sprintf("Policy(%d)", int(p)) + } +} + +// DefaultRatio is the Config.Ratio used when it is left at zero: twice as many +// insertions as elements at rest, within the 1.5 to 3 that resembles a +// database index. +const DefaultRatio = 2.0 + +// DefaultMaxBurst is the Config.MaxBurst used when it is left at zero. +const DefaultMaxBurst = 16 + +// Config configures [Build] and [Cycle]. The zero value is usable and selects +// the documented defaults. +type Config struct { + // Seed makes a stream reproducible: the same seed and configuration give + // the same stream. Zero is a seed like any other. + Seed uint64 + + // Ratio is the number of insertions per element at rest. Build inserts + // Ratio*target elements in all, of which target remain; Cycle inserts and + // deletes (Ratio-1)*target transient elements. Zero selects DefaultRatio. + // Build requires at least 1, Cycle more than 1. + Ratio float64 + + // MaxBurst bounds the length of a burst of insertions, which is drawn + // uniformly from 1 to MaxBurst. Zero selects DefaultMaxBurst. + MaxBurst int + + // LiveTarget is the number of transient elements the stream keeps present + // on average: the more are present, the longer the deletion bursts. Zero + // selects target/8, but no more than half the transient elements, and at + // least one. + LiveTarget int + + // Victims chooses which transient element a deletion removes. The zero + // value is Uniform. + Victims Policy +} + +// Build returns a stream that takes a structure from empty to holding exactly +// the IDs 0 to target-1, the way it would have got there in real use: with +// transient elements inserted and deleted along the way. +// +// Parameters: target is the number of elements at the end, at least one; c +// configures the stream, see [Config]. With Ratio r the stream holds r*target +// insertions, target of them permanent and the rest transient, and a deletion +// for each transient one. Bursts of insertions alternate with bursts of +// deletions, and the permanent elements are inserted in random order, spread +// across the whole stream. +// +// It returns the stream, or an error for a target below one, an invalid +// configuration, or more elements than fit in a uint32 ID. +// +// Use it to build a fixture with a realistic history, since a structure built +// by inserting its final elements in order can look quite different inside +// from one that has seen deletions, or to time the build itself. Replay it +// once, from the start; unlike a [Cycle] it does not return to where it began. +func Build(target int, c Config) ([]Op, error) { + c, transients, err := c.resolve(target, 1) + if err != nil { + return nil, err + } + s := newSimulation(c, target, transients) + s.permanent = shuffledIDs(target, &s.rng) + s.run() + return s.ops, nil +} + +// Cycle returns a stream that starts with the IDs 0 to target-1 present, +// inserts and deletes transient elements, and ends with exactly the IDs 0 to +// target-1 present again, so that it can be replayed endlessly. +// +// Parameters: target is the number of elements present at rest, at least +// one; c configures the stream, see [Config]. With Ratio r the cycle inserts +// and deletes (r-1)*target transient elements in alternating bursts, keeping +// about LiveTarget of them present at a time. The elements present at rest are +// never deleted. +// +// It returns the stream, or an error for a target below one, an invalid +// configuration, a ratio at which no transient element would be inserted, or +// more elements than fit in a uint32 ID. +// +// Use it for steady-state mutation benchmarks: build the structure with the +// IDs 0 to target-1, then replay the cycle through a [Replay] in the batch. +// Because the cycle ends in its start state, a batch can wrap around to the +// beginning, and the structure the last sample measures is the one the first +// sample measured. At 1M elements and r = 2 a cycle has 2M operations and +// takes 16 MB. +func Cycle(target int, c Config) ([]Op, error) { + c, transients, err := c.resolve(target, 0) + if err != nil { + return nil, err + } + if transients == 0 { + return nil, fmt.Errorf("workload: a cycle with Ratio %v over %d elements inserts no transient element; raise Ratio", c.Ratio, target) + } + s := newSimulation(c, target, transients) + s.run() + return s.ops, nil +} + +// resolve checks the configuration, fills in its defaults and returns the +// number of transient elements. minRatio is 1 for Build, where a ratio of one +// is a plain build, and 0 for Cycle, whose ratio must exceed one. +func (c Config) resolve(target int, minRatio float64) (Config, int, error) { + if target < 1 { + return c, 0, fmt.Errorf("workload: target must be at least 1, got %d", target) + } + if c.Ratio == 0 { + c.Ratio = DefaultRatio + } + if math.IsNaN(c.Ratio) || math.IsInf(c.Ratio, 0) || c.Ratio < 1 || (minRatio == 0 && c.Ratio <= 1) { + return c, 0, fmt.Errorf("workload: Ratio must be at least 1 for Build and above 1 for Cycle, got %v", c.Ratio) + } + if c.MaxBurst < 0 || c.LiveTarget < 0 { + return c, 0, fmt.Errorf("workload: MaxBurst and LiveTarget must not be negative, got %d and %d", c.MaxBurst, c.LiveTarget) + } + switch c.Victims { + case Uniform, FIFO, LIFO: + default: + return c, 0, fmt.Errorf("workload: unknown Victims policy %d", int(c.Victims)) + } + transients := math.Round((c.Ratio - 1) * float64(target)) + if float64(target)+transients > math.MaxUint32 { + return c, 0, fmt.Errorf("workload: %d elements and %.0f transient ones do not fit in uint32 IDs", target, transients) + } + if c.MaxBurst == 0 { + c.MaxBurst = DefaultMaxBurst + } + if c.LiveTarget == 0 { + c.LiveTarget = max(1, min(target/8, int(transients)/2)) + } + return c, int(transients), nil +} + +// shuffledIDs returns 0 to n-1 in random order, by Fisher-Yates. +func shuffledIDs(n int, rng *rtcompare.DPRNG) []uint32 { + ids := make([]uint32, n) + for i := range ids { + ids[i] = uint32(i) + } + for i := n - 1; i > 0; i-- { + j := rng.Uint32N(uint32(i + 1)) + ids[i], ids[j] = ids[j], ids[i] + } + return ids +} diff --git a/workload/workload_test.go b/workload/workload_test.go new file mode 100644 index 0000000..9857fb5 --- /dev/null +++ b/workload/workload_test.go @@ -0,0 +1,384 @@ +package workload + +import ( + "fmt" + "math" + "slices" + "strings" + "testing" +) + +// ids returns 0 to n-1. +func ids(n int) []uint32 { + out := make([]uint32, n) + for i := range out { + out[i] = uint32(i) + } + return out +} + +// counts returns the number of insertions and deletions and how often the +// stream switches between the two. +func counts(ops []Op) (inserts, deletes, switches int) { + for i, op := range ops { + if op.Kind == Insert { + inserts++ + } else { + deletes++ + } + if i > 0 && op.Kind != ops[i-1].Kind { + switches++ + } + } + return inserts, deletes, switches +} + +// TestStreamsAreValidByConstruction checks the central promise of the +// generator: every operation of every stream does real work. It belongs to the +// workload package, whose streams drive mutation benchmarks through rtcompare. +// For each ratio, size and deletion policy, a Build has to take an empty model +// to exactly the IDs 0 to target-1 and a Cycle has to return the model to that +// state, with no insertion of a present element and no deletion of an absent +// one; the insertion counts have to match the ratio, and insertions and +// deletions have to interleave rather than come in two blocks. +func TestStreamsAreValidByConstruction(t *testing.T) { + for _, target := range []int{1, 10, 1000, 20000} { + for _, ratio := range []float64{1, 1.01, 1.5, 2, 3} { + for _, policy := range []Policy{Uniform, FIFO, LIFO} { + name := fmt.Sprintf("target %d ratio %v %s", target, ratio, policy) + t.Run("build is valid for "+name, func(t *testing.T) { + checkStream(t, target, ratio, policy, true) + }) + if ratio > 1 && math.Round((ratio-1)*float64(target)) > 0 { + t.Run("cycle is valid for "+name, func(t *testing.T) { + checkStream(t, target, ratio, policy, false) + }) + } + } + } + } +} + +func checkStream(t *testing.T, target int, ratio float64, policy Policy, build bool) { + t.Helper() + c := Config{Seed: 3, Ratio: ratio, Victims: policy} + transients := int(math.Round((ratio - 1) * float64(target))) + var ops []Op + var err error + var start []uint32 + wantInserts := transients + if build { + ops, err = Build(target, c) + wantInserts += target + } else { + ops, err = Cycle(target, c) + start = ids(target) + } + if err != nil { + t.Fatalf("unexpected error: %v", err) + } + if err := Check(ops, start, ids(target)); err != nil { + t.Fatal(err) + } + inserts, deletes, switches := counts(ops) + if inserts != wantInserts || deletes != transients { + t.Errorf("got %d insertions and %d deletions, want %d and %d", inserts, deletes, wantInserts, transients) + } + // Interleaving: with enough transients the stream must change direction + // many times, not insert everything and then delete everything. + if transients >= 100 && switches < transients/20 { + t.Errorf("only %d switches between insertions and deletions for %d transients", switches, transients) + } +} + +// TestStreamsAreDeterministic checks that a benchmark built on a stream can be +// repeated exactly: the same seed and configuration have to give the same +// stream, including seed zero, and a different seed a different one. +func TestStreamsAreDeterministic(t *testing.T) { + for _, gen := range []struct { + name string + f func(int, Config) ([]Op, error) + }{{"Build", Build}, {"Cycle", Cycle}} { + t.Run(gen.name+" repeats for the same seed", func(t *testing.T) { + a, errA := gen.f(5000, Config{Seed: 0}) + b, errB := gen.f(5000, Config{Seed: 0}) + c, errC := gen.f(5000, Config{Seed: 1}) + if errA != nil || errB != nil || errC != nil { + t.Fatalf("unexpected errors: %v %v %v", errA, errB, errC) + } + if !slices.Equal(a, b) { + t.Error("the same seed gave two different streams") + } + if slices.Equal(a, c) { + t.Error("seeds 0 and 1 gave the same stream") + } + }) + } +} + +// TestPoliciesChooseTheirVictims checks that each deletion policy deletes what +// its name promises, since that is what makes a stream resemble a queue, a +// stack or a general-purpose set. Replaying a cycle, a FIFO deletion has to +// remove the oldest transient element present, a LIFO deletion the youngest, +// and a Uniform one neither consistently. +func TestPoliciesChooseTheirVictims(t *testing.T) { + for _, c := range []struct { + policy Policy + oldest, newest bool + name string + }{ + {FIFO, true, false, "FIFO deletes the oldest"}, + {LIFO, false, true, "LIFO deletes the youngest"}, + {Uniform, false, false, "Uniform deletes neither consistently"}, + } { + t.Run(c.name, func(t *testing.T) { + ops, err := Cycle(2000, Config{Seed: 11, Ratio: 2, Victims: c.policy}) + if err != nil { + t.Fatalf("unexpected error: %v", err) + } + var live []uint32 // transient elements present, oldest first + oldest, newest, deletes := 0, 0, 0 + for _, op := range ops { + if op.Kind == Insert { + live = append(live, op.ID) + continue + } + deletes++ + i := slices.Index(live, op.ID) + if i == 0 { + oldest++ + } + if i == len(live)-1 { + newest++ + } + live = slices.Delete(live, i, i+1) + } + if got := oldest == deletes; got != c.oldest { + t.Errorf("%d of %d deletions removed the oldest element", oldest, deletes) + } + if got := newest == deletes; got != c.newest { + t.Errorf("%d of %d deletions removed the youngest element", newest, deletes) + } + }) + } +} + +// TestCycleKeepsAboutLiveTargetPresent checks the steady state a cycle is +// meant to hold: the number of transient elements present settles around +// LiveTarget instead of growing with the length of the cycle. Over the middle +// of a long cycle, the median count has to lie within a factor of two of it. +func TestCycleKeepsAboutLiveTargetPresent(t *testing.T) { + const target, live = 10000, 500 + ops, err := Cycle(target, Config{Seed: 5, Ratio: 3, LiveTarget: live}) + if err != nil { + t.Fatalf("unexpected error: %v", err) + } + present, samples := 0, []int{} + for i, op := range ops { + if op.Kind == Insert { + present++ + } else { + present-- + } + if i > len(ops)/4 && i < 3*len(ops)/4 { + samples = append(samples, present) + } + } + slices.Sort(samples) + if median := samples[len(samples)/2]; median < live/2 || median > 2*live { + t.Errorf("median of %d transient elements present, want about %d", median, live) + } +} + +// TestConfigIsChecked checks that the generator refuses configurations that +// would produce something other than what was asked for, and names the +// problem. +func TestConfigIsChecked(t *testing.T) { + cases := []struct { + name string + build bool + n int + c Config + want string + }{ + {"returns error for an empty target", true, 0, Config{}, "target"}, + {"returns error for a ratio below one", true, 10, Config{Ratio: 0.5}, "Ratio"}, + {"returns error for a NaN ratio", true, 10, Config{Ratio: math.NaN()}, "Ratio"}, + {"returns error for a cycle at ratio one", false, 10, Config{Ratio: 1}, "Ratio"}, + {"returns error for a cycle without transients", false, 10, Config{Ratio: 1.01}, "no transient"}, + {"returns error for a negative burst", true, 10, Config{MaxBurst: -1}, "MaxBurst"}, + {"returns error for an unknown policy", true, 10, Config{Victims: Policy(7)}, "Victims"}, + {"returns error for more IDs than fit in uint32", false, 3_000_000_000, Config{Ratio: 2}, "uint32"}, + } + for _, c := range cases { + t.Run(c.name, func(t *testing.T) { + var err error + if c.build { + _, err = Build(c.n, c.c) + } else { + _, err = Cycle(c.n, c.c) + } + if err == nil || !strings.Contains(err.Error(), c.want) { + t.Errorf("got %v, want an error mentioning %q", err, c.want) + } + }) + } +} + +// TestCheckRejectsInvalidStreams checks the verifier itself, which tests of +// anything that produces a stream rely on: it has to catch every kind of +// invalid stream, name the offending operation, and accept a valid one. +func TestCheckRejectsInvalidStreams(t *testing.T) { + cases := []struct { + name string + ops []Op + start, end []uint32 + want string + }{ + {"accepts a valid stream", []Op{{5, Insert}, {1, Delete}}, []uint32{1}, []uint32{5}, ""}, + {"rejects inserting a present element", []Op{{1, Insert}}, []uint32{1}, []uint32{1}, "op 0 inserts element 1"}, + {"rejects deleting an absent element", []Op{{2, Insert}, {3, Delete}}, nil, []uint32{2}, "op 1 deletes element 3"}, + {"rejects an unknown kind", []Op{{1, Kind(9)}}, nil, nil, "Kind(9)"}, + {"rejects a missing element at the end", nil, []uint32{1}, []uint32{1, 2}, "element 2 should be present"}, + {"rejects an extra element at the end", []Op{{4, Insert}}, nil, nil, "first being 4"}, + {"rejects a duplicate start element", nil, []uint32{1, 1}, nil, "twice"}, + } + for _, c := range cases { + t.Run(c.name, func(t *testing.T) { + err := Check(c.ops, c.start, c.end) + if c.want == "" { + if err != nil { + t.Errorf("unexpected error: %v", err) + } + return + } + if err == nil || !strings.Contains(err.Error(), c.want) { + t.Errorf("got %v, want an error mentioning %q", err, c.want) + } + }) + } +} + +// model is a set that fails the test on any invalid operation, for replaying +// streams through a Cursor. +type model struct { + t *testing.T + present map[uint32]bool +} + +func newModel(t *testing.T, start []uint32) *model { + m := &model{t: t, present: map[uint32]bool{}} + for _, id := range start { + m.present[id] = true + } + return m +} + +func (m *model) apply(run []Op) { + for _, op := range run { + if (op.Kind == Insert) == m.present[op.ID] { + m.t.Fatalf("invalid %s of element %d", op.Kind, op.ID) + } + if op.Kind == Insert { + m.present[op.ID] = true + } else { + delete(m.present, op.ID) + } + } +} + +// TestCursorWrapsAroundTheCycle checks the replay a benchmark runs: batches of +// arbitrary size, smaller and larger than the cycle, have to continue exactly +// where the last one stopped and wrap around at the end, so that every +// operation stays valid, and Settle has to bring the structure back to the +// state the cycle starts from. +func TestCursorWrapsAroundTheCycle(t *testing.T) { + const target = 300 + ops, err := Cycle(target, Config{Seed: 9}) + if err != nil { + t.Fatalf("unexpected error: %v", err) + } + m := newModel(t, ids(target)) + cur := &Cursor{} + batch := cur.Batch(ops, m.apply) + applied := 0 + for _, n := range []uint64{1, 7, uint64(len(ops)) - 3, 0, uint64(2*len(ops) + 5), 13, uint64(len(ops))} { + batch(n) + applied += int(n) + if want := applied % len(ops); cur.Position() != want { + t.Fatalf("after %d operations the cursor is at %d, want %d", applied, cur.Position(), want) + } + } + cur.Settle(ops, m.apply) + if cur.Position() != 0 || len(m.present) != target { + t.Errorf("after Settle: position %d, %d elements present, want 0 and %d", cur.Position(), len(m.present), target) + } + for _, id := range ids(target) { + if !m.present[id] { + t.Fatalf("element %d missing after Settle", id) + } + } + cur.Settle(ops, func([]Op) { t.Error("Settle at the start of the cycle should do nothing") }) +} + +// TestCursorIgnoresAnEmptyStream checks that a batch over an empty stream does +// nothing rather than looping forever. +func TestCursorIgnoresAnEmptyStream(t *testing.T) { + cur := &Cursor{} + cur.Advance(nil, 10, func([]Op) { t.Error("apply called for an empty stream") }) + cur.Settle(nil, func([]Op) { t.Error("apply called for an empty stream") }) +} + +// TestStrings checks the names that appear in error messages and output. +func TestStrings(t *testing.T) { + for got, want := range map[string]string{ + Insert.String(): "insert", Delete.String(): "delete", Kind(4).String(): "Kind(4)", + Uniform.String(): "Uniform", FIFO.String(): "FIFO", LIFO.String(): "LIFO", Policy(9).String(): "Policy(9)", + } { + if got != want { + t.Errorf("got %q, want %q", got, want) + } + } +} + +// FuzzStreamsAndReplay checks the promises of the generator and the cursor on +// arbitrary configurations and batch sizes: whatever the seed, size, ratio, +// burst, live target and policy, a Build and a Cycle have to pass Check, and +// replaying the cycle in batches of the fuzzed sizes has to stay valid and +// settle back to the start. +func FuzzStreamsAndReplay(f *testing.F) { + f.Add(uint64(1), uint16(100), uint8(100), uint8(16), uint16(0), uint8(0), uint16(7), uint16(250)) + f.Add(uint64(0), uint16(1), uint8(200), uint8(1), uint16(1), uint8(1), uint16(1), uint16(1)) + f.Fuzz(func(t *testing.T, seed uint64, target uint16, extra, burst uint8, live uint16, policy uint8, n1, n2 uint16) { + c := Config{ + Seed: seed, + Ratio: 1 + float64(extra)/100, + MaxBurst: int(burst), + LiveTarget: int(live), + Victims: Policy(policy % 3), + } + size := int(target%2000) + 1 + ops, err := Build(size, c) + if err != nil { + t.Fatalf("Build: %v", err) + } + if err := Check(ops, nil, ids(size)); err != nil { + t.Fatalf("Build: %v", err) + } + cycle, err := Cycle(size, c) + if err != nil { + return // too few transients for a cycle at this ratio + } + if err := Check(cycle, ids(size), ids(size)); err != nil { + t.Fatalf("Cycle: %v", err) + } + m := newModel(t, ids(size)) + cur := &Cursor{} + cur.Advance(cycle, uint64(n1), m.apply) + cur.Advance(cycle, uint64(n2), m.apply) + cur.Settle(cycle, m.apply) + if len(m.present) != size { + t.Fatalf("%d elements present after Settle, want %d", len(m.present), size) + } + }) +}