Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
95 changes: 95 additions & 0 deletions HOWTO.md
Original file line number Diff line number Diff line change
Expand Up @@ -133,6 +133,101 @@ cost gets measured along with your code. See
[Attenuation](#attenuation-your-number-is-real-but-smaller-than-the-truth)
for what this costs you.

## Benchmarking insertions and deletions

If your candidates are data structures and the question is how fast they
*change* — insert, delete, grow, shrink — the obvious batch is a trap:

```go
Batch: func(n uint64) {
for i := range n {
m[key(i)] = value // insert
delete(m, key(i)) // and take it out again, so the map stays the same
}
}
```

That measures a structure that never changes shape. The element goes into
the same slot it just left; nothing splits, merges, resizes, rehashes or
leaves a tombstone, and those are exactly the costs that separate one data
structure from another in real use. The honest alternatives have their own
traps: picking what to delete inside the timed loop adds work to both
candidates and dilutes the difference; a stream that deletes elements that
aren't there times no-ops; a structure that keeps growing over the run is a
different structure at the last sample than at the first.

The `workload` package does all of this. You describe each data structure by
how to create an empty one and how to apply an operation to it:

```go
goMap := workload.Structure[map[uint64]struct{}]{
Name: "map",
New: func() map[uint64]struct{} { return make(map[uint64]struct{}, 100_000) },
Apply: func(m map[uint64]struct{}, run []workload.Op) {
for _, op := range run {
if op.Kind == workload.Insert {
m[uint64(op.ID)] = struct{}{}
} else {
delete(m, uint64(op.ID))
}
}
},
}
res, err := workload.Compare(100_000, goMap, otherSet, workload.Options{})
fmt.Println(res)
```

The IDs are abstract; map them to your own keys and values, for example
through a precomputed slice of keys. You get **two answers**, because there are
two different questions, and mixing them into one number would get both
wrong:

- **Steady state:** what does one insertion or deletion cost in a structure
that has been in use for a while? Both structures are filled to 100,000
elements, then a *cycle* is replayed on them: bursts of insertions and
deletions of 100,000 further, transient elements, about 12,500 of them present
at a time, ending exactly where it started, so it can repeat endlessly.
- **Build:** what does it cost to build such a structure from empty, with the
same kind of back-and-forth along the way? Every sample is one complete
build of a fresh structure.

The reason for the split: the first pass of a cycle is where a structure grows
to the largest size the cycle reaches. For a Go map that pass cost up to seven
times as much per operation as every pass after it. Timing it along with the
steady state would spread a one-time cost over the measurement, in a
proportion that depends on how long the run was. Dropping it would favour
structures whose growth is expensive. So it is measured where it belongs, in
the build, where creating the structure with or without a capacity hint, every
resize and the garbage all count, and the steady state starts after one
untimed cycle. If the two answers disagree, that is the result: one structure
can be faster to use and slower to build.

What `workload.Compare` takes care of, so that you don't have to:

- Every operation is valid by construction: nothing is inserted twice and
nothing is deleted that isn't there.
- The streams are precomputed; the timed loop only reads the next operation.
- Both structures are built in alternating chunks, so neither is the one built
last.
- The first pass of the cycle is untimed, and the position in the cycle stays
with each structure across batches.
- Whole builds take milliseconds, so the build comparison uses its own, smaller
defaults (31 repeats, 10 validation runs, about 1,300 builds in all) and
collects garbage between batches, so that one build's garbage is not
collected in the middle of the next.

`workload.Config` tunes the streams: `Ratio` (insertions per element at rest,
default 2), burst length, and `Victims`, which chooses what a deletion removes:
`Uniform` (the default, like a general-purpose map), `FIFO` (a queue or a
retention window), or `LIFO` (a stack or undo log).

For setups `Compare` does not cover, the parts are available on their own:
`workload.Cycle` and `workload.Build` make the streams, `workload.Replay`
replays a cycle on one structure (including the untimed first pass and
`Settle`, which returns the structure to its start state), and
`workload.Check` replays any stream against a model and reports the first
invalid operation.

## The long version: what `Compare` does, step by step

This is the sequence `Compare` runs automatically. Read it if you want to
Expand Down
8 changes: 8 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -24,6 +24,7 @@ Keywords: benchmarking, performance, bootstrap, runtime comparison, statistics,
- Detect a trend across a measurement run, which resampling structurally cannot see because it discards the order the samples arrived in.
- Resample in blocks when the measurements are correlated enough that treating them as independent would overstate confidence.
- Deterministic PRNG for reproducible inputs, and a crypto/rand-backed one where unpredictability is wanted.
- Compare mutable data structures under realistic insert/delete workloads in one call (`workload`), with separate answers for the steady state and for building to size; the streams are valid by construction, reproducible from a seed, and cyclic, so that a batch can replay them endlessly without rebuilding the structure.

## What this cannot tell you

Expand Down Expand Up @@ -162,6 +163,13 @@ Primitives:
- `SampleTime()` / `DiffTimeStamps()` — high-resolution timestamps, and `GetSampleTimePrecision()` for the smallest interval they can resolve here.
- `Median` / `QuickMedian` / `Statistics` — small statistics helpers.

Workloads for mutable data structures (`github.com/TomTonic/rtcompare/workload`):

- `workload.Compare(target, a, b, Options)` — the whole job in one call: each `Structure` says how to create an empty structure and apply operations to it, and you get two reports, one for the steady state (per insertion or deletion, after one untimed cycle) and one for building from empty (per whole build, growth included).
- `workload.Cycle(target, Config)` / `workload.Build(target, Config)` — the streams: a cycle of insertions and deletions that ends where it started, and a build from empty with a realistic history.
- `workload.Replay` — replays a cycle on one structure instance, with the untimed first pass, and `Settle` to return it to its start state. `workload.Cursor` is the bare position, for doing it by hand.
- `workload.Check(ops, start, end)` — replays a stream against a model and reports the first invalid operation.

A note on the threshold of `0.0`: every threshold is evaluated as `delta >= t`, so at zero the question is "at least as fast", not "faster". Quantized timings tie often, and every tie counts towards it. Ask for a threshold above zero if you mean strictly faster.

Note on negative `relativeGains`: Negative thresholds are allowed and are
Expand Down
2 changes: 1 addition & 1 deletion go.mod
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
module github.com/TomTonic/rtcompare

go 1.26.0
go 1.26.8

require (
github.com/TomTonic/Set3 v0.4.2
Expand Down
229 changes: 229 additions & 0 deletions workload/compare.go
Original file line number Diff line number Diff line change
@@ -0,0 +1,229 @@
package workload

import (
"fmt"
"runtime"
"strings"

"github.com/TomTonic/rtcompare"
)

// Structure describes one data structure under test for [Compare]: how to
// create an empty one, and how to apply operations to it.
type Structure[S any] struct {
// Name labels the structure in the results.
Name string

// New returns an empty structure, created the way the program would create
// it. Whether it is given a capacity hint is part of what the build
// comparison measures, so give it one exactly when the program would.
New func() S

// Apply performs the operations of run on s, in order. Keep the loop over
// run in this function, where the compiler can inline the structure's
// methods.
Apply func(s S, run []Op)
}

// check rejects a structure that cannot be used.
func (s Structure[S]) check(position string) error {
if s.New == nil || s.Apply == nil {
return fmt.Errorf("workload: structure %s (%q) needs both New and Apply", position, s.Name)
}
return nil
}

// Options configures [Compare]. The zero value is usable and selects the
// documented defaults.
type Options struct {
// Config configures the two streams, see [Config].
Config Config

// SteadyState are the options of the steady-state comparison, passed to
// [rtcompare.Compare] unchanged.
SteadyState rtcompare.CompareOptions

// Build are the options of the build comparison. Each of its samples is a
// whole build, which takes milliseconds rather than the microseconds of a
// calibrated batch, so a zero Collect.Repeats selects [BuildRepeats] and a
// zero ValidationRuns selects [BuildValidationRuns] rather than
// rtcompare's defaults, and Collect.GCBetween is always set, since every
// build leaves a whole structure of garbage behind that would otherwise be
// collected in the middle of the next candidate's build.
Build rtcompare.CompareOptions

// SkipBuild omits the build comparison. The steady-state result alone
// ignores what growing to size costs, and so favours structures whose
// growth is expensive; set it only when that cost is known not to matter.
SkipBuild bool
}

// BuildRepeats and BuildValidationRuns are the build comparison's defaults for
// Repeats and ValidationRuns. They are smaller than rtcompare's, because a
// sample is a whole build: with them a build comparison performs about 1,300
// builds, some ten seconds for 100,000 elements, where rtcompare's defaults
// would take several minutes.
const (
BuildRepeats = 31
BuildValidationRuns = 10
)

// Result holds the two answers [Compare] gives, which are answers to two
// different questions and are deliberately not combined into one number.
type Result struct {
// Target is the number of elements at rest, and CycleOps and BuildOps the
// lengths of the two streams.
Target int
CycleOps, BuildOps int

// SteadyState compares the cost of one insertion or deletion in a
// structure that has been in use for a while, averaged over the cycle. Its
// per-operation figures are nanoseconds per insertion or deletion.
SteadyState rtcompare.Report

// Build compares the cost of building a structure from empty to Target
// elements, with the transient insertions and deletions a real build sees,
// including creation, every resize and the garbage it leaves. Its
// per-operation figures are nanoseconds per whole build. It is the zero
// Report when Options.SkipBuild was set.
Build rtcompare.Report
}

// String renders both answers, each under the question it answers.
func (r Result) String() string {
var b strings.Builder
fmt.Fprintf(&b, "steady state, per insertion or deletion in a structure of %d elements (cycle of %d operations, first pass untimed):\n", r.Target, r.CycleOps)
b.WriteString(indent(r.SteadyState.String()))
if r.Build.SamplesA != nil {
fmt.Fprintf(&b, "\n\nbuild from empty to %d elements, per whole build (%d operations):\n", r.Target, r.BuildOps)
b.WriteString(indent(r.Build.String()))
}
return b.String()
}

func indent(s string) string {
return " " + strings.ReplaceAll(s, "\n", "\n ")
}

// Compare compares two data structures under a realistic mix of insertions
// and deletions, and answers two questions separately: what an operation costs
// once a structure is in use, and what it costs to build one.
//
// Parameters: target is the number of elements the structures hold at rest,
// at least one; a and b are the two structures, see [Structure]; opt
// configures the streams and both comparisons, see [Options].
//
// It returns both reports, or an error if a structure lacks New or Apply, if
// the configuration is invalid (see [Build] and [Cycle]), or if a comparison
// fails.
//
// Use it for any comparison of mutable data structures. It exists because
// doing this by hand has several traps, and each of them quietly changes the
// answer:
//
// - The steady-state comparison builds one structure of each kind to target
// elements with a [Build] stream, applied to the two alternately in chunks
// so that neither is built last, and then replays a [Cycle] on each
// through a [Replay]. The first pass of the cycle is untimed: it is where
// the structure grows to the cycle's peak size, a one-time cost that would
// otherwise be spread over the measurement in proportions that depend on
// how long the run was.
// - That growth is not dropped, it is measured where it belongs. The build
// comparison times whole builds, each from a structure fresh from New to
// target elements, so creation, capacity hints, every resize and the
// garbage are included. A structure that cannot be presized and has to
// grow pays for it here, and one that grows cheaply shows it here.
//
// The two answers can disagree, and then that is the result: one structure
// can be faster to use and slower to build. Read both.
//
// The cost is that of two [rtcompare.Compare] calls, the build one dominated
// by the builds it performs, about 1,300 at the defaults; see
// [BuildRepeats].
//
// res, err := workload.Compare(100_000,
// workload.Structure[map[uint64]struct{}]{Name: "map", New: newMap, Apply: applyToMap},
// workload.Structure[*set3.Set3[uint64]]{Name: "Set3", New: newSet3, Apply: applyToSet3},
// workload.Options{})
// if err != nil { ... }
// fmt.Println(res)
func Compare[SA, SB any](target int, a Structure[SA], b Structure[SB], opt Options) (Result, error) {
if err := a.check("A"); err != nil {
return Result{}, err
}
if err := b.check("B"); err != nil {
return Result{}, err
}
cycle, err := Cycle(target, opt.Config)
if err != nil {
return Result{}, err
}
build, err := Build(target, opt.Config)
if err != nil {
return Result{}, err
}
res := Result{Target: target, CycleOps: len(cycle), BuildOps: len(build)}

res.SteadyState, err = compareSteadyState(a, b, build, cycle, opt.SteadyState)
if err != nil {
return res, fmt.Errorf("workload: steady-state comparison: %w", err)
}
if !opt.SkipBuild {
res.Build, err = rtcompare.Compare(buildCandidate(a, build), buildCandidate(b, build), buildOptions(opt.Build))
if err != nil {
return res, fmt.Errorf("workload: build comparison: %w", err)
}
}
return res, nil
}

// compareSteadyState builds one structure of each kind, alternately, and
// compares the cycle replayed on each.
func compareSteadyState[SA, SB any](a Structure[SA], b Structure[SB], build, cycle []Op, opt rtcompare.CompareOptions) (rtcompare.Report, error) {
sa, sb := a.New(), b.New()
applyA := func(run []Op) { a.Apply(sa, run) }
applyB := func(run []Op) { b.Apply(sb, run) }
buildAlternately(build, applyA, applyB)
ra, rb := NewReplay(cycle, applyA), NewReplay(cycle, applyB)
return rtcompare.Compare(ra.Candidate(a.Name), rb.Candidate(b.Name), opt)
}

// alternateChunk is how many operations of the build stream go to one
// structure before the other gets its turn.
const alternateChunk = 256

// buildAlternately applies the stream to both structures in alternating
// chunks. Built one after the other, the structure built second would be the
// one in the caches and in fresher memory when the measurement starts, which
// has been measured to make an identical structure a few percent faster.
func buildAlternately(ops []Op, applyA, applyB func([]Op)) {
for start := 0; start < len(ops); start += alternateChunk {
run := ops[start:min(start+alternateChunk, len(ops))]
applyA(run)
applyB(run)
}
}

// buildCandidate times whole builds: each operation of its batch creates a
// structure and applies the entire build stream to it.
func buildCandidate[S any](s Structure[S], build []Op) rtcompare.Candidate {
return rtcompare.Candidate{Name: s.Name, Batch: func(n uint64) {
for range n {
fresh := s.New()
s.Apply(fresh, build)
runtime.KeepAlive(fresh)
}
}}
}

// buildOptions fills in the build comparison's own defaults.
func buildOptions(opt rtcompare.CompareOptions) rtcompare.CompareOptions {
if opt.Collect.Repeats == 0 {
opt.Collect.Repeats = BuildRepeats
}
if opt.ValidationRuns == 0 {
opt.ValidationRuns = BuildValidationRuns
}
opt.Collect.GCBetween = true
return opt
}
Loading
Loading