From 135207a57427dc1ca13f49e691427ad2422cbfab Mon Sep 17 00:00:00 2001 From: Tom Tonic Date: Sun, 27 Sep 2026 19:04:19 +0200 Subject: [PATCH 1/3] Update Go version to 1.26.8 in go.mod --- go.mod | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/go.mod b/go.mod index 740d1df..402711d 100644 --- a/go.mod +++ b/go.mod @@ -1,6 +1,6 @@ module github.com/TomTonic/rtcompare -go 1.26.0 +go 1.26.8 require ( github.com/TomTonic/Set3 v0.4.2 From 47ba095457f32cffce96c0651a813e6ab7a07da7 Mon Sep 17 00:00:00 2001 From: Tom Tonic Date: Sun, 27 Sep 2026 19:56:00 +0200 Subject: [PATCH 2/3] Stop Compare from favouring the candidate that ran alone last Compare validated candidate A, then candidate B, and then measured both after a single warm-up pair. B therefore entered the measurement with the last-level cache full of its own data. With a working set near the size of that cache and short calibrated batches, the head start survived into the medians: identical code measured up to 66% apart, reported as resolved (#111). - ValidatePair validates two candidates together, interleaved batch by batch in the order A, B, B', A' and back. In that order each half of each A/A experiment has the same history. Compare uses it. - Collect warms up for at least CollectOptions.WarmupDuration (DefaultWarmupDuration, 300 ms) in addition to the batch count, and alternates the order of the candidates while it does. A validation pays for the duration once, not once per run. - Report.DriftRatio tests the ratio B/A of neighbouring batches for a trend. A warning appears when the trend is significant and larger than the result's resolution. - HOWTO and README describe the pitfall and the checks against it. Measured with cmd/rtcompare-aa, which follows in the next commit, on a Ryzen 9 7900 (32 MB L3): two identical 16 MB pointer chases, 5 processes in both role assignments, Repeats 41, ValidationRuns 2. The pooled delta went from -29.4% to +0.35%. With the defaults it went from -0.5% to +0.34%, and the in-cache control stayed within 0.1%. The difference that remains follows the build order of the two fixtures and not the roles. A 2 s warm-up does not change it, so it is memory layout (#109). Compare now costs 0.6 s more, or 0.3 s with SkipValidation. Co-Authored-By: Claude Opus 5.5 --- HOWTO.md | 98 ++++++++++++++- README.md | 3 +- collect.go | 261 +++++++++++++++++++++++++++++---------- collect_test.go | 82 ++++++++++++- compare.go | 95 +++++++++++---- compare_test.go | 69 ++++++++++- validate.go | 311 ++++++++++++++++++++++++++++++++++------------- validate_test.go | 169 +++++++++++++++++++++++++ 8 files changed, 908 insertions(+), 180 deletions(-) diff --git a/HOWTO.md b/HOWTO.md index f44accd..d25ac8b 100644 --- a/HOWTO.md +++ b/HOWTO.md @@ -171,7 +171,7 @@ duration — see [Troubleshooting](#troubleshooting-what-to-do-when). ### Step 2 — Find out what your own machine invents from nothing -**Function:** `ValidateHarness` +**Function:** `ValidatePair` (or `ValidateHarness` for a single candidate) This is the step every other benchmarking approach skips, and it's the reason rtcompare exists. @@ -199,6 +199,17 @@ collection more than the other — and your comparison is only as trustworthy as the *worse* of the two. `Compare` does this automatically and uses the higher (worse) of the two floors. +**And do it for both *together*,** with `ValidatePair(a, b, ...)`, rather than +calling `ValidateHarness(a)` and then `ValidateHarness(b)`. Validating one +candidate on its own means running it alone for a while, and whatever runs +alone last leaves the CPU's caches full of its own data. If you validate A and +then B and then compare them, B starts the comparison warm and A starts cold. +See [Whatever ran last starts warm](#whatever-ran-last-starts-warm) for how +much that can matter. `ValidatePair` interleaves the two validations batch by +batch, so that neither candidate ever has the machine to itself, and the noise +floors are measured with the other candidate competing for the caches, just +as it will in the real comparison. + **What you get back**, and what each number tells you: - **`NoiseFloor`** — the practical floor. If your measured difference is @@ -227,6 +238,12 @@ against a machine that's changing over time — if it's getting slower as the run progresses, both candidates are equally exposed to that instead of whichever one happens to run later. +Before the first measured batch, `Collect` warms both candidates up, again +alternating between them, for at least `CollectOptions.WarmupDuration` (300 ms +by default). That is not about the one-time costs you might expect a warm-up +for, such as page faults; a single batch would cover those. It is about the +caches, see [below](#whatever-ran-last-starts-warm). + This gives you two lists of numbers — `SamplesA` and `SamplesB` — one measurement per batch. Everything from here on works from these two lists. @@ -244,7 +261,11 @@ bootstrap is structurally blind to it, because it discards the order the numbers arrived in. `DetectDrift` looks specifically for that trend, on each candidate's series -separately. If it finds one, it's worth knowing even though ABBA ordering +separately. `Compare` also runs it on the ratio B/A of each pair of +neighbouring batches (`Report.DriftRatio`), which catches something the two +series on their own cannot: one candidate that started with an advantage and +loses it as the run goes on. A machine that slows down slows both candidates +and leaves the ratio flat; a head start moves the ratio and nothing else. If it finds one, it's worth knowing even though ABBA ordering already protects the *comparison* from being biased by it — a real trend means your measurements are less independent of each other than the statistics assume, which affects how much you should trust *any* interval or @@ -375,6 +396,63 @@ honestly — as the cost of that whole region, not as an isolated number for your function alone. Read a result as "how much faster is this measured region," not "how much faster is this one function in isolation." +## Whatever ran last starts warm + +Your CPU keeps recently used memory in its caches, and the last-level cache is +the one that matters here: tens of megabytes, shared by everything. Whatever +touched its data last, for long enough, owns most of that cache. That can be +one of your candidates, for reasons that have nothing to do with its speed: + +- it was **built** last — you set up A's data structure, then B's; +- it was **validated** last, if you called `ValidateHarness` once per + candidate; +- it was **calibrated** last, since calibration runs each candidate on its own. + +The candidate that ran last then starts the measurement warm and the other +cold. When both candidates' data fit in the cache together, this evens out +within a batch or two and does not matter. When each one's data is about the +size of the last-level cache, it matters a lot. Short, calibrated batches may +not touch enough memory to turn the cache over, so the head start can survive +into the medians. Two *identical* 16 MB pointer-chasing structures, compared +on a machine with 32 MB of L3, came out 5 to 50% apart, always in favour of +the one that ran alone last, with narrow intervals that reported the +difference as real. (That is [issue #111](https://github.com/TomTonic/rtcompare/issues/111); +`cmd/rtcompare-aa` reproduces it.) + +rtcompare deals with this in three ways, and `Compare` uses all of them: + +1. `ValidatePair` validates both candidates together, so validation leaves + neither of them ahead. +2. `Collect` warms both candidates up, alternately, for at least + `CollectOptions.WarmupDuration`, 300 ms by default. In the experiment + above, about 250 ms were enough to erase the head start. This is what a + comparison now costs extra: 0.3 s per `Collect` and once more for the + validation, 0.6 s per `Compare`. For small data that stays in the cache you + can lower it, and `time.Nanosecond` warms up by batch count alone. +3. `Compare` checks the ratio B/A for a trend across the run + (`Report.DriftRatio`) and warns when it is larger than the result's + resolution — the sign that the candidates had not settled when the + measurement began. If you see that warning, raise `WarmupDuration`. + +What none of this can remove is a difference in **where** the data landed in +memory. Two structures with identical content, built one after the other, are +not laid out identically: different addresses, different pages, different +cache sets. In the experiment above, after the warm-up, the structure built +second was still 3 to 5% faster, whichever role it played, and a warm-up of +2 s instead of 0.3 s did not change that. That is a real difference between the +two data structures as they sit in memory, not a measurement artefact, and it +is why comparing large data in a single process is not enough; see +[issue #109](https://github.com/TomTonic/rtcompare/issues/109). + +The practical checks, for any comparison of large data: + +- **Swap the roles.** Compare (A, B) and then (B, A). A bias towards a role + flips the sign of the delta when you swap; a real difference does not. +- **Swap the build order.** Build B's data first and A's second. If the + result moves with the build order, you are measuring layout. +- **Do not validate candidates one at a time** before comparing them. Use + `ValidatePair`, or `Compare`, which does. + ## Troubleshooting: what to do, when This is the part of the "long protocol" that's normally invisible — the @@ -427,6 +505,22 @@ mid-batch. move on, or find a quieter machine (a dedicated benchmark server, a CI runner with less contention) if this comparison matters enough. +### A warning says the candidates had not reached a steady state + +**Symptom:** `Report.Warnings` says that "the ratio B/A shifted" during the +run. + +**What it means:** one candidate was ahead at the start of the run and lost +that advantage as it went on, usually because it started with the caches full +of its own data. The median of the run then depends on how long the run was, +and the difference it reports is partly an artefact. See +[Whatever ran last starts warm](#whatever-ran-last-starts-warm). + +**What to do:** raise `CollectOptions.WarmupDuration`, for example to one or +two seconds, and check that nothing runs one candidate alone between your +setup and the comparison. Then swap the roles of the two candidates and see +whether the result keeps its sign. + ### A drift warning appears **Symptom:** `DriftReport.Drifted(...)` returns true, or `Report.Warnings` diff --git a/README.md b/README.md index 01e4686..c4164d9 100644 --- a/README.md +++ b/README.md @@ -140,7 +140,7 @@ The individual steps, for when the summary is not enough: Measuring: - `Candidate` / `Batch` — one implementation under test, with optional `Setup` and `Teardown` that run outside the measured region. -- `Collect(a, b, CollectOptions)` — runs both candidates and returns one timing sample per repeat each, ready to hand to `CompareSamples`. Owns measurement order, warm-up, GC placement and batch sizing. +- `Collect(a, b, CollectOptions)` — runs both candidates and returns one timing sample per repeat each, ready to hand to `CompareSamples`. Owns measurement order, warm-up, GC placement and batch sizing. The warm-up alternates the candidates for at least `WarmupDuration` (300 ms by default), so that neither starts the measurement with the caches to itself. - `CalibrateInnerLoops(candidate, CalibrationOptions)` — sizes batches for a target quantization error. Called automatically when `CollectOptions.InnerLoops` is left at zero. Judging: @@ -154,6 +154,7 @@ Judging: Checking the measurement itself: - `ValidateHarness(candidate, ValidationOptions)` — runs a candidate against itself and reports the noise floor, the tie rate, the drift rate and the autocorrelation. `Resolves(difference)` answers whether a result clears that floor. The floor is the 90th percentile of the differences observed on identical code, not their maximum, so that it converges as you validate longer instead of growing; roughly one A/A run in ten exceeds it. Validate both candidates and use the worse floor. +- `ValidatePair(a, b, ValidationOptions)` — validates two candidates together, interleaved batch by batch, and returns one `HarnessValidation` each. Use it instead of two `ValidateHarness` calls before comparing the two: a candidate validated on its own last would start the comparison with a warm cache. - `DetectDrift(samples)` — tests a sample series for a trend across the run. Primitives: diff --git a/collect.go b/collect.go index 8ddfc19..113b593 100644 --- a/collect.go +++ b/collect.go @@ -5,6 +5,7 @@ import ( "math" "runtime" "runtime/debug" + "time" ) // Batch runs the code under test exactly n times. @@ -148,10 +149,34 @@ func (o Order) String() string { // resolution. See CollectOptions.MaxQuantizationError. const DefaultRepeats = 101 -// DefaultWarmup is the number of unmeasured batches [Collect] runs per candidate -// before collecting samples, when CollectOptions.Warmup is left at zero. +// DefaultWarmup is the least number of unmeasured batches [Collect] runs per +// candidate before collecting samples, when CollectOptions.Warmup is left at +// zero. The warm-up also runs for at least [DefaultWarmupDuration], which is the +// bound that usually decides how long it takes. const DefaultWarmup = 1 +// DefaultWarmupDuration is the least wall-clock time [Collect] spends warming +// up both candidates, together, before it collects samples, when +// CollectOptions.WarmupDuration is left at zero. +// +// A fixed number of batches is not a warm-up when the working set is about the +// size of the last-level cache. Whatever ran alone last before the measurement, +// be it an A/A validation, a calibration or the caller's own fixture build, +// leaves the cache full of its own data. Calibrated batches are short, often a +// fraction of a millisecond, and a whole run of them may touch too little memory +// to turn the cache over, so the head start survives into the medians. Two +// identical 16 MB pointer chases on a machine with 32 MB of L3 measured the +// candidate that ran alone last 5 to 50% faster after one warm-up pair, and +// within noise of each other after about 400 pairs of 0.3 ms batches, some +// 250 ms in all. See issue #111 and the reproduction in cmd/rtcompare-aa. +// +// Time is the bound rather than a batch count because what has to happen, the +// cache turning over, takes a roughly fixed amount of work, and batches are +// calibrated to a duration rather than to a volume. The cost is paid once per +// [Collect] and once per validation, not once per validation run; see +// [ValidateHarness]. +const DefaultWarmupDuration = 300 * time.Millisecond + // CollectOptions configures a [Collect] run. // // Every count in this struct is a Go int with the usual Go convention for @@ -216,16 +241,29 @@ type CollectOptions struct { // is [OrderABBA]. Order Order - // Warmup is the number of unmeasured batches run per candidate before sample - // collection starts, to fault in pages, grow stacks and train branch + // Warmup is the least number of unmeasured batches run per candidate before + // sample collection starts, to fault in pages, grow stacks and train branch // predictors and caches. Zero selects [DefaultWarmup]. To run no warm-up at // all, set SkipWarmup rather than passing a negative number. + // + // The warm-up continues until both this count and WarmupDuration have been + // reached, and it alternates the order of the candidates the way the + // measurement does, so that neither of them ends it with the caches to + // itself. Warmup int - // SkipWarmup disables warm-up entirely. This exists as its own field so that - // Warmup keeps a single unambiguous meaning; "no warm-up" is a mode, not a - // count. Measuring without warm-up means the first samples include one-time - // costs such as page faults and stack growth. + // WarmupDuration is the least wall-clock time the warm-up runs for, across + // both candidates together. Zero selects [DefaultWarmupDuration], whose + // documentation explains why a count alone is not enough. Negative values + // are rejected. To warm up by count alone, set it to [time.Nanosecond]. + WarmupDuration time.Duration + + // SkipWarmup disables warm-up entirely, both the count and the duration. + // This exists as its own field so that Warmup keeps a single unambiguous + // meaning; "no warm-up" is a mode, not a count. Measuring without warm-up + // means the first samples include one-time costs such as page faults and + // stack growth, and that whichever candidate ran alone last before Collect + // starts with a warm cache. SkipWarmup bool // GCBetween requests an explicit garbage collection before every batch, @@ -308,10 +346,18 @@ type CollectOptions struct { // affected all of them. Treat a result below roughly 1% as "not resolved" rather // than as a small but real effect. // -// Collect returns an error if either candidate has a nil Batch, if Repeats or -// Warmup is negative, if Repeats is below [MinimumDataPoints], if Order is not -// one of the defined constants, or if automatic calibration of InnerLoops fails. -// It does not otherwise inspect the collected samples. +// A note on what ran before: the warm-up alternates the candidates for at least +// CollectOptions.WarmupDuration so that neither starts the measurement with the +// caches to itself, whichever of them the caller built, validated or calibrated +// last. See [DefaultWarmupDuration] for why a count of batches is not enough, +// and [ValidatePair] for validating both candidates without handing one of them +// that head start in the first place. +// +// Collect returns an error if either candidate has a nil Batch, if Repeats, +// Warmup or WarmupDuration is negative, if Repeats is below +// [MinimumDataPoints], if Order is not one of the defined constants, or if +// automatic calibration of InnerLoops fails. It does not otherwise inspect the +// collected samples. func Collect(a, b Candidate, opt CollectOptions) (samplesA, samplesB []float64, err error) { if a.Batch == nil { return nil, nil, fmt.Errorf("rtcompare: candidate %s has a nil Batch function", a.label("A")) @@ -319,99 +365,182 @@ func Collect(a, b Candidate, opt CollectOptions) (samplesA, samplesB []float64, if b.Batch == nil { return nil, nil, fmt.Errorf("rtcompare: candidate %s has a nil Batch function", b.label("B")) } + s, err := opt.schedule() + if err != nil { + return nil, nil, err + } + + if opt.DisableGC { + previous := debug.SetGCPercent(-1) + defer debug.SetGCPercent(previous) + } + + if s.innerLoops == 0 { + // DisableGC is already in effect for the whole call, so it is not + // repeated here; the rest of the conditions must match the real run. + s.innerLoops, err = calibratePair(a, b, CalibrationOptions{ + MaxQuantizationError: opt.MaxQuantizationError, + MaxInnerLoops: opt.MaxInnerLoops, + GCBetween: opt.GCBetween, + }) + if err != nil { + return nil, nil, err + } + } + + samples := measureInTurn([]Candidate{a, b}, s) + return samples[0], samples[1], nil +} + +// schedule is a CollectOptions with every default resolved and every check +// passed, so that the loops that run it have nothing left to decide. It exists +// because [Collect] and the A/A validations measure different sets of +// candidates under the same options and must not interpret them differently. +type schedule struct { + repeats int + innerLoops uint64 // zero until calibrated + order Order + gc bool + disableGC bool + warmupRounds int + warmupDuration time.Duration + seed uint64 +} + +// schedule checks the options and resolves their defaults. InnerLoops is +// passed through as it is, zero included, because calibrating it needs the +// candidates. +func (opt CollectOptions) schedule() (schedule, error) { if opt.Repeats < 0 { - return nil, nil, fmt.Errorf("rtcompare: Repeats must not be negative, got %d", opt.Repeats) + return schedule{}, fmt.Errorf("rtcompare: Repeats must not be negative, got %d", opt.Repeats) } if opt.Repeats == 0 { opt.Repeats = DefaultRepeats } if uint64(opt.Repeats) < MinimumDataPoints { - return nil, nil, fmt.Errorf("rtcompare: Repeats must be at least %d, got %d", MinimumDataPoints, opt.Repeats) + return schedule{}, fmt.Errorf("rtcompare: Repeats must be at least %d, got %d", MinimumDataPoints, opt.Repeats) } if opt.Warmup < 0 { - return nil, nil, fmt.Errorf("rtcompare: Warmup must not be negative, got %d; use SkipWarmup to disable warm-up", opt.Warmup) + return schedule{}, fmt.Errorf("rtcompare: Warmup must not be negative, got %d; use SkipWarmup to disable warm-up", opt.Warmup) + } + if opt.WarmupDuration < 0 { + return schedule{}, fmt.Errorf("rtcompare: WarmupDuration must not be negative, got %v; use SkipWarmup to disable warm-up", opt.WarmupDuration) } switch opt.Order { case OrderABBA, OrderRandom, OrderSequential: default: - return nil, nil, fmt.Errorf("rtcompare: unknown Order %d", int(opt.Order)) + return schedule{}, fmt.Errorf("rtcompare: unknown Order %d", int(opt.Order)) } - warmup := opt.Warmup - if warmup == 0 { - warmup = DefaultWarmup + s := schedule{ + repeats: opt.Repeats, + innerLoops: opt.InnerLoops, + order: opt.Order, + gc: opt.GCBetween, + disableGC: opt.DisableGC, + warmupRounds: opt.Warmup, + warmupDuration: opt.WarmupDuration, + seed: opt.Seed, + } + if s.warmupRounds == 0 { + s.warmupRounds = DefaultWarmup + } + if s.warmupDuration == 0 { + s.warmupDuration = DefaultWarmupDuration } if opt.SkipWarmup { - warmup = 0 + s.warmupRounds, s.warmupDuration = 0, 0 } + return s, nil +} - if opt.DisableGC { - previous := debug.SetGCPercent(-1) - defer debug.SetGCPercent(previous) +// calibratePair sizes the batches for two candidates and returns the larger of +// the two sizes. The cheaper candidate needs the larger batch, so the maximum +// leaves both at or above the target duration while keeping the operation count +// identical for both. +func calibratePair(a, b Candidate, opt CalibrationOptions) (uint64, error) { + calA, err := CalibrateInnerLoops(a, opt) + if err != nil { + return 0, fmt.Errorf("calibrating candidate %s: %w", a.label("A"), err) } - - if opt.InnerLoops == 0 { - // DisableGC is already in effect for the whole call, so it is not - // repeated here; the rest of the conditions must match the real run. - calOpt := CalibrationOptions{ - MaxQuantizationError: opt.MaxQuantizationError, - MaxInnerLoops: opt.MaxInnerLoops, - GCBetween: opt.GCBetween, - } - calA, err := CalibrateInnerLoops(a, calOpt) - if err != nil { - return nil, nil, fmt.Errorf("calibrating candidate %s: %w", a.label("A"), err) - } - calB, err := CalibrateInnerLoops(b, calOpt) - if err != nil { - return nil, nil, fmt.Errorf("calibrating candidate %s: %w", b.label("B"), err) - } - // The cheaper candidate needs the larger batch. Using the maximum for - // both keeps the operation count identical and leaves both batches at - // or above the target duration. - opt.InnerLoops = max(calA.InnerLoops, calB.InnerLoops) + calB, err := CalibrateInnerLoops(b, opt) + if err != nil { + return 0, fmt.Errorf("calibrating candidate %s: %w", b.label("B"), err) } + return max(calA.InnerLoops, calB.InnerLoops), nil +} - for range warmup { - runBatch(a, opt.InnerLoops, opt.GCBetween) - runBatch(b, opt.InnerLoops, opt.GCBetween) +// measureInTurn warms the candidates up and then measures them in rounds, each +// candidate once per round, and returns one sample series per candidate. +// +// It is the loop behind [Collect], generalised from two candidates to any +// number so that the A/A validations of two candidates can be interleaved +// rather than run one after the other; see [ValidatePair]. Each round goes +// through the candidates forwards or backwards, as the order decides. For two +// candidates that is exactly ABBA, Random or Sequential. For more it keeps the +// property that matters: under ABBA every candidate holds every position equally +// often, and none of them runs alone for longer than two batches. +func measureInTurn(cands []Candidate, s schedule) [][]float64 { + if s.disableGC { + previous := debug.SetGCPercent(-1) + defer debug.SetGCPercent(previous) } + warmUp(cands, s) + var rng DPRNG - if opt.Order == OrderRandom { - if opt.Seed == 0 { + if s.order == OrderRandom { + if s.seed == 0 { rng = NewDPRNG() } else { - rng = NewDPRNG(opt.Seed) + rng = NewDPRNG(s.seed) } } - samplesA = make([]float64, 0, opt.Repeats) - samplesB = make([]float64, 0, opt.Repeats) - - for i := range opt.Repeats { - aFirst := true - switch opt.Order { + samples := make([][]float64, len(cands)) + for i := range samples { + samples[i] = make([]float64, 0, s.repeats) + } + for round := range s.repeats { + forward := true + switch s.order { case OrderABBA: - aFirst = i%2 == 0 + forward = round%2 == 0 case OrderRandom: // Uint32N uses the high bits of the scrambled state, which mix // better than the low bit would. - aFirst = rng.Uint32N(2) == 0 + forward = rng.Uint32N(2) == 0 case OrderSequential: - aFirst = true + forward = true + } + for k := range cands { + j := turn(k, len(cands), forward) + samples[j] = append(samples[j], runBatch(cands[j], s.innerLoops, s.gc)) } + } + return samples +} - if aFirst { - samplesA = append(samplesA, runBatch(a, opt.InnerLoops, opt.GCBetween)) - samplesB = append(samplesB, runBatch(b, opt.InnerLoops, opt.GCBetween)) - } else { - samplesB = append(samplesB, runBatch(b, opt.InnerLoops, opt.GCBetween)) - samplesA = append(samplesA, runBatch(a, opt.InnerLoops, opt.GCBetween)) +// warmUp runs unmeasured rounds until both the round count and the duration of +// the schedule have been reached. It reverses the order every round, like +// OrderABBA, because a warm-up that always ends on the same candidate hands that +// candidate the caches. +func warmUp(cands []Candidate, s schedule) { + start := time.Now() + for round := 0; round < s.warmupRounds || time.Since(start) < s.warmupDuration; round++ { + for k := range cands { + runBatch(cands[turn(k, len(cands), round%2 == 0)], s.innerLoops, s.gc) } } +} - return samplesA, samplesB, nil +// turn returns the index of the k-th candidate to run in a round of n, going +// forwards or backwards. +func turn(k, n int, forward bool) int { + if forward { + return k + } + return n - 1 - k } // timeBatch runs one candidate's lifecycle for a single batch of n operations diff --git a/collect_test.go b/collect_test.go index c1541e9..1c269cb 100644 --- a/collect_test.go +++ b/collect_test.go @@ -6,6 +6,7 @@ import ( "runtime/debug" "strings" "testing" + "time" ) // collectSink absorbs results of measured work so the compiler cannot eliminate it. @@ -49,6 +50,7 @@ func TestCollectRejectsBadOptions(t *testing.T) { {"negative Repeats", CollectOptions{Repeats: -1, InnerLoops: 1}, "Repeats must not be negative"}, {"too few Repeats", CollectOptions{Repeats: 10, InnerLoops: 1}, "at least"}, {"negative Warmup", CollectOptions{Repeats: 11, InnerLoops: 1, Warmup: -1}, "SkipWarmup"}, + {"negative WarmupDuration", CollectOptions{Repeats: 11, InnerLoops: 1, WarmupDuration: -time.Millisecond}, "WarmupDuration"}, {"unknown Order", CollectOptions{Repeats: 11, InnerLoops: 1, Order: Order(42)}, "Order"}, } for _, c := range cases { @@ -152,10 +154,12 @@ func TestCollectWarmupCounts(t *testing.T) { return &n, Candidate{Batch: func(uint64) { n++ }} } - // Warmup: 0 selects DefaultWarmup, so each candidate runs Repeats+DefaultWarmup times. + // Warmup: 0 selects DefaultWarmup, so each candidate runs Repeats+DefaultWarmup + // times. A nanosecond of WarmupDuration is met by the first round, which + // leaves the count to decide. na, ca := count() nb, cb := count() - if _, _, err := Collect(ca, cb, CollectOptions{Repeats: 11, InnerLoops: 1}); err != nil { + if _, _, err := Collect(ca, cb, CollectOptions{Repeats: 11, InnerLoops: 1, WarmupDuration: time.Nanosecond}); err != nil { t.Fatalf("unexpected error: %v", err) } if want := 11 + DefaultWarmup; *na != want || *nb != want { @@ -165,7 +169,7 @@ func TestCollectWarmupCounts(t *testing.T) { // Explicit warm-up count. na, ca = count() nb, cb = count() - if _, _, err := Collect(ca, cb, CollectOptions{Repeats: 11, InnerLoops: 1, Warmup: 3}); err != nil { + if _, _, err := Collect(ca, cb, CollectOptions{Repeats: 11, InnerLoops: 1, Warmup: 3, WarmupDuration: time.Nanosecond}); err != nil { t.Fatalf("unexpected error: %v", err) } if *na != 14 || *nb != 14 { @@ -183,6 +187,76 @@ func TestCollectWarmupCounts(t *testing.T) { } } +// TestCollectWarmupRunsForItsDuration checks that a comparison does not start +// measuring before the candidates have run long enough to settle. It belongs to +// Collect's warm-up, which is bounded by time as well as by a batch count +// because a count of short batches cannot turn over a large cache. With +// batches that take no time at all, a Collect call has to last at least the +// requested WarmupDuration, and SkipWarmup has to switch the duration off along +// with the count. +func TestCollectWarmupRunsForItsDuration(t *testing.T) { + const want = 40 * time.Millisecond + start := time.Now() + if _, _, err := Collect(noopCandidate(), noopCandidate(), CollectOptions{Repeats: 11, InnerLoops: 1, WarmupDuration: want}); err != nil { + t.Fatalf("unexpected error: %v", err) + } + if got := time.Since(start); got < want { + t.Errorf("Collect returned after %v, before the %v warm-up could have run", got, want) + } + + start = time.Now() + if _, _, err := Collect(noopCandidate(), noopCandidate(), CollectOptions{Repeats: 11, InnerLoops: 1, WarmupDuration: time.Hour, SkipWarmup: true}); err != nil { + t.Fatalf("unexpected error: %v", err) + } + if got := time.Since(start); got > time.Minute { + t.Errorf("SkipWarmup should disable the warm-up duration too, but Collect took %v", got) + } +} + +// TestCollectWarmupAlternatesOrder checks that neither candidate gets the last +// word before the measurement starts. It belongs to Collect's warm-up, which +// reverses the order every round like OrderABBA does, because a warm-up that +// always ends on the same candidate hands that candidate the caches. The +// expectation is the exact sequence A B B A A B B A for four rounds, followed +// by the measurement's own ABBA. +func TestCollectWarmupAlternatesOrder(t *testing.T) { + log, a, b := recorder() + if _, _, err := Collect(a, b, CollectOptions{Repeats: 12, InnerLoops: 1, Warmup: 4, WarmupDuration: time.Nanosecond}); err != nil { + t.Fatalf("unexpected error: %v", err) + } + if got, want := strings.Join((*log)[:8], ""), "ABBAABBA"; got != want { + t.Errorf("warm-up order: got %s, want %s", got, want) + } + if got, want := strings.Join((*log)[8:], ""), strings.Repeat("ABBA", 6); got != want { + t.Errorf("measurement order after warm-up: got %s, want %s", got, want) + } +} + +// TestMeasureInTurnBalancesManyCandidates checks the loop that lets several +// candidates share one measurement, which is what the paired A/A validation +// runs on. Under OrderABBA every round is the previous one reversed, so every +// candidate holds every position equally often, and each gets its own sample +// series of the requested length. +func TestMeasureInTurnBalancesManyCandidates(t *testing.T) { + var log []string + mk := func(name string) Candidate { + return Candidate{Name: name, Batch: func(uint64) { log = append(log, name) }} + } + s := schedule{repeats: 4, innerLoops: 1} + samples := measureInTurn([]Candidate{mk("a"), mk("b"), mk("c")}, s) + if got, want := strings.Join(log, ""), "abccbaabccba"; got != want { + t.Errorf("order: got %s, want %s", got, want) + } + if len(samples) != 3 { + t.Fatalf("expected one series per candidate, got %d", len(samples)) + } + for i, series := range samples { + if len(series) != 4 { + t.Errorf("series %d has %d samples, want 4", i, len(series)) + } + } +} + func TestCollectSetupTeardownOrderAndCount(t *testing.T) { var log []string c := Candidate{ @@ -215,7 +289,7 @@ func TestCollectSetupTeardownRunDuringWarmup(t *testing.T) { Batch: func(uint64) {}, Teardown: func() { teardowns++ }, } - if _, _, err := Collect(c, noopCandidate(), CollectOptions{Repeats: 11, InnerLoops: 1, Warmup: 2}); err != nil { + if _, _, err := Collect(c, noopCandidate(), CollectOptions{Repeats: 11, InnerLoops: 1, Warmup: 2, WarmupDuration: time.Nanosecond}); err != nil { t.Fatalf("unexpected error: %v", err) } if want := 13; setups != want || teardowns != want { diff --git a/compare.go b/compare.go index ab32a34..c50aa0d 100644 --- a/compare.go +++ b/compare.go @@ -100,6 +100,17 @@ type Report struct { // means the test could not be run. DriftA, DriftB DriftReport + // DriftRatio tests the ratio B/A of each pair of neighbouring batches for a + // trend across the run. A zero N means the test could not be run. + // + // It sees what DriftA and DriftB cannot: a machine that slows down slows + // both candidates alike and leaves the ratio flat, while one candidate that + // started with the caches to itself and loses them over the run moves the + // ratio and nothing else. A trend here means the two had not reached a + // steady state when the measurement began, so the difference depends on how + // long the run was; see [DefaultWarmupDuration]. + DriftRatio DriftReport + // Resolved is the short answer: the difference is both statistically // distinguishable from zero and larger than what this setup invents on its // own. It is deliberately conservative, and false does not mean the @@ -172,10 +183,16 @@ func sortedKeys(m map[float64]float64) []float64 { // what makes differences far below the clock's resolution measurable. // 2. Run each candidate against itself, repeatedly, to find out what this setup // reports as a difference when there is provably none. That is the noise -// floor, and it is the number every result has to be read against. -// 3. Measure the two candidates against each other, interleaved. -// 4. Test each series for a trend across the run, which resampling cannot see -// because it discards the order the samples arrived in. +// floor, and it is the number every result has to be read against. The two +// candidates' validations are interleaved batch by batch, see +// [ValidatePair], so that neither enters the comparison with the caches to +// itself. +// 3. Warm both candidates up, alternately, for at least +// CollectOptions.WarmupDuration, then measure them against each other, +// interleaved. +// 4. Test each series, and the ratio between them, for a trend across the run, +// which resampling cannot see because it discards the order the samples +// arrived in. // 5. Resample in blocks if the measurements turned out to be correlated with // their neighbours, and as single observations if they did not. // 6. Report the difference, an interval around it, and whether it clears both @@ -199,6 +216,20 @@ func sortedKeys(m map[float64]float64) []float64 { // pay only for the measurement, accepting that the result then has nothing to be // read against. // +// Two warm-ups of CollectOptions.WarmupDuration come on top, one before the +// validation and one before the measurement, which at [DefaultWarmupDuration] is +// 0.6 s in all, or 0.3 s with SkipValidation. +// +// # What ran before it +// +// Whatever touched one candidate's data last before Compare, such as building +// its fixture, starts that candidate with a warm cache. The warm-ups are there +// to wash this out, and with a working set near the size of the last-level +// cache they need the time they take. Building the two candidates' data in an +// interleaved or balanced order costs nothing and removes the question. When a +// head start survives anyway, Report.DriftRatio shows it as a trend and +// Report.Warnings says so. +// // # What it still cannot tell you // // The measured difference is that of the whole batch body, not of the isolated @@ -236,36 +267,27 @@ func Compare(a, b Candidate, opt CompareOptions) (Report, error) { // measurement have to describe the same setup, and leaving InnerLoops at // zero would have each of the three runs calibrate for itself. if co.InnerLoops == 0 { - calOpt := CalibrationOptions{ + var err error + co.InnerLoops, err = calibratePair(a, b, CalibrationOptions{ MaxQuantizationError: co.MaxQuantizationError, MaxInnerLoops: co.MaxInnerLoops, GCBetween: co.GCBetween, DisableGC: co.DisableGC, - } - calA, err := CalibrateInnerLoops(a, calOpt) - if err != nil { - return Report{}, fmt.Errorf("calibrating candidate %s: %w", a.label("A"), err) - } - calB, err := CalibrateInnerLoops(b, calOpt) + }) if err != nil { - return Report{}, fmt.Errorf("calibrating candidate %s: %w", b.label("B"), err) + return Report{}, err } - // The cheaper candidate needs the larger batch, so the maximum leaves - // both at or above the target duration. - co.InnerLoops = max(calA.InnerLoops, calB.InnerLoops) } r := Report{BlockLength: 1} if !opt.SkipValidation { + // Together rather than one after the other: a candidate validated on + // its own would enter the measurement with the caches to itself. vo := ValidationOptions{Collect: co, Runs: opt.ValidationRuns, Resamples: opt.Resamples} - va, err := ValidateHarness(a, vo) - if err != nil { - return Report{}, fmt.Errorf("validating candidate %s: %w", a.label("A"), err) - } - vb, err := ValidateHarness(b, vo) + va, vb, err := ValidatePair(a, b, vo) if err != nil { - return Report{}, fmt.Errorf("validating candidate %s: %w", b.label("B"), err) + return Report{}, fmt.Errorf("validating candidates %s and %s: %w", a.label("A"), b.label("B"), err) } r.Validated = true r.ValidationA, r.ValidationB = va, vb @@ -286,6 +308,9 @@ func Compare(a, b Candidate, opt CompareOptions) (Report, error) { if d, err := DetectDrift(sb); err == nil { r.DriftB = d } + if d, err := DetectDrift(pairRatios(sa, sb)); err == nil { + r.DriftRatio = d + } // Without validation there is no A/A estimate of the dependence, so fall // back to the run itself. It is the same quantity measured on one sample @@ -315,6 +340,25 @@ func Compare(a, b Candidate, opt CompareOptions) (Report, error) { return r, nil } +// pairRatios returns b[i]/a[i] for each pair of samples taken next to each +// other, which is the series a head start of one candidate shows up in. A zero +// in a produces a non-finite ratio, which DetectDrift then declines. +func pairRatios(a, b []float64) []float64 { + ratios := make([]float64, min(len(a), len(b))) + for i := range ratios { + ratios[i] = b[i] / a[i] + } + return ratios +} + +// resolution is the smallest relative change that could move this report's +// conclusion: the larger of the noise floor and half the interval's width. It +// exists so that a trend is only worth a warning when it is large enough to +// matter to the question asked, not merely significant. +func (r Report) resolution() float64 { + return math.Max(r.NoiseFloor, (r.Estimate.High-r.Estimate.Low)/2) +} + // warnings lists the things that undermine a report, in plain sentences. func (r Report) warnings() []string { var w []string @@ -347,6 +391,15 @@ func (r Report) warnings() []string { } } + // Unlike drift in one series, a trend in the ratio is a bias: it means one + // candidate was warmer than the other for part of the run. + if d := r.DriftRatio; d.N > 0 && d.Drifted(DriftLevel) && math.Abs(d.RelativeShift) > r.resolution() { + w = append(w, fmt.Sprintf( + "the ratio B/A shifted %+.2f%% from the first half of the run to the second, so the candidates had not reached a steady state and the difference depends on how long the run was; "+ + "a common cause is a head start for whichever candidate ran alone last before the comparison or had its data built last, which a longer CollectOptions.WarmupDuration removes", + d.RelativeShift*100)) + } + if r.Validated { if tie := math.Max(r.ValidationA.TieRate, r.ValidationB.TieRate); tie > 0.05 { w = append(w, fmt.Sprintf( diff --git a/compare_test.go b/compare_test.go index fca19e4..c0ba6ce 100644 --- a/compare_test.go +++ b/compare_test.go @@ -4,6 +4,7 @@ import ( "math" "strings" "testing" + "time" ) var compareSink uint64 @@ -39,9 +40,13 @@ func scaledCandidate(name string, mult uint64) Candidate { // trials, and cost is still small: validation is the dominant cost of a // comparison, but these run against a fixed 20000-operation batch rather than // calibrating, so ten of them stay well under a second. +// +// The warm-up is cut to 10 ms. The default is sized for working sets near the +// size of the last-level cache, and these candidates touch no memory at all, +// so the default would only add 0.6 s to every comparison. func fastCompare() CompareOptions { return CompareOptions{ - Collect: CollectOptions{Repeats: 51, InnerLoops: 20000}, + Collect: CollectOptions{Repeats: 51, InnerLoops: 20000, WarmupDuration: 10 * time.Millisecond}, ValidationRuns: 10, Resamples: 600, } @@ -439,6 +444,26 @@ func TestReportWarnings(t *testing.T) { }, want: "candidate B drifted", }, + { + name: "ratio trend larger than the resolution", + report: Report{ + Validated: true, NoiseFloor: 0.01, + Estimate: Estimate{Delta: -0.3, Low: -0.35, High: -0.25}, + DriftRatio: DriftReport{N: 41, PValue: 0.001, RelativeShift: 0.12}, + }, + want: "had not reached a steady state", + }, + { + name: "ratio trend smaller than the resolution", + report: Report{ + Validated: true, NoiseFloor: 0.02, + Estimate: Estimate{Delta: 0.5, Low: 0.4, High: 0.6}, + DriftRatio: DriftReport{N: 101, PValue: 0.001, RelativeShift: 0.01}, + ValidationA: HarnessValidation{Level: 0.95}, + ValidationB: HarnessValidation{Level: 0.95}, + }, + absent: "steady state", + }, { name: "coarse measurement", report: Report{ @@ -486,3 +511,45 @@ func TestReportWarnings(t *testing.T) { }) } } + +// TestCompareNeverLetsOneCandidateRunAlone checks the property that keeps a +// comparison of identical code from favouring one side: from the first +// validation batch to the last measured one, the two candidates share the +// machine. Compare validates both, warms both up and measures both, and any +// stretch in which one of them runs alone for long leaves it with the caches +// to itself (issue #111). The expectation is that no candidate ever runs more +// than twice in a row over the whole call. +func TestCompareNeverLetsOneCandidateRunAlone(t *testing.T) { + var log []string + mk := func(name string) Candidate { + work := scaledCandidate(name, 1) + return Candidate{Name: name, Batch: func(n uint64) { + log = append(log, name) + work.Batch(n) + }} + } + opt := fastCompare() + opt.ValidationRuns = 3 + opt.Collect.Repeats = 11 + opt.Collect.Warmup = 4 + if _, err := Compare(mk("a"), mk("b"), opt); err != nil { + t.Fatalf("unexpected error: %v", err) + } + if got := longestRun(log); got > 2 { + t.Errorf("a candidate ran %d times in a row during Compare; at most 2 keeps the caches shared", got) + } +} + +// TestPairRatios checks the series in which a head start of one candidate +// becomes visible: the ratio B/A of neighbouring samples, pair by pair, cut to +// the shorter input. A zero in A has to come out non-finite, so that +// DetectDrift declines the series rather than reporting on it. +func TestPairRatios(t *testing.T) { + got := pairRatios([]float64{1, 2, 4}, []float64{2, 2, 2, 9}) + if want := []float64{2, 1, 0.5}; len(got) != 3 || got[0] != want[0] || got[1] != want[1] || got[2] != want[2] { + t.Errorf("pairRatios: got %v, want %v", got, want) + } + if _, err := DetectDrift(pairRatios([]float64{0, 1, 1, 1}, []float64{1, 1, 1, 1})); err == nil { + t.Error("a zero sample in A should make the ratio series unusable for DetectDrift") + } +} diff --git a/validate.go b/validate.go index 5183ab1..27eeb67 100644 --- a/validate.go +++ b/validate.go @@ -311,16 +311,24 @@ type ValidationOptions struct { // each direction, because splitting ties needs the confidence both ways. The // resampling dominates. Measured on a candidate calibrated to 48 microsecond // batches, the whole validation took 0.76 s at ten runs and 3.03 s at forty, of -// which the measurement itself was under a tenth. +// which the measurement itself was under a tenth. On top of that comes one +// warm-up of CollectOptions.WarmupDuration, [DefaultWarmupDuration] unless set, +// before the first run only; the later runs warm up by count alone, since the +// candidate has been running all along. // // A worked use, and the reason the API exists: measure the noise floor first, // then require a real result to clear it. // // Validate both candidates, not just one: they need not be equally well -// behaved, and a comparison is only as trustworthy as the worse of them. +// behaved, and a comparison is only as trustworthy as the worse of them. When +// the two are about to be compared, validate them together with [ValidatePair] +// rather than calling this function twice. A candidate validated on its own +// runs alone for the whole validation and leaves the caches full of its own +// data, so the one validated last starts the comparison warm and the other +// cold; with a working set near the size of the last-level cache that alone has +// produced differences of 10 to 50% between identical code. // -// va, err := rtcompare.ValidateHarness(fast, rtcompare.ValidationOptions{Collect: opts}) -// vb, err := rtcompare.ValidateHarness(slow, rtcompare.ValidationOptions{Collect: opts}) +// va, vb, err := rtcompare.ValidatePair(fast, slow, rtcompare.ValidationOptions{Collect: opts}) // floor := max(va.NoiseFloor, vb.NoiseFloor) // // sa, sb, err := rtcompare.Collect(fast, slow, opts) @@ -332,14 +340,121 @@ func ValidateHarness(c Candidate, opt ValidationOptions) (HarnessValidation, err if c.Batch == nil { return HarnessValidation{}, fmt.Errorf("rtcompare: candidate %s has a nil Batch function", c.label("under validation")) } + opt, s, err := opt.resolve() + if err != nil { + return HarnessValidation{}, err + } + if s.innerLoops == 0 { + // Calibrate once. Recalibrating per run would let the batch size drift + // between runs and make their noise levels incomparable. + cal, err := CalibrateInnerLoops(c, opt.calibration()) + if err != nil { + return HarnessValidation{}, fmt.Errorf("calibrating candidate %s: %w", c.label("under validation"), err) + } + s.innerLoops = cal.InnerLoops + } + + runs := newAARuns(opt) + for run := range opt.Runs { + samples := measureInTurn([]Candidate{c, c}, s.forRun(run)) + runs.add(samples[0], samples[1]) + } + return runs.result(s.innerLoops), nil +} + +// ValidatePair runs the A/A validation of two candidates together, and returns +// the same [HarnessValidation] for each that [ValidateHarness] would. +// +// Parameters: a and b are the two candidates that are about to be compared, +// and opt holds the options the comparison will use, exactly as for +// ValidateHarness. When opt.Collect.InnerLoops is zero, both candidates are +// calibrated and the larger batch size is used for both validations, as +// [Compare] and [Collect] do it, so that the floors describe the batch size the +// comparison will measure at. +// +// Use it in place of two calls to ValidateHarness whenever the validation is +// followed by a comparison of the same two candidates, which is what [Compare] +// does. The difference is the order in which the batches run. Called one after +// the other, the two validations each let one candidate run alone for a long +// time, and whichever ran last starts the comparison with the caches full of its +// own data. For a working set near the size of the last-level cache that is a +// head start large enough to report identical code as tens of percent apart +// (issue #111). ValidatePair instead runs every A/A experiment of a +// interleaved with one of b, batch by batch in the balanced order A, B, B′, A′ +// and back, so neither candidate ever runs alone for more than two batches. The +// floors are measured with the other candidate competing for the caches, as it +// will during the comparison; see pairSlots for why the order is that one. +// +// The cost is the same as the two separate calls, minus one warm-up. +// +// An error is returned if either candidate has a nil Batch, for the option +// errors of ValidateHarness, or if calibration fails. +// +// va, vb, err := rtcompare.ValidatePair(a, b, rtcompare.ValidationOptions{Collect: opts}) +// if err != nil { ... } +// floor := max(va.NoiseFloor, vb.NoiseFloor) +// sa, sb, err := rtcompare.Collect(a, b, opts) +func ValidatePair(a, b Candidate, opt ValidationOptions) (va, vb HarnessValidation, err error) { + if a.Batch == nil { + return va, vb, fmt.Errorf("rtcompare: candidate %s has a nil Batch function", a.label("A")) + } + if b.Batch == nil { + return va, vb, fmt.Errorf("rtcompare: candidate %s has a nil Batch function", b.label("B")) + } + opt, s, err := opt.resolve() + if err != nil { + return va, vb, err + } + if s.innerLoops == 0 { + if s.innerLoops, err = calibratePair(a, b, opt.calibration()); err != nil { + return va, vb, err + } + } + + both := [2]Candidate{a, b} + cands := make([]Candidate, len(pairSlots)) + var halves [2][]int // the two slots of each candidate + for slot, who := range pairSlots { + cands[slot] = both[who] + halves[who] = append(halves[who], slot) + } + runs := [2]*aaRuns{newAARuns(opt), newAARuns(opt)} + for run := range opt.Runs { + samples := measureInTurn(cands, s.forRun(run)) + for who, h := range halves { + runs[who].add(samples[h[0]], samples[h[1]]) + } + } + return runs[0].result(s.innerLoops), runs[1].result(s.innerLoops), nil +} + +// pairSlots says which candidate, 0 for a and 1 for b, runs in each of the four +// slots of a paired validation round. Slots 0 and 3 are a's two halves of its +// A/A experiment, slots 1 and 2 are b's. +// +// The order matters because an A/A experiment is only unbiased if its two +// halves are measured under the same conditions, and one condition is what ran +// just before. Rounds alternate forwards and backwards, so the slot that ends +// one round also starts the next and runs twice in a row. With A, B, A′, B′ that +// would always be the same slot, A in one direction and B′ in the other: A +// would follow itself half the time and A′ never, which lent A a warm start +// that A′ did not have and inflated the floors. With A, B, B′, A′ each half of +// each candidate is preceded by its own candidate in one direction and by the +// other in the other, so both halves see the same history. +var pairSlots = [4]int{0, 1, 1, 0} + +// resolve checks the options, fills in their defaults and derives the +// measurement schedule, so that the single and the paired validation cannot +// interpret them differently. +func (opt ValidationOptions) resolve() (ValidationOptions, schedule, error) { if opt.Runs < 0 { - return HarnessValidation{}, fmt.Errorf("rtcompare: Runs must not be negative, got %d", opt.Runs) + return opt, schedule{}, fmt.Errorf("rtcompare: Runs must not be negative, got %d", opt.Runs) } if opt.Runs == 0 { opt.Runs = DefaultValidationRuns } if opt.Runs < 2 { - return HarnessValidation{}, fmt.Errorf("rtcompare: Runs must be at least 2 for a noise floor to mean anything, got %d", opt.Runs) + return opt, schedule{}, fmt.Errorf("rtcompare: Runs must be at least 2 for a noise floor to mean anything, got %d", opt.Runs) } if opt.Resamples == 0 { opt.Resamples = DefaultResamples @@ -348,111 +463,137 @@ func ValidateHarness(c Candidate, opt ValidationOptions) (HarnessValidation, err opt.Level = DefaultValidationLevel } if opt.Level <= 0.5 || opt.Level >= 1 { - return HarnessValidation{}, fmt.Errorf("rtcompare: Level must be in (0.5, 1), got %v", opt.Level) + return opt, schedule{}, fmt.Errorf("rtcompare: Level must be in (0.5, 1), got %v", opt.Level) } + s, err := opt.Collect.schedule() + return opt, s, err +} - co := opt.Collect - if co.InnerLoops == 0 { - // Calibrate once. Recalibrating per run would let the batch size drift - // between runs and make their noise levels incomparable. - cal, err := CalibrateInnerLoops(c, CalibrationOptions{ - MaxQuantizationError: co.MaxQuantizationError, - MaxInnerLoops: co.MaxInnerLoops, - GCBetween: co.GCBetween, - DisableGC: co.DisableGC, - }) - if err != nil { - return HarnessValidation{}, fmt.Errorf("calibrating candidate %s: %w", c.label("under validation"), err) - } - co.InnerLoops = cal.InnerLoops +// calibration derives the options for sizing the validation's batches from the +// measurement options, so that calibration sees the conditions the runs will. +func (opt ValidationOptions) calibration() CalibrationOptions { + return CalibrationOptions{ + MaxQuantizationError: opt.Collect.MaxQuantizationError, + MaxInnerLoops: opt.Collect.MaxInnerLoops, + GCBetween: opt.Collect.GCBetween, + DisableGC: opt.Collect.DisableGC, + } +} + +// forRun returns the schedule for one A/A run. Only the first run warms up for +// the full duration: the later ones follow directly on the batches of the run +// before, so the caches are already in the state a duration would buy, and +// paying for it again would multiply the cost of a validation by the run count. +func (s schedule) forRun(run int) schedule { + if run > 0 { + s.warmupDuration = 0 } + return s +} - deltas := make([]float64, 0, opt.Runs) - confidences := make([]float64, 0, opt.Runs) - tieRates := make([]float64, 0, opt.Runs) - driftShifts := make([]float64, 0, 2*opt.Runs) - autocorrelations := make([]float64, 0, 2*opt.Runs) - driftedRuns := 0 +// aaRuns accumulates the per-run figures of repeated A/A experiments on one +// candidate and summarises them. It exists so that the single and the paired +// validation compute their reports in exactly the same way. +type aaRuns struct { + opt ValidationOptions + deltas []float64 + confidences []float64 + tieRates []float64 + driftShifts []float64 + autocorrelations []float64 + driftedRuns int +} - for run := range opt.Runs { - sampleA, sampleB, err := Collect(c, c, co) +func newAARuns(opt ValidationOptions) *aaRuns { + return &aaRuns{ + opt: opt, + deltas: make([]float64, 0, opt.Runs), + confidences: make([]float64, 0, opt.Runs), + tieRates: make([]float64, 0, opt.Runs), + driftShifts: make([]float64, 0, 2*opt.Runs), + autocorrelations: make([]float64, 0, 2*opt.Runs), + } +} + +// add records one A/A run from its two sample series. +func (r *aaRuns) add(sampleA, sampleB []float64) { + medA, medB := Median(sampleA), Median(sampleB) + delta := 0.0 + if medB != 0 && !math.IsNaN(medA) && !math.IsNaN(medB) { + delta = 1 - medA/medB + } + r.deltas = append(r.deltas, delta) + + // Both directions, so that ties can be split. With + // confAB = P(medA < medB) + P(tie) + // confBA = P(medA > medB) + P(tie) + // and the three probabilities summing to one, (confAB + 1 - confBA)/2 + // collapses to P(medA < medB) + P(tie)/2, and confAB + confBA - 1 + // recovers the tie rate. Both follow from the public API alone. + confAB := BootstrapConfidence(sampleA, sampleB, []float64{0.0}, r.opt.Resamples, 0)[0.0] + confBA := BootstrapConfidence(sampleB, sampleA, []float64{0.0}, r.opt.Resamples, 0)[0.0] + r.confidences = append(r.confidences, (confAB+1-confBA)/2) + r.tieRates = append(r.tieRates, math.Max(0, confAB+confBA-1)) + + // Drift is a property of the order the samples arrived in, which the + // bootstrap above has already discarded. Both series are examined; a run + // counts as drifting if either did. + drifted := false + for _, series := range [][]float64{sampleA, sampleB} { + d, err := DetectDrift(series) if err != nil { - return HarnessValidation{}, fmt.Errorf("A/A run %d of %d: %w", run+1, opt.Runs, err) - } - medA, medB := Median(sampleA), Median(sampleB) - delta := 0.0 - if medB != 0 && !math.IsNaN(medA) && !math.IsNaN(medB) { - delta = 1 - medA/medB - } - deltas = append(deltas, delta) - - // Both directions, so that ties can be split. With - // confAB = P(medA < medB) + P(tie) - // confBA = P(medA > medB) + P(tie) - // and the three probabilities summing to one, (confAB + 1 - confBA)/2 - // collapses to P(medA < medB) + P(tie)/2, and confAB + confBA - 1 - // recovers the tie rate. Both follow from the public API alone. - confAB := BootstrapConfidence(sampleA, sampleB, []float64{0.0}, opt.Resamples, 0)[0.0] - confBA := BootstrapConfidence(sampleB, sampleA, []float64{0.0}, opt.Resamples, 0)[0.0] - confidences = append(confidences, (confAB+1-confBA)/2) - tieRates = append(tieRates, math.Max(0, confAB+confBA-1)) - - // Drift is a property of the order the samples arrived in, which the - // bootstrap above has already discarded. Both series are examined; a run - // counts as drifting if either did. - drifted := false - for _, series := range [][]float64{sampleA, sampleB} { - d, err := DetectDrift(series) - if err != nil { - // Too few samples to look for a trend, or a non-finite value. - // Neither is a reason to fail the validation. - continue - } - driftShifts = append(driftShifts, math.Abs(d.RelativeShift)) - autocorrelations = append(autocorrelations, lag1Autocorrelation(series)) - if d.Drifted(DriftLevel) { - drifted = true - } + // Too few samples to look for a trend, or a non-finite value. + // Neither is a reason to fail the validation. + continue } - if drifted { - driftedRuns++ + r.driftShifts = append(r.driftShifts, math.Abs(d.RelativeShift)) + r.autocorrelations = append(r.autocorrelations, lag1Autocorrelation(series)) + if d.Drifted(DriftLevel) { + drifted = true } } + if drifted { + r.driftedRuns++ + } +} +// result summarises the recorded runs. +func (r *aaRuns) result(innerLoops uint64) HarnessValidation { // Sorted so that the floor can be read off as a quantile. This is a private // copy, so the caller's Deltas keep the order the runs were performed in. - absDeltas := make([]float64, len(deltas)) - for i, d := range deltas { + absDeltas := make([]float64, len(r.deltas)) + for i, d := range r.deltas { absDeltas[i] = math.Abs(d) } slices.Sort(absDeltas) var sum float64 falseSignals := 0 - for _, conf := range confidences { + for _, conf := range r.confidences { sum += conf - if conf > opt.Level || conf < 1-opt.Level { + if conf > r.opt.Level || conf < 1-r.opt.Level { falseSignals++ } } + runs := len(r.deltas) return HarnessValidation{ - Runs: opt.Runs, - InnerLoops: co.InnerLoops, + Runs: runs, + InnerLoops: innerLoops, NoiseFloor: quantileOfSorted(absDeltas, NoiseFloorQuantile), - MaxObservedNoise: absDeltas[len(absDeltas)-1], + MaxObservedNoise: absDeltas[runs-1], TypicalNoise: Median(absDeltas), - MeanConfidence: sum / float64(len(confidences)), - MedianConfidence: Median(confidences), - TieRate: Median(tieRates), - FalseSignalRate: float64(falseSignals) / float64(len(confidences)), - Level: opt.Level, - DriftRate: float64(driftedRuns) / float64(opt.Runs), - MedianDriftShift: medianOrZero(driftShifts), - Autocorrelation: medianOrZero(autocorrelations), - Deltas: deltas, - Confidences: confidences, - }, nil + MeanConfidence: sum / float64(runs), + MedianConfidence: Median(r.confidences), + TieRate: Median(r.tieRates), + FalseSignalRate: float64(falseSignals) / float64(runs), + Level: r.opt.Level, + DriftRate: float64(r.driftedRuns) / float64(runs), + MedianDriftShift: medianOrZero(r.driftShifts), + Autocorrelation: medianOrZero(r.autocorrelations), + Deltas: r.deltas, + Confidences: r.confidences, + } } // medianOrZero is Median with an empty input mapping to zero rather than to the diff --git a/validate_test.go b/validate_test.go index c624dd2..ac910ea 100644 --- a/validate_test.go +++ b/validate_test.go @@ -5,6 +5,7 @@ import ( "slices" "strings" "testing" + "time" ) var validateSink uint64 @@ -308,3 +309,171 @@ func TestNoiseFloorIsAQuantileNotAMaximum(t *testing.T) { t.Errorf("expected %d recorded deltas, got %d", v.Runs, len(v.Deltas)) } } + +// longestRun returns the longest stretch of identical consecutive entries. +func longestRun(log []string) int { + longest, run := 0, 0 + for i := range log { + if i > 0 && log[i] == log[i-1] { + run++ + } else { + run = 1 + } + longest = max(longest, run) + } + return longest +} + +// TestValidatePairInterleavesTheCandidates checks that validating two +// candidates before comparing them does not leave either of them with the +// caches to itself. It belongs to the A/A validation, which Compare runs for +// both candidates before measuring them against each other; validated one after +// the other, the second one entered the comparison warm (issue #111). The +// expectation is that the batches of the two candidates alternate throughout, +// so that no candidate ever runs more than twice in a row, and that each +// candidate still gets a full validation of its own. +func TestValidatePairInterleavesTheCandidates(t *testing.T) { + var log []string + mk := func(name string) Candidate { + return Candidate{Name: name, Batch: func(n uint64) { + log = append(log, name) + rng := NewDPRNG(0x2468) + var acc uint64 + for range n { + acc ^= rng.Uint64() + } + validateSink ^= acc + }} + } + opt := quickValidation(3) + opt.Collect.WarmupDuration = time.Millisecond + va, vb, err := ValidatePair(mk("a"), mk("b"), opt) + if err != nil { + t.Fatalf("unexpected error: %v", err) + } + if got := longestRun(log); got > 2 { + t.Errorf("a candidate ran %d times in a row during the paired validation; at most 2 keeps the caches shared", got) + } + for name, v := range map[string]HarnessValidation{"a": va, "b": vb} { + if v.Runs != 3 || len(v.Deltas) != 3 || v.InnerLoops != opt.Collect.InnerLoops { + t.Errorf("candidate %s: expected 3 runs at %d inner loops, got %d runs, %d deltas, %d inner loops", + name, opt.Collect.InnerLoops, v.Runs, len(v.Deltas), v.InnerLoops) + } + } +} + +// TestValidatePairCalibratesToTheLargerBatch checks that the two noise floors +// describe the batch size the comparison will actually measure at. Within the +// A/A validation, leaving InnerLoops at zero calibrates both candidates, and +// the cheaper one needs the larger batch; the pair is expected to use that +// larger size for both validations. +func TestValidatePairCalibratesToTheLargerBatch(t *testing.T) { + opt := ValidationOptions{ + Collect: CollectOptions{Repeats: 11, MaxQuantizationError: 0.01, WarmupDuration: time.Millisecond}, + Runs: 2, + Resamples: 300, + } + va, vb, err := ValidatePair(steadyCandidate(1), steadyCandidate(8), opt) + if err != nil { + t.Fatalf("unexpected error: %v", err) + } + if va.InnerLoops != vb.InnerLoops { + t.Errorf("the two validations used different batch sizes: %d and %d", va.InnerLoops, vb.InnerLoops) + } + cheap, err := CalibrateInnerLoops(steadyCandidate(1), CalibrationOptions{MaxQuantizationError: 0.01}) + if err != nil { + t.Fatalf("unexpected error: %v", err) + } + // Calibration is a noisy search, so only the order of magnitude is + // compared: the pair must be sized for the cheap candidate, not the costly. + if float64(va.InnerLoops) < float64(cheap.InnerLoops)/3 { + t.Errorf("the pair used %d inner loops, far below the %d the cheaper candidate needs", va.InnerLoops, cheap.InnerLoops) + } +} + +// TestValidatePairRejectsBadInput checks that a paired validation fails loudly +// rather than measuring something unintended. It shares its option checks with +// ValidateHarness, and is expected to name the candidate at fault when one of +// them has no Batch function. +func TestValidatePairRejectsBadInput(t *testing.T) { + good := steadyCandidate(1) + cases := []struct { + name string + a, b Candidate + opt ValidationOptions + want string + }{ + {"returns error for nil batch in A", Candidate{Name: "empty"}, good, quickValidation(3), `A ("empty")`}, + {"returns error for nil batch in B", good, Candidate{Name: "empty"}, quickValidation(3), `B ("empty")`}, + {"returns error for a single run", good, good, ValidationOptions{Runs: 1}, "at least 2"}, + {"returns error for a bad level", good, good, ValidationOptions{Runs: 3, Level: 0.4}, "Level"}, + {"returns error for bad collect options", good, good, ValidationOptions{Runs: 3, Collect: CollectOptions{Repeats: 3}}, "Repeats"}, + {"returns error when calibration fails", good, Candidate{Name: "ignores n", Batch: func(uint64) { validateSink++ }}, + ValidationOptions{Runs: 2, Collect: CollectOptions{Repeats: 11, MaxInnerLoops: 1000}}, "calibrating"}, + } + for _, c := range cases { + t.Run(c.name, func(t *testing.T) { + _, _, err := ValidatePair(c.a, c.b, c.opt) + if err == nil { + t.Fatal("expected an error, got nil") + } + if !strings.Contains(err.Error(), c.want) { + t.Errorf("error %q does not mention %q", err.Error(), c.want) + } + }) + } +} + +// TestValidationWarmsUpForItsDurationOnlyOnce checks that a validation's cost +// does not grow with the run count through the warm-up. Within the A/A +// validation, only the first run is preceded by a warm-up of the full +// duration; the later runs follow directly on the batches before them and are +// expected to warm up by count alone. +func TestValidationWarmsUpForItsDurationOnlyOnce(t *testing.T) { + s := schedule{warmupRounds: 3, warmupDuration: time.Second} + if got := s.forRun(0); got.warmupDuration != time.Second || got.warmupRounds != 3 { + t.Errorf("first run: got %+v, want the full warm-up", got) + } + if got := s.forRun(5); got.warmupDuration != 0 || got.warmupRounds != 3 { + t.Errorf("later run: got %+v, want the count without the duration", got) + } +} + +// TestPairSlotsGiveBothHalvesTheSameHistory checks that a paired validation +// measures each candidate's two A/A halves under the same conditions. Within +// the A/A validation, what ran just before a batch changes what the caches +// hold, so a half that more often follows its own candidate starts warmer and +// makes identical code look different. Replaying the order the measurement +// loop uses, each candidate's two slots are expected to be preceded by that +// same candidate equally often. +func TestPairSlotsGiveBothHalvesTheSameHistory(t *testing.T) { + var order []int // slot indices in run order + for round := range 8 { + for k := range pairSlots { + order = append(order, turn(k, len(pairSlots), round%2 == 0)) + } + } + // Read cyclically: over an even number of rounds the sequence repeats, and + // the first batch of a run follows the last one of the warm-up, which ran + // in the same alternating order. + selfPreceded := map[int]int{} + for i := range order { + slot, before := order[i], order[(i+len(order)-1)%len(order)] + if pairSlots[slot] == pairSlots[before] { + selfPreceded[slot]++ + } + } + halves := map[int][]int{} + for slot, who := range pairSlots { + halves[who] = append(halves[who], slot) + } + for who, slots := range halves { + if len(slots) != 2 { + t.Fatalf("candidate %d has %d slots, want 2", who, len(slots)) + } + if a, b := selfPreceded[slots[0]], selfPreceded[slots[1]]; a != b { + t.Errorf("candidate %d: slot %d follows its own candidate %d times, slot %d %d times", + who, slots[0], a, slots[1], b) + } + } +} From 3d3b666052b2929936535aa7259f8dc609eaa473 Mon Sep 17 00:00:00 2001 From: Tom Tonic Date: Sun, 27 Sep 2026 19:56:00 +0200 Subject: [PATCH 3/3] Add cmd/rtcompare-aa to reproduce the warm-start bias The command compares two identical, separately allocated pointer chases, so the true difference is zero. Its modes and flags: - -mode prefix runs Collect after validating no candidate, only A, only B, or A then B, one candidate at a time. - -mode compare runs Compare in both role assignments and prints the drift tests and warnings. - -build ab|ba|mixed selects the allocation order, which separates layout effects from role effects. Issue #111 asks for this reproduction to live in the repository, and #109 will extend the command. Co-Authored-By: Claude Opus 5.5 --- cmd/rtcompare-aa/main.go | 264 +++++++++++++++++++++++++++++++++++++++ 1 file changed, 264 insertions(+) create mode 100644 cmd/rtcompare-aa/main.go diff --git a/cmd/rtcompare-aa/main.go b/cmd/rtcompare-aa/main.go new file mode 100644 index 0000000..7be54cf --- /dev/null +++ b/cmd/rtcompare-aa/main.go @@ -0,0 +1,264 @@ +// Command rtcompare-aa compares two identical, separately allocated data +// structures with each other, so that the true difference is zero by +// construction and everything rtcompare reports is an artefact of the +// measurement. +// +// It exists to reproduce, and to re-check fixes for, two effects that only show +// up once the working set outgrows the caches: +// +// - a systematic head start for whichever candidate ran alone last before the +// measurement (issue #111), which -mode prefix isolates and -mode compare +// shows end to end; +// - scatter between processes far beyond the interval a single process +// reports (issue #109), for which the command is meant to be run several +// times. +// +// Pick -n so that one instance is roughly 0.5 to 1 times the size of the +// machine's last-level cache; a node is 64 bytes, so the default of 262,144 +// nodes is 16 MB. -n 4096 fits in the caches and serves as the control, where +// nothing should be reported. +package main + +import ( + "flag" + "fmt" + "os" + "time" + + "github.com/TomTonic/rtcompare" +) + +// node is one cache line: a value, a pointer to follow and padding. +type node struct { + val uint64 + next *node + _ [6]uint64 +} + +// fixture is a random pointer chase: nodes allocated in random order with +// random successors, and a fixed sequence of probes over them. +type fixture struct { + nodes []*node + probe []int32 +} + +// sink keeps the chases observable to the compiler. +var sink uint64 + +// build allocates both fixtures from the same seed, so that they hold the same +// content, in the order the flag asks for. Whatever is allocated last is also +// what the caches hold when the measurement starts, which is one of the head +// starts this command is about. +func build(n int, order string) (a, b *fixture, err error) { + fa, fb := newFixture(n), newFixture(n) + perm, next, probe := layout(n) + switch order { + case "ab": + fa.fill(perm, next, probe) + fb.fill(perm, next, probe) + case "ba": + fb.fill(perm, next, probe) + fa.fill(perm, next, probe) + case "mixed": + fillMixed(fa, fb, perm, next, probe) + default: + return nil, nil, fmt.Errorf("unknown build order %q, want ab, ba or mixed", order) + } + return fa, fb, nil +} + +func newFixture(n int) *fixture { + return &fixture{nodes: make([]*node, n), probe: make([]int32, 1<<20)} +} + +// layout draws the shared content of both fixtures: the order nodes are +// allocated in, each node's successor and the probe sequence. +func layout(n int) (perm, next []int, probe []int32) { + rng := rtcompare.NewDPRNG(1) + perm = make([]int, n) + for i := range perm { + perm[i] = i + } + for i := n - 1; i > 0; i-- { + j := int(rng.Uint32N(uint32(i + 1))) + perm[i], perm[j] = perm[j], perm[i] + } + next = make([]int, n) + for i := range next { + next[i] = int(rng.Uint32N(uint32(n))) + } + probe = make([]int32, 1<<20) + even := uint32(n &^ 1) // probe^1 has to stay in range + for i := range probe { + probe[i] = int32(rng.Uint32N(even)) + } + return perm, next, probe +} + +func (f *fixture) fill(perm, next []int, probe []int32) { + for _, i := range perm { + f.nodes[i] = &node{val: uint64(i)} + } + f.link(next, probe) +} + +func (f *fixture) link(next []int, probe []int32) { + for i, j := range next { + f.nodes[i].next = f.nodes[j] + } + copy(f.probe, probe) +} + +// fillMixed allocates the two fixtures node by node in turn, so that neither +// was built last. +func fillMixed(a, b *fixture, perm, next []int, probe []int32) { + for _, i := range perm { + a.nodes[i] = &node{val: uint64(i)} + b.nodes[i] = &node{val: uint64(i)} + } + a.link(next, probe) + b.link(next, probe) +} + +// candidate loads a node per operation and follows two pointers from it. Each +// operation depends on the one before, like a lookup whose branches depend on +// the data it loads, so every cache miss is paid in full. Every candidate keeps +// its own cursor into the probe sequence. +func candidate(name string, f *fixture) rtcompare.Candidate { + j := 0 + return rtcompare.Candidate{Name: name, Batch: func(n uint64) { + var acc uint64 + for range n { + x := f.nodes[int(f.probe[j])^int(acc&1)] + acc += x.val + x.next.val + x.next.next.val + if j++; j == len(f.probe) { + j = 0 + } + } + sink += acc + }} +} + +type config struct { + n int + mode string + order string + repeats int + validation int + warmup int + warmupDur time.Duration + skipVal bool +} + +func main() { + var c config + flag.IntVar(&c.n, "n", 1<<18, "nodes per instance (64 bytes each)") + flag.StringVar(&c.mode, "mode", "compare", "compare: Compare in both role assignments; prefix: Collect after validating none, A, B, or A then B") + flag.StringVar(&c.order, "build", "ab", "allocation order of the two fixtures: ab, ba or mixed") + flag.IntVar(&c.repeats, "repeats", 0, "CollectOptions.Repeats (0: default)") + flag.IntVar(&c.validation, "validation", 0, "CompareOptions.ValidationRuns (0: default)") + flag.IntVar(&c.warmup, "warmup", 0, "CollectOptions.Warmup (0: default)") + flag.DurationVar(&c.warmupDur, "warmupdur", 0, "CollectOptions.WarmupDuration (0: default)") + flag.BoolVar(&c.skipVal, "skipvalidation", false, "CompareOptions.SkipValidation") + flag.Parse() + + if err := run(c); err != nil { + fmt.Fprintln(os.Stderr, "rtcompare-aa:", err) + os.Exit(1) + } + + // sink is only ever written, so that the compiler cannot drop the chases; + // in a package main the linter can see that nothing reads it. This read + // tells it what the writes already told the compiler, as in + // cmd/rtcompare-example. + _ = sink +} + +func run(c config) error { + fa, fb, err := build(c.n, c.order) + if err != nil { + return err + } + switch c.mode { + case "compare": + return compare(c, fa, fb) + case "prefix": + return prefix(c, fa, fb) + default: + return fmt.Errorf("unknown mode %q, want compare or prefix", c.mode) + } +} + +func collectOptions(c config) rtcompare.CollectOptions { + return rtcompare.CollectOptions{MaxQuantizationError: 1e-4, Repeats: c.repeats, Warmup: c.warmup, WarmupDuration: c.warmupDur} +} + +// compare runs Compare with each fixture in each role. Delta is 1 - A/B, so a +// negative delta means A measured slower. +func compare(c config, fa, fb *fixture) error { + opt := rtcompare.CompareOptions{ + Collect: collectOptions(c), + ValidationRuns: c.validation, + SkipValidation: c.skipVal, + } + for _, roles := range []struct { + name string + a, b *fixture + }{{"A=first B=second", fa, fb}, {"A=second B=first", fb, fa}} { + rep, err := rtcompare.Compare(candidate("A", roles.a), candidate("B", roles.b), opt) + if err != nil { + return err + } + fmt.Printf("n=%d build=%s %s: A %.2f ns, B %.2f ns, delta %+.2f%% [%+.2f, %+.2f] floor %.2f%% resolved=%v\n", + c.n, c.order, roles.name, rep.NsPerOpA, rep.NsPerOpB, + 100*rep.Estimate.Delta, 100*rep.Estimate.Low, 100*rep.Estimate.High, 100*rep.NoiseFloor, rep.Resolved) + fmt.Printf(" drift A %+.1f%% (p %.3f), B %+.1f%% (p %.3f), B/A %+.1f%% (p %.3f)\n", + 100*rep.DriftA.RelativeShift, rep.DriftA.PValue, 100*rep.DriftB.RelativeShift, rep.DriftB.PValue, + 100*rep.DriftRatio.RelativeShift, rep.DriftRatio.PValue) + for _, w := range rep.Warnings { + fmt.Println(" warning:", w) + } + } + return nil +} + +// prefix runs Collect at a fixed, calibrated batch size after validating no +// candidate, only A, only B, or A and then B, which is what Compare did before +// v0.7.0, and prints the ratio of the medians B/A. 1.00 is the truth. The +// validations are deliberately run one candidate at a time, as a caller who +// does not use ValidatePair would, so that the remaining ratio shows what +// Collect's warm-up washes out on its own. +func prefix(c config, fa, fb *fixture) error { + co := collectOptions(c) + cal, err := rtcompare.CalibrateInnerLoops(candidate("A", fa), rtcompare.CalibrationOptions{MaxQuantizationError: co.MaxQuantizationError}) + if err != nil { + return err + } + co.InnerLoops = cal.InnerLoops + vo := rtcompare.ValidationOptions{Collect: co, Runs: c.validation} + for _, pre := range []string{"", "A", "B", "AB"} { + a, b := candidate("A", fa), candidate("B", fb) + for _, who := range pre { + switch who { + case 'A': + _, err = rtcompare.ValidateHarness(a, vo) + case 'B': + _, err = rtcompare.ValidateHarness(b, vo) + } + if err != nil { + return err + } + } + sa, sb, err := rtcompare.Collect(a, b, co) + if err != nil { + return err + } + label := pre + if label == "" { + label = "none" + } + fmt.Printf("n=%d build=%s validated %-4s then Collect (InnerLoops %d): median B/A = %.3f\n", + c.n, c.order, label, co.InnerLoops, rtcompare.Median(sb)/rtcompare.Median(sa)) + } + return nil +}