Stop Compare from favouring the candidate that ran alone last - #112
Merged
Merged
Conversation
Compare validated candidate A, then candidate B, and then measured both after a single warm-up pair. B therefore entered the measurement with the last-level cache full of its own data. With a working set near the size of that cache and short calibrated batches, the head start survived into the medians: identical code measured up to 66% apart, reported as resolved (#111). - ValidatePair validates two candidates together, interleaved batch by batch in the order A, B, B', A' and back. In that order each half of each A/A experiment has the same history. Compare uses it. - Collect warms up for at least CollectOptions.WarmupDuration (DefaultWarmupDuration, 300 ms) in addition to the batch count, and alternates the order of the candidates while it does. A validation pays for the duration once, not once per run. - Report.DriftRatio tests the ratio B/A of neighbouring batches for a trend. A warning appears when the trend is significant and larger than the result's resolution. - HOWTO and README describe the pitfall and the checks against it. Measured with cmd/rtcompare-aa, which follows in the next commit, on a Ryzen 9 7900 (32 MB L3): two identical 16 MB pointer chases, 5 processes in both role assignments, Repeats 41, ValidationRuns 2. The pooled delta went from -29.4% to +0.35%. With the defaults it went from -0.5% to +0.34%, and the in-cache control stayed within 0.1%. The difference that remains follows the build order of the two fixtures and not the roles. A 2 s warm-up does not change it, so it is memory layout (#109). Compare now costs 0.6 s more, or 0.3 s with SkipValidation. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The command compares two identical, separately allocated pointer chases, so the true difference is zero. Its modes and flags: - -mode prefix runs Collect after validating no candidate, only A, only B, or A then B, one candidate at a time. - -mode compare runs Compare in both role assignments and prints the drift tests and warnings. - -build ab|ba|mixed selects the allocation order, which separates layout effects from role effects. Issue #111 asks for this reproduction to live in the repository, and #109 will extend the command. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This was referenced Sep 27, 2026
TomTonic
added a commit
that referenced
this pull request
Sep 27, 2026
Using the workload package correctly meant remembering several things. The first pass of a cycle had to be played untimed. A cursor had to be kept with its structure and settled afterwards. The two structures had to be built so that neither came last. And leaving out the first pass must not mean dropping growth altogether, or structures whose growth is expensive would look better than they are. - Compare(target, a, b, Options) takes two Structures, each only New and Apply, and answers two questions separately. SteadyState is the cost of one insertion or deletion, measured on a cycle after one untimed pass. Build is the cost of one whole build of a fresh structure, so creation, capacity hints, resizing and garbage all count. The build comparison has its own defaults (31 repeats, 10 validation runs, about 1,300 builds) and always sets GCBetween. - Both steady-state structures are built in alternating chunks of 256 operations, so neither is built last. - Replay bundles a cycle, the apply function and the cursor for one structure. Its candidate plays one cycle untimed in Setup before the first batch. Settle restores the start state. - Check sizes its model for the largest state the stream reaches. The example now compares map with Set3 through Compare. At 100,000 elements the build comparison resolves Set3 as 43% faster to build. On main, the steady-state comparison still shows the start-of-Collect settling that #112's warm-up removes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #111.
Problem
CompareranValidateHarness(a), thenValidateHarness(b), thenCollectwith a single warm-up pair. When the working set was near the size of the last-level cache, B started the measurement with the cache full of its own data. Short, calibrated batches did not turn the cache over during the run, so identical code came out tens of percent apart and was reported as resolved.Changes
ValidatePair(a, b, ValidationOptions)(new, exported) validates both candidates together, interleaved batch by batch in the order A, B, B′, A′ and back. In that order each half of an A/A experiment has the same history. The first version used A, B, A′, B′, where A followed itself and A′ never did; a test now checks the history balance.CompareusesValidatePair, andValidateHarnessis unchanged for direct callers.CollectOptions.WarmupDuration(new;DefaultWarmupDuration= 300 ms). The warm-up runs until both the count and the duration are reached, and alternates the order of the candidates (ABBA) instead of always running a then b. A validation pays for the duration only before its first run.Report.DriftRatiorunsDetectDrifton B/A for each pair of neighbouring batches. A warning appears only when the trend is significant and larger than the resolution, i.e. max(noise floor, half the interval width).Compare,Collect,CollectOptions,ValidateHarnessand the README.cmd/rtcompare-aa(A/A pointer chase,-mode prefix|compare,-build ab|ba|mixed), as the issue asks for. Intervals and noise floor cover only within-process noise; results scatter between processes far beyond them #109 will extend it.Internally,
Collectnow runs through aschedule(the resolved options) andmeasureInTurn(a balanced loop for any number of candidates). With two candidates that loop produces exactly the old ABBA, Random and Sequential orders.Measurements
Ryzen 9 7900 (32 MB L3), WSL2, Go 1.27.1. Two identical 16 MB fixtures, 5 processes, each in both role assignments. Delta = 1 − A/B, pooled across both role assignments:
Repeats 41,ValidationRuns 2, build abRepeats 41,ValidationRuns 2, build baWhat remains: after the fix, the fixture built second is 3–5 % faster in both roles, and the sign flips with
-build ba. A 2 s warm-up leaves that unchanged, so it comes from the memory layout of the two fixtures, not from a warm cache. That is #109, and the HOWTO now explains how to tell the two apart (swap the roles, swap the build order).The noise floors in the 41/2 case also went down, from 3–9 % to 1–4 %. For the example's candidates with
GCBetween,ValidatePairand twoValidateHarnesscalls gave the same floors within noise (6 repetitions × 40 runs each).Behaviour change
Comparetakes 0.6 s longer (0.3 s withSkipValidation), and eachCollect0.3 s longer.Warmup: Nnow also get the minimum duration.WarmupDuration: time.Nanosecondwarms up by count alone.fastComparein the tests cuts the warm-up to 10 ms, because those candidates stay in the cache.Checks
go test ./...andgo test ./... -racepass, andgolangci-lint runreports 0 issues.measureInTurnwith 3 candidates,ValidatePair(interleaving, calibration to the larger batch, errors), history balance ofpairSlots, warm-up once per validation, no long single-candidate stretches inCompare,pairRatios, and warning cases forDriftRatio.🤖 Generated with Claude Code