Skip to content

Stop Compare from favouring the candidate that ran alone last - #112

Merged
TomTonic merged 3 commits into
mainfrom
fix/111-warm-start-bias
Sep 27, 2026
Merged

TomTonic merged 3 commits into
mainfrom
fix/111-warm-start-bias

Conversation

@TomTonic

Copy link
Copy Markdown
Owner

Fixes #111.

Problem

Compare ran ValidateHarness(a), then ValidateHarness(b), then Collect with a single warm-up pair. When the working set was near the size of the last-level cache, B started the measurement with the cache full of its own data. Short, calibrated batches did not turn the cache over during the run, so identical code came out tens of percent apart and was reported as resolved.

Changes

  • ValidatePair(a, b, ValidationOptions) (new, exported) validates both candidates together, interleaved batch by batch in the order A, B, B′, A′ and back. In that order each half of an A/A experiment has the same history. The first version used A, B, A′, B′, where A followed itself and A′ never did; a test now checks the history balance. Compare uses ValidatePair, and ValidateHarness is unchanged for direct callers.
  • Warm-up by duration: CollectOptions.WarmupDuration (new; DefaultWarmupDuration = 300 ms). The warm-up runs until both the count and the duration are reached, and alternates the order of the candidates (ABBA) instead of always running a then b. A validation pays for the duration only before its first run.
  • Diagnosis: Report.DriftRatio runs DetectDrift on B/A for each pair of neighbouring batches. A warning appears only when the trend is significant and larger than the resolution, i.e. max(noise floor, half the interval width).
  • Docs: a new HOWTO section "Whatever ran last starts warm", plus updates to Compare, Collect, CollectOptions, ValidateHarness and the README.
  • Reproduction: cmd/rtcompare-aa (A/A pointer chase, -mode prefix|compare, -build ab|ba|mixed), as the issue asks for. Intervals and noise floor cover only within-process noise; results scatter between processes far beyond them #109 will extend it.

Internally, Collect now runs through a schedule (the resolved options) and measureInTurn (a balanced loop for any number of candidates). With two candidates that loop produces exactly the old ABBA, Random and Sequential orders.

Measurements

Ryzen 9 7900 (32 MB L3), WSL2, Go 1.27.1. Two identical 16 MB fixtures, 5 processes, each in both role assignments. Delta = 1 − A/B, pooled across both role assignments:

Setting before after
Repeats 41, ValidationRuns 2, build ab −29.4 % (7 of 10 resolved) +0.35 %
Defaults, build ab −0.5 % +0.34 %
Repeats 41, ValidationRuns 2, build ba – +0.19 %
In-cache control (4K nodes) ±0.2 % −0.06 %

What remains: after the fix, the fixture built second is 3–5 % faster in both roles, and the sign flips with -build ba. A 2 s warm-up leaves that unchanged, so it comes from the memory layout of the two fixtures, not from a warm cache. That is #109, and the HOWTO now explains how to tell the two apart (swap the roles, swap the build order).

The noise floors in the 41/2 case also went down, from 3–9 % to 1–4 %. For the example's candidates with GCBetween, ValidatePair and two ValidateHarness calls gave the same floors within noise (6 repetitions × 40 runs each).

Behaviour change

  • Compare takes 0.6 s longer (0.3 s with SkipValidation), and each Collect 0.3 s longer.
  • Callers who set Warmup: N now also get the minimum duration. WarmupDuration: time.Nanosecond warms up by count alone.
  • Test suite: 61 s → 69 s. fastCompare in the tests cuts the warm-up to 10 ms, because those candidates stay in the cache.

Checks

  • go test ./... and go test ./... -race pass, and golangci-lint run reports 0 issues.
  • New tests (documented outside-in): warm-up duration and order, measureInTurn with 3 candidates, ValidatePair (interleaving, calibration to the larger batch, errors), history balance of pairSlots, warm-up once per validation, no long single-candidate stretches in Compare, pairRatios, and warning cases for DriftRatio.

🤖 Generated with Claude Code

TomTonic and others added 3 commits September 27, 2026 19:04
Compare validated candidate A, then candidate B, and then measured both
after a single warm-up pair. B therefore entered the measurement with the
last-level cache full of its own data. With a working set near the size
of that cache and short calibrated batches, the head start survived into
the medians: identical code measured up to 66% apart, reported as
resolved (#111).

- ValidatePair validates two candidates together, interleaved batch by
  batch in the order A, B, B', A' and back. In that order each half of
  each A/A experiment has the same history. Compare uses it.
- Collect warms up for at least CollectOptions.WarmupDuration
  (DefaultWarmupDuration, 300 ms) in addition to the batch count, and
  alternates the order of the candidates while it does. A validation
  pays for the duration once, not once per run.
- Report.DriftRatio tests the ratio B/A of neighbouring batches for a
  trend. A warning appears when the trend is significant and larger than
  the result's resolution.
- HOWTO and README describe the pitfall and the checks against it.

Measured with cmd/rtcompare-aa, which follows in the next commit, on a
Ryzen 9 7900 (32 MB L3): two identical 16 MB pointer chases, 5 processes
in both role assignments, Repeats 41, ValidationRuns 2. The pooled delta
went from -29.4% to +0.35%. With the defaults it went from -0.5% to
+0.34%, and the in-cache control stayed within 0.1%. The difference that
remains follows the build order of the two fixtures and not the roles.
A 2 s warm-up does not change it, so it is memory layout (#109).

Compare now costs 0.6 s more, or 0.3 s with SkipValidation.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The command compares two identical, separately allocated pointer chases,
so the true difference is zero. Its modes and flags:

- -mode prefix runs Collect after validating no candidate, only A, only
  B, or A then B, one candidate at a time.
- -mode compare runs Compare in both role assignments and prints the
  drift tests and warnings.
- -build ab|ba|mixed selects the allocation order, which separates
  layout effects from role effects.

Issue #111 asks for this reproduction to live in the repository, and
#109 will extend the command.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
TomTonic added a commit that referenced this pull request Sep 27, 2026
Using the workload package correctly meant remembering several things.
The first pass of a cycle had to be played untimed. A cursor had to be
kept with its structure and settled afterwards. The two structures had
to be built so that neither came last. And leaving out the first pass
must not mean dropping growth altogether, or structures whose growth
is expensive would look better than they are.

- Compare(target, a, b, Options) takes two Structures, each only New
  and Apply, and answers two questions separately. SteadyState is the
  cost of one insertion or deletion, measured on a cycle after one
  untimed pass. Build is the cost of one whole build of a fresh
  structure, so creation, capacity hints, resizing and garbage all
  count. The build comparison has its own defaults (31 repeats, 10
  validation runs, about 1,300 builds) and always sets GCBetween.
- Both steady-state structures are built in alternating chunks of 256
  operations, so neither is built last.
- Replay bundles a cycle, the apply function and the cursor for one
  structure. Its candidate plays one cycle untimed in Setup before the
  first batch. Settle restores the start state.
- Check sizes its model for the largest state the stream reaches.

The example now compares map with Set3 through Compare. At 100,000
elements the build comparison resolves Set3 as 43% faster to build. On
main, the steady-state comparison still shows the start-of-Collect
settling that #112's warm-up removes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@TomTonic
TomTonic merged commit c04f9b9 into main Sep 27, 2026
5 of 6 checks passed
@TomTonic
TomTonic deleted the fix/111-warm-start-bias branch September 27, 2026 19:19
@TomTonic
TomTonic restored the fix/111-warm-start-bias branch September 27, 2026 19:19
@TomTonic
TomTonic deleted the fix/111-warm-start-bias branch September 27, 2026 19:20
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Compare favours the candidate that ran last on its own (validation order and a one-batch warm-up)

1 participant