Skip to content

Pool comparisons across processes, since one process is one observation - #113

Merged
TomTonic merged 3 commits into
mainfrom
fix/109-between-process-scatter
Sep 27, 2026
Merged

TomTonic merged 3 commits into
mainfrom
fix/109-between-process-scatter

Conversation

@TomTonic

@TomTonic TomTonic commented Sep 27, 2026 •

Copy link
Copy Markdown
Owner

Fixes #109. Stacked on #112 (#111): this PR reuses Report.resolution(), DriftRatio and cmd/rtcompare-aa. Merge #112 first; this PR's base then switches to main.

Problem

The interval, noise floor and Resolved of a Report describe only the noise within one process. For large or pointer-heavy data, the memory layout moves the difference by several points. The layout is fixed for the lifetime of a process and differs in the next one.

Changes

  • Combine(reports, level) (Pooled, error) returns:

    • a Student t interval over the per-process deltas (k ≥ 3);
    • SpreadBetween / SpreadWithin / Inflation, i.e. how many times more the processes scatter than their own intervals imply;
    • Cochran's Q and I2;
    • warnings for too few processes, opposite-sign resolutions and suspended processes.

    The t quantile is computed in-package, checked against table values to 1e-5, so there is no new dependency. I chose the t interval over DerSimonian-Laird because DL is known to be too narrow at k ≈ 5.

  • multiproc package (new):

    • Run(Options, suite) re-executes the binary as child processes, one at a time. Children are recognised through environment variables and pass their named reports back as JSON in a file, so user output on stdout does not interfere.
    • The adaptive stop rule from the field report: min 5, max 20 processes, stop once every interval is within ±2 pp or ±10 % of |Δ|.
    • Rotation (default 2): the stop rule is only checked after complete rotations, so that alternating build orders stay balanced (see measurements).
    • Process offers Index, Seed, Record, PerturbHeap() and Rand().
    • Works in go test too, with Args: -test.run=^TestX$; the package's own tests do exactly that.
  • PerturbHeap(seed) *Spacers: seeded filler in every small size class, with and without pointers, plus one large block of up to 16 MB.

  • DPRNG.Shuffle: Fisher–Yates.

  • Report.Suspended: sleep detection via wall clock vs. monotonic clock (> 1 s), with a warning.

  • Drift warning for A/B only when |shift| > resolution (max(floor, half-width)). According to the field report it used to fire in 399 of 406 comparisons.

  • Docs: HOWTO section "One process is one observation", troubleshooting for suspend, glossary, README, and notes on Report.Estimate and EstimateDifference.

  • cmd/rtcompare-aa: -workload scan (the range scan from the issue), -perturb, -build random and -multi.

Measurements

Ryzen 9 7900, WSL2. A/A comparison of identical fixtures (true difference = 0), -workload scan, defaults.

Setup Processes Pooled result (role 1 / role 2) Inflation I²
1M, fixed layout (build ab) 8 −2.95 % / +3.08 %, all 8 with the same sign, 7 of 8 resolved – –
1M, PerturbHeap + random build order 8 −1.69 % [−3.45, +0.08] / +1.30 % [−0.55, +3.15] 2.9–3.3× 0.90
1M, PerturbHeap + alternating build order 14 +0.29 % [−1.59, +2.17] / −0.21 % [−2.16, +1.75] 4.5× 0.96
4K, PerturbHeap + alternating 6 −0.46 % [−1.31, +0.38] / +0.38 % [−0.37, +1.13] 1.7–1.8× 0.68–0.75

What the measurements show:

  • Single processes mislead. In the alternating 1M run, 25 of 28 single-process comparisons were resolved at ±3–4 %, with both signs. The pooled intervals cover 0. The factors (4.5× out of cache, about 2× in cache) match the observations in the issue.
  • Build order dominates. Whichever list is built second is about 3 % faster, in every process and regardless of PerturbHeap. That is the flip from the issue ("A before or after B: −2.2 % vs +2.2 %"). A random order over 8 processes produced 6:2, so the pooled result leaned. Hence the docs recommend alternating by p.Index for two fixtures, and Rotation keeps the stop rule balanced.
  • What PerturbHeap does: on its own it removed no measurable bias in this reproduction. It stays in because it breaks the determinism of the layout, which the process results would otherwise repeat. The HOWTO says honestly that the build order matters more.

Not addressed (as agreed)

  • A noise floor across processes (direction 5 in the issue); the pooled interval already covers that variance.
  • Quantization warnings at ~50 ns/op.
  • Actively keeping the machine awake (caffeinate etc.); it is documented, and sleep is detected.

Checks

  • go test ./... and go test ./... -race pass, and golangci-lint run reports 0 issues.
  • New tests (documented outside-in):
    • t quantiles against table values;
    • Combine on scattered and agreeing processes, its warnings and its input errors;
    • Precise, heterogeneity without SE;
    • Shuffle: permutation, determinism, uniformity;
    • PerturbHeap: determinism and bounds;
    • suspend gap;
    • drift threshold;
    • multiproc end to end with real child processes: pooling, seed reproducibility, stop rule, Rotation, child errors, option checks.

Addendum: multiproc takes care of the precautions itself (f23243b)

Correct use of multiproc required several things to be remembered: PerturbHeap + KeepAlive, build order via p.Index%2, returning early on res.Child, and -test.run in tests. Each forgotten step quietly weakens the result. multiproc now does all of it on its own:

  • Heap perturbation: Run perturbs the heap in every child before the suite runs, and keeps the spacers alive. Process.PerturbHeap has been removed.
  • Pair{Name, A, B func() rtcompare.Candidate, Options}: the driver calls the builders itself, in alternating order (A first in even processes, B first in odd ones), then compares and records the result.
  • Main(opt, pairs...): a complete benchmark program with progress output and the pooled result. The child exits by itself.
  • RunTest(t, opt, pairs...): children run only the calling test. In the child the rest of the test is skipped, and errors → t.Fatal.
  • Report.LiveHeap + warning: above LargeHeapThreshold (16 MB of live data, via runtime/metrics), Compare warns that one process is not enough and points to multiproc.

A benchmark program is now just:

multiproc.Main(multiproc.Options{}, multiproc.Pair{Name: "lookup", A: buildA, B: buildB})

The HOWTO section shrinks from "remember X" to "this happens automatically". Tests (go test ./..., -race) and lint are green. I also merged all three branches locally and tested them together: go test ./... is green there too.

🤖 Generated with Claude Code

TomTonic and others added 2 commits September 27, 2026 20:24
A single Report's interval covers the noise within one process. Each
process lays its data out differently, and for large or pointer-heavy
data that layout moves the difference by several points, far beyond
the interval, and differently in the next process (#109).

- Combine pools per-process reports into a Pooled result. It gives a
  Student t interval over the per-process deltas, the scatter between
  processes against the scatter within one (Inflation), Cochran's Q and
  I², and warnings when processes resolve with opposite signs. The t
  quantile is computed in-package, so there is no new dependency.
- The multiproc package re-executes the binary as child processes, one
  at a time, and collects each child's named reports through a file,
  not stdout. It pools them and applies an adaptive stop rule: at least
  5 processes, at most 20, until every pooled interval is within ±2
  points or ±10% of |delta|. The rule is only checked after a whole
  Rotation (default 2), so alternating build orders stay balanced.
- PerturbHeap allocates seeded filler in every small size class and one
  large block, so that each process gets its own heap layout.
  DPRNG.Shuffle permutes build orders.
- Report.Suspended detects a machine that slept during Compare, from
  the gap between the wall clock and the monotonic clock.
- A drift warning for a single candidate now needs the shift to exceed
  the result's resolution as well as being significant. It used to
  fire on almost every run.
- HOWTO, README, Report and EstimateDifference say what the interval
  covers and how to pool.

Measured with cmd/rtcompare-aa -workload scan (100 consecutive elements
of a linked list whose nodes were allocated in random order) on a Ryzen
9 7900, comparing identical fixtures, so the true difference is zero:

- 1M nodes, fixed layout, 8 processes: every process put the same sign
  on the difference, -2.95% and +3.08% in the two roles, 7 of 8
  resolved. Whichever list was built second was faster, and
  PerturbHeap did not change that.
- 1M nodes, perturbed heap and build order alternating by process: 14
  processes, pooled +0.29% [-1.59, +2.17] and -0.21% [-2.16, +1.75].
  The processes scattered 4.5 times as widely as their own intervals
  (I² 0.96). 25 of 28 single-process results were resolved at ±3-4%,
  with both signs.
- 4K nodes, same setup: 6 processes, pooled within ±0.5%, scatter 1.7
  to 1.8 times the single-process interval.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- -workload scan walks 100 consecutive elements of a linked list with
  randomly placed nodes. It is the range scan from #109.
- -perturb applies PerturbHeap before the fixtures are built.
- -build random picks ab or ba from -seed.
- -multi runs the comparison in several processes through multiproc,
  perturbing each process's heap and alternating the build order.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Using multiproc correctly meant remembering to perturb the heap, keep
the spacers alive, alternate the build order by process index, return
early in a child and restrict a test's children to that test. Each one
forgotten silently weakens the result, so the package now takes care
of all of them.

- Run perturbs the heap from the process seed in every child before
  the suite runs, and keeps the spacers alive until the suite is done.
  Process.PerturbHeap is gone.
- Pair describes a comparison by two builder functions. Pairs, Main and
  RunTest build them in an order that alternates with the process
  index, compare them and record the report.
- Main is a whole benchmark program. It prints progress and the pooled
  results, and exits in a child.
- RunTest restricts the children to the calling test, skips the rest
  of the test in a child and fails the test on errors.
- Report.LiveHeap records the live heap from runtime/metrics. Above
  LargeHeapThreshold (16 MB) Compare warns that one process is not
  enough and points to multiproc.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@TomTonic
TomTonic deleted the branch main September 27, 2026 19:19
@TomTonic TomTonic closed this Sep 27, 2026
@TomTonic TomTonic reopened this Sep 27, 2026
@TomTonic
TomTonic changed the base branch from fix/111-warm-start-bias to main September 27, 2026 19:19
@TomTonic
TomTonic merged commit 7590fe8 into main Sep 27, 2026
1 check passed
@TomTonic
TomTonic deleted the fix/109-between-process-scatter branch September 27, 2026 19:20
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Intervals and noise floor cover only within-process noise; results scatter between processes far beyond them

1 participant