Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
113 changes: 113 additions & 0 deletions HOWTO.md
Original file line number Diff line number Diff line change
Expand Up @@ -453,6 +453,96 @@ The practical checks, for any comparison of large data:
- **Do not validate candidates one at a time** before comparing them. Use
`ValidatePair`, or `Compare`, which does.

## One process is one observation

Everything `Compare` reports — the interval, the noise floor, `Resolved` —
describes the noise **within one run of your program**. Run the same program
again and you get a new process, and a new process lays its data out in memory
differently: different addresses, different pages, different cache sets. For
small data that barely matters. For data larger than the caches, or full of
pointers (trees, linked structures, maps of heap objects), it can move the
difference between A and B by several points, and that shift is fixed for the
whole life of the process. No amount of `Repeats` inside the process sees it,
and the A/A validation cannot either, because both halves of an A/A run share
the same layout.

How big this gets was measured on a 1M-key ordered index
([issue #109](https://github.com/TomTonic/rtcompare/issues/109)): the same
comparison, repeated in separate processes, scattered 4 to 10 times more
widely than each process's own interval said it should. One process reported
`+14.3% [+13.2, +15.2]`, the next `+24.8%`. Merely adding an unrelated
structure to the fixture build moved another comparison from `+2.9%` to
`-27.3%`, both resolved with narrow intervals. Even for data that fits in the
cache, the intervals were about twice too narrow.

So the rule is: **when your data is large or pointer-heavy, one process is one
observation.** Run several and pool them. `Compare` reminds you: when the
program holds more than 16 MB of live data, its warnings say so.

### How

The `multiproc` package does all of it. You say how to build each candidate;
it does the rest:

```go
func main() {
multiproc.Main(multiproc.Options{}, multiproc.Pair{
Name: "lookup",
A: func() rtcompare.Candidate { return lookupIn(buildTreeA()) },
B: func() rtcompare.Candidate { return lookupIn(buildTreeB()) },
})
}
```

In a test, `multiproc.RunTest(t, multiproc.Options{}, pair)` does the same and
returns the pooled results for your assertions.

What happens behind that call, so that you don't have to remember any of it:

- **Your program is started again as child processes**, one after another, so
they don't disturb each other. In a test, the children run only that test.
- **Each child gets a different heap layout.** Before anything is built, it
fills the heap with a seeded amount of filler, so that your data lands at
different addresses in each process (`rtcompare.PerturbHeap`). Otherwise
every process would repeat the same layout, and its bias with it.
- **The build order alternates.** Even-numbered processes build A's data
first, odd-numbered ones B's. This matters more than the heap: in the
reproduction in `cmd/rtcompare-aa`, whichever of two identical 1M-node lists
was built second was about 3% faster in every process, however the heap was
perturbed. That is also why the builders are functions: they have to run
inside each process, after the perturbation, in the right order. Build
everything the measurement depends on inside them.
- **It stops when the answer is precise enough.** At least 5 processes, then
more until every comparison's pooled interval is within ±2 percentage
points, or within ±10% of the difference itself, and at most 20. It only
stops after an even number, so both build orders count equally. In the
measurements behind issue #109, five were enough for data in the cache and
13 to 15 were needed for data far out of it.

Keep the machine awake while this runs: a laptop that goes to sleep pauses the
measurement for as long as it sleeps. rtcompare notices that
(`Report.Suspended`) but cannot give you the time back.

For anything the pairs don't cover, `multiproc.Run` takes a suite function of
your own, and still perturbs the heap and pools for you; alternate the build
order by `p.Index` yourself there. If you run the processes by some other
means, `rtcompare.Combine(reports, 0)` does the pooling part on its own.

### Reading a pooled result

`Combine` treats each process as one number, its delta, and puts a Student t
interval around their mean. With few processes that interval is wide, and it
should be: five observations are five observations. Next to it you get:

- **`Inflation`** — how many times more the processes scatter than one
process's interval implies. Near 1, a single process would have told you the
truth. Well above 1, it would not have, and the pooled result is the one to
quote.
- **`I2`** — the share of the scatter that the per-process intervals do not
explain. Above about 0.5, the differences between processes dominate.
- **Warnings** — among them, when processes resolved the difference with
opposite signs, each of them confident.

## Troubleshooting: what to do, when

This is the part of the "long protocol" that's normally invisible — the
Expand Down Expand Up @@ -526,6 +616,11 @@ whether the result keeps its sign.
**Symptom:** `DriftReport.Drifted(...)` returns true, or `Report.Warnings`
mentions a candidate that "drifted during the run."

This warning only appears when the trend is both significant and larger than
what the result can resolve (the noise floor, or half the interval, whichever
is larger). A long run finds shifts of a tenth of a percent significant, and a
warning that fires on nearly every run tells you nothing.

**What it means:** the machine changed behavior over the course of the run —
usually getting slower, most often from thermal throttling as sustained load
heats up the CPU. Because measurement order is interleaved by default (ABBA),
Expand All @@ -542,6 +637,20 @@ beyond what the reported interval shows.
result from it with a bit more skepticism than the headline confidence
suggests.

### A warning says the machine was suspended

**Symptom:** `Report.Suspended` is non-zero, or a warning says the machine
"appears to have been suspended."

**What it means:** the wall clock moved further than the monotonic clock
during the comparison, which on Linux and macOS happens when the machine
sleeps. The samples straddle a pause, possibly a long one, after which caches,
clock speeds and everything else started cold.

**What to do:** repeat the comparison with the machine kept awake — plugged
in, and with `caffeinate -i` on macOS, `systemd-inhibit --what=idle:sleep` on
Linux, or the power settings on Windows.

### The autocorrelation is high

**Symptom:** `HarnessValidation.Autocorrelation` (or `Report.Autocorrelation`)
Expand Down Expand Up @@ -656,6 +765,10 @@ package-level sink variable.
- **Autocorrelation** — how much each measurement resembles the one right
before it. High values mean measurements aren't fully independent of one
another.
- **Layout effect** — the part of a measured difference that comes from where
the data happens to lie in memory in this particular process, not from the
code. It is fixed for the life of a process and changes with the next one;
see [One process is one observation](#one-process-is-one-observation).
- **Attenuation** — the true difference between two pieces of code getting
diluted in the measured number because of fixed overhead (a loop, an
accumulator) that both candidates carry equally. See
Expand Down
13 changes: 12 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -23,6 +23,7 @@ Keywords: benchmarking, performance, bootstrap, runtime comparison, statistics,
- Measure the harness against itself, so that a result can be compared with the difference the same setup reports between two runs of identical code.
- Detect a trend across a measurement run, which resampling structurally cannot see because it discards the order the samples arrived in.
- Resample in blocks when the measurements are correlated enough that treating them as independent would overstate confidence.
- Run a comparison in several processes, each with its own heap layout, and pool the results into an interval that covers the scatter between processes (`multiproc`, `Combine`, `PerturbHeap`).
- Deterministic PRNG for reproducible inputs, and a crypto/rand-backed one where unpredictability is wanted.

## What this cannot tell you
Expand All @@ -31,6 +32,8 @@ Two limits are worth knowing before the first run, because neither is visible in

**Attenuation.** What is measured is the loop, not the function. Whatever fixed per-operation cost the batch body carries — the loop itself, an accumulator, regenerating an input the candidate mutates — is present in both candidates and shrinks the difference between them. In a controlled experiment where the true difference was exactly 50%, the measured difference was 35%, because 1.81 ns/op of loop overhead sat on top of 2.13 ns/op of real work. Subtracting an empty-loop baseline does not repair it: the compiler optimizes an empty loop differently, and that correction recovered 2 of the 15 missing percentage points. Read a result as the speedup of the measured region, not of the isolated function.

**The process.** Every interval a single run reports covers the noise within that one process. A process fixes its memory layout for its whole life, and for data larger than the caches or full of pointers that layout alone has moved a difference by several points — 4 to 10 times the reported interval, and sometimes past zero. Where that applies, one process is one observation: run several with the `multiproc` package and read the pooled result. See [One process is one observation](HOWTO.md#one-process-is-one-observation).

**The noise floor.** Resampling quantifies how much an estimate would move if the same measurements were drawn again. It cannot see a bias that affected all of them equally, and will report a tight confidence around one. Measured on identical code, this package has seen apparent differences from a few tenths of a percent to well over one, carried with high confidence. `ValidateHarness` exists to measure that floor for your machine and your options; a result below it has resolved nothing, however confident the number looks.

## Install
Expand Down Expand Up @@ -156,10 +159,18 @@ Checking the measurement itself:
- `ValidateHarness(candidate, ValidationOptions)` — runs a candidate against itself and reports the noise floor, the tie rate, the drift rate and the autocorrelation. `Resolves(difference)` answers whether a result clears that floor. The floor is the 90th percentile of the differences observed on identical code, not their maximum, so that it converges as you validate longer instead of growing; roughly one A/A run in ten exceeds it. Validate both candidates and use the worse floor.
- `ValidatePair(a, b, ValidationOptions)` — validates two candidates together, interleaved batch by batch, and returns one `HarnessValidation` each. Use it instead of two `ValidateHarness` calls before comparing the two: a candidate validated on its own last would start the comparison with a warm cache.
- `DetectDrift(samples)` — tests a sample series for a trend across the run.
- `Report.Suspended` — how long the machine slept during a `Compare`, from the gap between the wall clock and the monotonic clock.
- `Report.LiveHeap` — the program's live data; above 16 MB, `Compare` warns that one process is not enough and points to `multiproc`.

Across processes:

- `multiproc.Main(Options, pairs...)` / `multiproc.RunTest(t, Options, pairs...)` — the whole job in one call: each `Pair` says how to build candidate A and B, and the program is re-executed as child processes, one at a time, each with its own heap layout and with the build order alternating, until every pooled interval is within ±2 points or ±10% of the difference (at least 5, at most 20 processes). `multiproc.Run(Options, suite)` does the same for a suite function of your own.
- `Combine(reports, level)` — pools per-process reports of one comparison into a `Pooled` result: a t interval over the per-process deltas, plus how far the processes scatter beyond their own intervals (`Inflation`, Cochran's Q, I²).
- `PerturbHeap(seed)` — allocates seeded filler in every small size class and one large block, so that data built afterwards lands at different addresses in each process.

Primitives:

- `DPRNG` / `CPRNG` — deterministic and cryptographic generators with `Uint64`, `Float64` and `Uint32N`.
- `DPRNG` / `CPRNG` — deterministic and cryptographic generators with `Uint64`, `Float64` and `Uint32N`. `DPRNG.Shuffle` permutes, e.g. the order in which fixtures are built.
- `SampleTime()` / `DiffTimeStamps()` — high-resolution timestamps, and `GetSampleTimePrecision()` for the smallest interval they can resolve here.
- `Median` / `QuickMedian` / `Statistics` — small statistics helpers.

Expand Down
Loading
Loading