Build performance benchmarks from your own account history instead of borrowing an industry average — so you stop investigating noise and start catching real change.
flowchart LR
A["<b>1 · Your own history</b><br/>12+ periods of one metric<br/>per channel"]
B["<b>2 · The builder</b><br/>clean → baseline → control<br/>limits → run rules"]
C["<b>3 · A baseline register</b><br/>normal / watch / signal zones<br/>and when to recalibrate"]
A --> B --> C
style A fill:#f6f8fa,stroke:#57606a,color:#1f2328
style B fill:#ddf4ff,stroke:#0969da,color:#1f2328
style C fill:#dafbe1,stroke:#1a7f37,color:#1f2328
A month at 0.71% CTR is normal against your own baseline. Against a borrowed industry benchmark of 0.45% it reads as a 58% outperformance, and someone spends a week working out what they did right.
Built on the Historical Performance Benchmarking Framework.
- Put
SKILL.mdin your AI tool's instructions field — a Claude Project, a Custom GPT, a Gemini Gem, or an API system prompt. - Put
reference/volatility-multipliers.mdinto its knowledge base or file uploads. - Give it 12 or more consecutive periods of one metric, for one channel.
| You give it | 12+ consecutive periods of history for one metric, one channel, with volume per period |
| You get back | A baseline register — mean, control limits, three zones, run rules, and when to recalibrate |
Missing a required input, it asks instead of guessing.
- Picks a window of genuinely stable history, starting after any structural change — a baseline spanning a replatform is measuring two different businesses
- Excludes one-time events, and tells you what each exclusion did to the mean rather than quietly applying it
- Sets control limits using a multiplier matched to how volatile that metric actually is
- Adds run rules, which catch the slow deterioration that control limits miss — and slow deterioration is what actually kills accounts
- Runs a leakage test to confirm no ad's own performance is feeding the benchmark it gets scored against
Three zones, and the important one is Normal. Most of what marketing teams investigate is noise, and the cost isn't the investigation — it's the changes made in response to it.
There's also a protocol for a brand with no history at all: use an industry anchor for four weeks, clearly labeled provisional, build your own baseline over the next eight, and only then turn control limits on. A provisional benchmark presented like an established one gets acted on as though it were.
| File | What it is |
|---|---|
SKILL.md |
The skill. This is the thing you paste in |
reference/volatility-multipliers.md |
How wide the limits should be per metric type, and why |
marketing-measurement-validator → run first · ad-performance-scorecard-builder → uses this
MIT © Raneq Barber