Skip to content

feat(eval): bounded parameter sweep with a trade-off dashboard - #199

Merged
mrsibe merged 1 commit into
feat/eval-context-metricsfrom
feat/eval-sweep-dashboard
Sep 30, 2026
Merged

mrsibe merged 1 commit into
feat/eval-context-metricsfrom
feat/eval-sweep-dashboard

Conversation

@mrsibe

@mrsibe mrsibe commented Sep 30, 2026

Copy link
Copy Markdown
Owner

Stacked on #198 (feat/eval-context-metrics). Retarget to main after that merges.

What does this PR do?

Adds a bounded parameter sweep with one dashboard showing quality, context precision/recall, prompt size, index size and latency side by side. Child 7 of #192.

npm run eval:sweep runs the real harness over strategy × candidateK {5,10,20,40} × contextK {3,5,8} — 24 runs, chunking fixed — and writes docs/eval/sweep-v1.6.{json,md}.

The harness gained contextChars per question: the size of the context window, reported as characters, not tokens, because the harness pins an embedding model and no generation tokenizer. Calling it "tokens" would be a made-up number.

What the dashboard already shows

Two of the three axes are decided by this corpus, and one of them decisively.

candidateK changes nothing. Every metric is identical from 5 to 40, because the corpus is 19 chunks and the relevant passages are already inside the top 5. This is the saturation problem #192 describes, made visible on the axis it affects.

contextK has a clear optimum here: 3.

contextK Context P Context R Context chars
3 0.3556 1.0000 2079
5 0.2133 1.0000 3327
8 0.1333 1.0000 5312

Recall is 1.0000 at every width while precision falls and the window grows. Wider adds prompt cost and noise for no recall. The production default is already 3; this is the first evidence that it is the right 3 rather than a guess.

hybrid beats dense on nDCG@10 (0.9561 vs 0.9437) and MAP@10 (0.9389 vs 0.9222) at every setting, with no metric regressing — consistent with #197.

What it deliberately does not do

The report labels the grid maximum as not a recommendation. Selecting the maximum on the same questions is how a benchmark becomes a lookup table; the adoption rule and the validation/test split are what keep the decision honest.

p95 is reported per row but is noisy at this sample size. It is in the table so a latency cost can be seen, not so it can be ranked.

Testing

  • npm run typecheck — clean
  • npm test — 497 pass
  • npm run eval — regenerated (contextChars added to perQuestion)
  • npm run eval:sweep — 24 runs, dashboard written

Related

Part of #192. Child 7 (parameter sweep + dashboard).

The epic asks for a grid over `candidateK` and `contextK` with latency, index size and
context size recorded next to quality, and explicitly not for a single aggregate
"RAG score". Child 7 of #192.

`npm run eval:sweep` runs the real harness over
`strategy × candidateK {5,10,20,40} × contextK {3,5,8}` (24 runs, chunking fixed) and
writes `docs/eval/sweep-v1.6.{json,md}`. The harness gained `contextChars` per
question — the size of the context window, reported as **characters, not tokens**,
because the harness pins an embedding model and no generation tokenizer.

## What the dashboard already shows

Two of the three axes are decided by this corpus, and one of them decisively:

- **`candidateK` changes nothing.** Every metric is identical from 5 to 40, because
  the corpus is 19 chunks and the relevant passages are already inside the top 5. This
  is the saturation problem #192 describes, now visible on the axis it affects.
- **`contextK` has a clear optimum here: 3.** Context recall is 1.0000 at every width,
  while precision falls `0.3556 → 0.2133 → 0.1333` and the window grows
  `2079 → 3327 → 5312` characters as it widens. Wider adds prompt cost and noise for
  **no** recall. The production default is already 3, and this is the first evidence
  that it is the right 3 rather than a guess.
- **hybrid beats dense** on nDCG@10 (0.9561 vs 0.9437) and MAP@10 (0.9389 vs 0.9222) at
  every setting, with no metric regressing — consistent with #197.

The report labels the grid maximum as **not** a recommendation, because selecting on the
same questions is how a benchmark becomes a lookup table; the adoption rule and the
`validation`/`test` split are what keep the decision honest.

Timing p95 is reported per row but is noisy at this sample size; it is in the table so
a latency cost can be seen, not so it can be ranked.

## Testing

- `npm run typecheck` — clean
- `npm test` — 497 pass
- `npm run eval` — regenerated (`contextChars` added to `perQuestion`)
- `npm run eval:sweep` — 24 runs, dashboard written

Part of #192 (child 7).
@github-actions github-actions Bot added the enhancement New feature or request label Sep 30, 2026
@mrsibe
mrsibe merged commit 5a2b0d1 into feat/eval-context-metrics Sep 30, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant