feat(eval): bounded parameter sweep with a trade-off dashboard - #199
Merged
Merged
Conversation
The epic asks for a grid over `candidateK` and `contextK` with latency, index size and context size recorded next to quality, and explicitly not for a single aggregate "RAG score". Child 7 of #192. `npm run eval:sweep` runs the real harness over `strategy × candidateK {5,10,20,40} × contextK {3,5,8}` (24 runs, chunking fixed) and writes `docs/eval/sweep-v1.6.{json,md}`. The harness gained `contextChars` per question — the size of the context window, reported as **characters, not tokens**, because the harness pins an embedding model and no generation tokenizer. ## What the dashboard already shows Two of the three axes are decided by this corpus, and one of them decisively: - **`candidateK` changes nothing.** Every metric is identical from 5 to 40, because the corpus is 19 chunks and the relevant passages are already inside the top 5. This is the saturation problem #192 describes, now visible on the axis it affects. - **`contextK` has a clear optimum here: 3.** Context recall is 1.0000 at every width, while precision falls `0.3556 → 0.2133 → 0.1333` and the window grows `2079 → 3327 → 5312` characters as it widens. Wider adds prompt cost and noise for **no** recall. The production default is already 3, and this is the first evidence that it is the right 3 rather than a guess. - **hybrid beats dense** on nDCG@10 (0.9561 vs 0.9437) and MAP@10 (0.9389 vs 0.9222) at every setting, with no metric regressing — consistent with #197. The report labels the grid maximum as **not** a recommendation, because selecting on the same questions is how a benchmark becomes a lookup table; the adoption rule and the `validation`/`test` split are what keep the decision honest. Timing p95 is reported per row but is noisy at this sample size; it is in the table so a latency cost can be seen, not so it can be ranked. ## Testing - `npm run typecheck` — clean - `npm test` — 497 pass - `npm run eval` — regenerated (`contextChars` added to `perQuestion`) - `npm run eval:sweep` — 24 runs, dashboard written Part of #192 (child 7).
This was referenced Sep 30, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #198 (
feat/eval-context-metrics). Retarget tomainafter that merges.What does this PR do?
Adds a bounded parameter sweep with one dashboard showing quality, context precision/recall, prompt size, index size and latency side by side. Child 7 of #192.
npm run eval:sweepruns the real harness overstrategy × candidateK {5,10,20,40} × contextK {3,5,8}— 24 runs, chunking fixed — and writesdocs/eval/sweep-v1.6.{json,md}.The harness gained
contextCharsper question: the size of the context window, reported as characters, not tokens, because the harness pins an embedding model and no generation tokenizer. Calling it "tokens" would be a made-up number.What the dashboard already shows
Two of the three axes are decided by this corpus, and one of them decisively.
candidateKchanges nothing. Every metric is identical from 5 to 40, because the corpus is 19 chunks and the relevant passages are already inside the top 5. This is the saturation problem #192 describes, made visible on the axis it affects.contextKhas a clear optimum here: 3.Recall is 1.0000 at every width while precision falls and the window grows. Wider adds prompt cost and noise for no recall. The production default is already 3; this is the first evidence that it is the right 3 rather than a guess.
hybrid beats dense on nDCG@10 (0.9561 vs 0.9437) and MAP@10 (0.9389 vs 0.9222) at every setting, with no metric regressing — consistent with #197.
What it deliberately does not do
The report labels the grid maximum as not a recommendation. Selecting the maximum on the same questions is how a benchmark becomes a lookup table; the adoption rule and the
validation/testsplit are what keep the decision honest.p95 is reported per row but is noisy at this sample size. It is in the table so a latency cost can be seen, not so it can be ranked.
Testing
npm run typecheck— cleannpm test— 497 passnpm run eval— regenerated (contextCharsadded toperQuestion)npm run eval:sweep— 24 runs, dashboard writtenRelated
Part of #192. Child 7 (parameter sweep + dashboard).