feat(eval): measure the context window, and name what still needs a model - #198
Merged
Merged
Conversation
…odel The harness measured the retriever but never the window the prompt actually gets. `evidenceK: 5` was a separate constant from the `contextK: 3` production uses, so the one metric that looked at a window looked at a different one than the product does. Child 5 of #192. **`contextK` is now the window for both context metrics.** - `contextPrecision@contextK` — of the first `contextK` passages, the share covering ground truth. This is what `evidencePrecisionAt5` was, at the production width. - `contextRecall@contextK` — the share of needed ground-truth blocks that made it into that window. Distinct from `Recall@10`: a block found at rank 4 is invisible when `contextK = 3`, and that is a product fact, not a ranking fact. Both are **deterministic**: the dataset says which blocks answer the question, so no model is needed to score a window. Together they are the trade-off a `contextK` decision makes — a wider window finds more and carries more noise — which is what the sweep in the next child needs. Baseline: `contextPrecision@3 = 0.3556`, `contextRecall@3 = 1.0000` — the needed evidence is always inside the top 3 on this corpus, but only about a third of what is inside the window is relevant. That second number is the one that says the window is paying for passages that do not answer the question. **What is deliberately not here.** Faithfulness, completeness, answer correctness and noise sensitivity need a generative model. The harness runs offline with only the pinned embedding model — the same constraint that keeps the reranker unmeasured (#170) — so this PR does not add them and does not fake them: the report's Definitions section now says so, and the "Not evaluated" line in `eval-retrieval.mjs` points at the same constraint. Adding an LLM judge is a separate change that has to solve model pinning first, not a line of code. `evidenceK` is removed from the harness config, so there is exactly one context width. ## Testing - `npm run typecheck` — clean - `npm test` — 497 pass - `npm run eval` twice — byte-identical `docs/eval/baseline-v1.6.json` - `npm run eval:retrieval` — regenerated; hybrid still clears the amended rule Part of #192 (child 5, deterministic half).
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #197 (
feat/eval-saturation-rule). Retarget tomainafter that merges.What does this PR do?
Measures the window the prompt actually receives, and states plainly what still needs a model. Child 5 of #192 — the deterministic half.
The gap
The harness measured the retriever but never the context window.
evidenceK: 5was a separate constant from thecontextK: 3production uses, so the one metric that looked at a window looked at a different one than the product does.The change
contextKis now the window for both context metrics:contextPrecision@contextK— of the firstcontextKpassages, the share covering ground truth. This is whatevidencePrecisionAt5was, at the production width.contextRecall@contextK— the share of needed ground-truth blocks that made it into that window. Distinct fromRecall@10: a block found at rank 4 is invisible whencontextK = 3, and that is a product fact, not a ranking fact.Both are deterministic: the dataset says which blocks answer the question, so no model is needed to score a window. Together they are the trade-off a
contextKdecision makes — a wider window finds more and carries more noise — which is what the sweep in the next child needs.Baseline:
contextPrecision@3 = 0.3556,contextRecall@3 = 1.0000. The needed evidence is always inside the top 3 on this corpus, but only about a third of what is inside the window is relevant. The second number is the one that says the window is paying for passages that do not answer the question.evidenceKis removed from the harness config, so there is exactly one context width.What is deliberately not here
Faithfulness, completeness, answer correctness and noise sensitivity are not added. They need a generative model, and the harness runs offline with only the pinned embedding model — the same constraint that keeps the reranker unmeasured (#170). This PR does not fake them. The report's Definitions section now says so explicitly, and the experiment script's "Not evaluated" note points at the same constraint.
Adding an LLM judge is a separate change that has to solve model pinning first. It is not a line of code, and shipping a judge that silently degrades to "no metrics" in CI would be worse than shipping none.
Testing
npm run typecheck— cleannpm test— 497 passnpm run evaltwice — byte-identicaldocs/eval/baseline-v1.6.jsonnpm run eval:retrieval— regenerated; hybrid still clears the amended ruleRelated
Part of #192. Child 5 (deterministic half: context precision/recall).