Skip to content

fix(eval): refuse a context window wider than the retrieval depth - #201

Merged
mrsibe merged 1 commit into
feat/eval-sweep-dashboardfrom
fix/eval-sweep-context-window-invariant
Sep 30, 2026
Merged

mrsibe merged 1 commit into
feat/eval-sweep-dashboardfrom
fix/eval-sweep-context-window-invariant

Conversation

@mrsibe

@mrsibe mrsibe commented Sep 30, 2026

Copy link
Copy Markdown
Owner

Stacked on #199 (feat/eval-sweep-dashboard). Retarget to main after the chain merges.

What does this PR do?

Refuses a context window wider than the retrieval depth, instead of reporting it as if it were measured. Found in the #192 review.

The bug

The harness retrieves candidateK passages and the context metrics look at the first contextK of them. candidateK=5, contextK=8 therefore reports on five passages while claiming eight. The dashboard showed:

candidateK contextK Context chars Context P
5 5 3327 0.2133
5 8 3327 0.2133

Identical — "8 is as good as 5", apparently. It was not a measurement at all: passages 6–8 did not exist. evidencePrecisionAtK divides by the passages actually retrieved, so nothing exposed it either.

The fix

A wrong number that looks like a measurement is worse than a failure, so the invariant is enforced rather than documented.

  • The harness refuses it. contextK > candidateK throws before indexing, naming both numbers:
    contextK (8) cannot exceed candidateK (5): the harness retrieves candidateK passages, so a wider context window can never be filled.
  • The sweep skips those cells and says so. The report lists them in a "Skipped cells" section rather than quietly omitting rows, because "we did not measure this" and "this measured the same as its neighbour" are different statements — and the old output showed the second while meaning the first.

docs/eval/sweep-v1.6.* is regenerated: 22 rows instead of 24, and the misleading candidateK=5 / contextK=8 row is gone rather than silently equal to 5.

The production cells (candidateK=20, contextK ∈ {3,5,8}) are unaffected.

Testing

  • npm run typecheck — clean
  • npm test — 502 pass, including 3 new harness-invariant tests that need no vector store, because the check runs before any database work: the refusal, the message naming both numbers, and contextK == candidateK being accepted
  • electron . --eval-harness --eval-candidate-k=5 --eval-context-k=8 — refused with the message above
  • npm run eval:sweep — 22 rows, 2 skipped and listed

Related

Part of #192 (review follow-up).

The sweep contained cells the harness cannot fill. The harness retrieves `candidateK`
passages and the context metrics look at the first `contextK` of them, so
`candidateK=5, contextK=8` reports on five passages while claiming eight. The dashboard
showed `contextK=5` and `contextK=8` at `candidateK=5` as **identical** — and because
`evidencePrecisionAtK` divides by the passages actually retrieved, nothing exposed it.

Found in the #192 review. A wrong number that looks like a measurement is worse than a
failure, so the invariant is enforced rather than documented:

- **The harness refuses it.** `contextK > candidateK` throws before indexing, naming
  both numbers:
  `contextK (8) cannot exceed candidateK (5): the harness retrieves candidateK
  passages, so a wider context window can never be filled.`
- **The sweep skips those cells** and says so. The report lists the skipped
  combinations in a "Skipped cells" section instead of quietly omitting rows, because
  "we did not measure this" and "this measured the same as its neighbour" are different
  statements and the old output showed the second while meaning the first.

`docs/eval/sweep-v1.6.*` is regenerated: 22 rows instead of 24, and the misleading
`candidateK=5 / contextK=8` row is gone rather than silently equal to `5`.

The production cells (`candidateK=20, contextK ∈ {3,5,8}`) are unaffected.

## Testing

- `npm run typecheck` — clean
- `npm test` — 502 pass, including 3 new harness-invariant tests that need no vector
  store, because the check runs before any database work: the refusal, the message
  naming both numbers, and `contextK == candidateK` being accepted
- `node_modules/electron/dist/electron.exe . --eval-harness --eval-candidate-k=5
  --eval-context-k=8` — refused with the message above
- `npm run eval:sweep` — 22 rows, 2 skipped and listed

Part of #192 (review follow-up).
@github-actions github-actions Bot added the bug Something isn't working label Sep 30, 2026
@mrsibe
mrsibe merged commit fb5c6e6 into feat/eval-sweep-dashboard Sep 30, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant