Skip to content

fix(eval): run the harness on Windows and keep the baseline POSIX - #193

Merged
mrsibe merged 1 commit into
mainfrom
fix/eval-windows-gates
Sep 30, 2026
Merged

mrsibe merged 1 commit into
mainfrom
fix/eval-windows-gates

Conversation

@mrsibe

@mrsibe mrsibe commented Sep 30, 2026

Copy link
Copy Markdown
Owner

What does this PR do?

Unblocks the two eval gates that make the v1.5 harness usable on Windows. No behaviour change on Linux/macOS; no metric changes.

Part of #192 — these are prerequisites, not one of its child issues: none of the evaluation work in that epic can be verified from a Windows checkout until these two are fixed.

The two bugs

1. npm run eval could not start on Windows

Error: spawn ... EINVAL
    errno: -4071, code: 'EINVAL', syscall: 'spawn'

All three eval scripts resolve node_modules/.bin/electron.cmd and spawn it. Node 24 refuses to spawn a .cmd/.bat without shell: true (the CVE-2024-27980 mitigation), so the harness was unrunnable on Windows with the Node version this repo declares (>=24 <25).

The electron package exports the path to the real executable the .bin wrapper runs, so the scripts spawn that directly — no shell, no .cmd, works on every platform.

2. A Windows run rewrote the committed baseline

-    "corpus": "eval/corpus",
+    "corpus": "eval\\corpus",

path.relative returns backslashes on Windows, so config.corpus was written with them. The file's stated contract — in the source comment, in eval/README.md, and in the CI step that diffs the JSON — is that it is identical on every machine. A Windows run silently broke that, so npm run eval on Windows could never produce a clean git diff. The label is normalised to POSIX separators.

How was this tested

  • npm run typecheck — clean
  • npm test — 478 pass
  • npm run eval on Windows with the pinned model — runs, and leaves docs/eval/baseline-v1.5.json byte-identical to the committed file, which neither bug allowed before
  • Metrics unchanged: {"recallAt1":0.833333,"recallAt5":1,"recallAt10":1,"mrr":0.927778,"ndcgAt10":0.94375,"evidencePrecisionAt5":0.213333}

docs/eval/baseline-v1.5.md differs only in the informational timing line (indexing ms, p50/p95), which is deliberately not frozen and is not diffed by CI.

Related

Part of #192.

Two gates the v1.5 eval path needs to be trustworthy on every platform.

**`npm run eval` could not start on Windows.** The three scripts resolve
`node_modules/.bin/electron.cmd` and spawn it. Node 24 refuses to spawn a
`.cmd`/`.bat` without `shell: true` and fails with `EINVAL`, so the harness was
unrunnable on the platform this project is developed on. The `electron` package
exports the path to the real executable the wrapper runs, which spawns directly
on every platform and needs no shell.

**A Windows run rewrote the committed baseline.** `path.relative` returns
backslashes on Windows, so `config.corpus` was written as `eval\corpus` instead
of `eval/corpus`. The file's whole contract is that it is identical on every
machine — the CI determinism check diffs it — and a Windows run silently broke
that. The label is now normalised to POSIX separators.

Verified on Windows with the pinned model: `npm run eval` runs and leaves
`docs/eval/baseline-v1.5.json` byte-identical to the committed file.
@github-actions github-actions Bot added the bug Something isn't working label Sep 30, 2026
@mrsibe
mrsibe merged commit 2d3c302 into main Sep 30, 2026
6 of 7 checks passed
@mrsibe
mrsibe deleted the fix/eval-windows-gates branch September 30, 2026 10:26
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant