Repository navigation
feat(eval): freeze a reproducible v1.4 retrieval baseline - #115
Merged
Merged
Conversation
There was no way to answer "did this change make retrieval better?", which makes every v1.5 hybrid/reranker/chunking claim folklore. This adds the harness and the reference point they must be measured against. - `eval/corpus` + `eval/questions.jsonl`: a first-party corpus with ground truth. Relevance is expressed in **corpus identity** (`document` relative path + `block` ordinal + optional `quote`), never a runtime `documentId`/`blockId`. Runtime ids are random and `blockId` embeds them, so a dataset keyed on them would break — rather than measure — the chunking change #78 is about. - `src/main/eval/`: metrics (pure, unit-tested), the harness, and the report renderer. It indexes the corpus through the normal ingestion path and runs the real `Retriever`, so it measures the shipped pipeline, not a shortcut. - Runs as a main-process entry (`--eval-harness`) like the packaged smoke test, because the DB layer, vector store and loaders do not exist outside Electron. It uses a throwaway profile and never opens the developer's database. - `npm run eval:prepare` is the only networked step: it bootstraps the pinned `multilingual-e5-small` revision. `npm run eval` is offline, refuses to download, and forces the local backend regardless of the developer's embedding config. - The committed JSON excludes wall-clock timing, so two runs are byte-identical and a PR can diff them. Latency p50/p95 is reported in the markdown as informational. - CI caches the pinned model and asserts two consecutive runs produce identical output. Baseline on this machine: Recall@1 0.9231, Recall@5 1.0000, MRR 0.9487, nDCG@10 0.9615. Refs #75 (frozen baseline, absorbed from #76)
…he dataset Two corrections to the benchmark before it is frozen. 1. `citationRecall` was not recall and not about answers: it was the share of the top-k retrieved passages covering ground truth, i.e. retrieval precision. It also never ran a model, so the old name misdescribed what was measured. Rename it to `evidencePrecisionAt5` (`evidencePrecisionAtK`), document that each retrieved passage counts once however many blocks it covers, and leave answer-level citation correctness to the resolver (#70). A model-driven answer eval is a separate deliverable. 2. The dataset saturated: 12 of 13 questions were rank 1 and Recall@5 was 1.0, which makes the frozen adoption rule ("Recall@5 improves and nDCG@10 does not regress") unsatisfiable. Add six near-duplicate distractor documents (lake vs river monitoring, supercapacitors vs lithium cells, wild vs hive pollinators, green roofs vs street canopy, vinegar vs lactic fermentation, wave vs tidal energy) and a Chinese document, and grow to 30 questions: similar-document discrimination, multi-location ground truth, paraphrases, Chinese questions and a Chinese question over an English source. New baseline: Recall@1 0.7667, Recall@5 0.9333, Recall@10 1.0000, MRR 0.8492, nDCG@10 0.8899, evidence precision@5 0.2000. Two runs are byte-identical. Refs #75
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What does this PR do?
Adds the RAG eval harness and commits the frozen v1.4 retrieval baseline. It indexes a first-party corpus through the normal ingestion path, runs the real
Retriever, and writes a deterministic metrics JSON plus a markdown summary.npm run evalis offline;npm run eval:prepareis the one-time networked model bootstrap.Why?
There was no way to answer "did this change make retrieval better?", so every future v1.5 claim about BM25, fusion, rerankers or chunking would be folklore. #75 is the reference point #77 and #78 must be measured against (and absorbs #76's "freeze the baseline" DoD).
Related issue
Fixes #75
What changed?
eval/corpus/(13 first-party markdown documents) +eval/questions.jsonl(30 questions) +eval/README.md.src/main/eval/metrics.ts— pure Recall@1/5/10, MRR, nDCG@10, evidence precision, percentile; unit-tested.src/main/eval/harness.ts/run.ts/report.ts/types.ts— ingestion, ground-truth resolution, retrieval, report.npm run eval:prepareandnpm run eval;scripts/eval.mjslauncher.docs/eval/baseline-v1.4.{json,md}— the frozen report.docs/architecture.md/CONTRIBUTING.md— document the measurement seam and commands.Design notes
{ document: "river-monitoring.md", page: null, block: 4, quote: "..." }. RuntimedocumentIds are random andblockIds embed them, so a dataset keyed on them would break instead of measure a chunking change. The runner maps corpus → runtime ids after ingestion; production id generation is untouched. The optionalquotefails the run if a parser change moves the block ordinal.evidencePrecisionAt5, not citation recall. The harness runs no model and produces no answer, so a metric named "answer citation recall" would be a false claim. It is the share of the first 5 retrieved passages that cover ground truth — retrieval precision. Answer-level citation correctness stays with the resolver ([Feat] Resolve and validate citations; never render an ungrounded [n] #70); a model-driven answer eval is a separate deliverable.eval:preparevsevalsplit: a clean checkout meanseval:prepare && eval— no user model/provider configuration, but not "no download".evalrefuses to download and forces the local backend, so a developer's remote embedding config cannot leak into the baseline.--eval-harness): the DB layer, vector store and loaders do not exist outside Electron, and a plain Node runner would have to reimplement the ingestion path it is supposed to measure. It uses a throwaway profile and never opens the developer's database.Frozen numbers
Recall@5 has room to move, so the adoption rule ("adopt only if Recall@5 improves and nDCG@10 does not regress") is now satisfiable.
How was this tested?
npm run typecheck— passes.npm test— passes (metrics cases updated for the rename).npm run build— passes.node scripts/eval.mjstwice: byte-identicalbaseline-v1.4.json.npm run eval:prepare— no-op with a warm cache;evalwith a cold cache exits 1 with the documented message.npm run build:unpackandnpm run smoke:packaged— pass (18 checks).Not verified: cross-machine reproducibility. The committed baseline was produced on the author's Linux machine; CI asserts two consecutive runs on ubuntu agree, not that they equal this file. If the first CI run disagrees, the committed snapshot should be regenerated on CI's platform.
Screenshots / recordings
Not applicable (no UI).
Checklist
npm run typecheckpasses.npm run buildpasses.Desktop / build changes
npm run build:unpackpasses.npm run smoke:packagedpasses.