Skip to content

feat(provenance): stabilise source identity, refine PDF blocks, extract the Retrieval seam - #111

Merged
mrsibe merged 5 commits into
mainfrom
feat/v1.4-pre-citation-provenance
Sep 25, 2026
Merged

mrsibe merged 5 commits into
mainfrom
feat/v1.4-pre-citation-provenance

Conversation

@mrsibe

@mrsibe mrsibe commented Sep 25, 2026 •

Copy link
Copy Markdown
Owner

What does this PR do?

Closes the provenance work that must land before structured citations (#69), in the order the roadmap calls for:

  1. reindexDocument() rebuilds the derived index in place instead of importing a second document.
  2. PDF blocks get paragraph granularity, and the paragraph bbox reaches RetrievedEvidence, so a citation can point at a paragraph (and highlight it) rather than only a page.
  3. Retrieval gets a Retriever seam that returns RetrievedEvidence[] with provenance attached, and provenance hydration becomes batch.
  4. Re-indexing a pre-upgrade document lazily recovers its parse structure instead of downgrading its provenance.

Why?

  • Re-index was not a re-index. It deleted the old index, marked the old document processing forever, then called addDocument(), which minted a new documentId and dropped localFilePath. Once [Feat] Structured citations through retrieval → prompt → chat_messages.metadata.citations[] #69 persists a documentId in a citation, that silently repoints every historical answer at a different source — the exact thing the provenance epic exists to prevent.
  • PDF provenance stopped one level too coarse. PdfLoader joined every TextItem on a page into one string, so the data could say "page 12" but never "the paragraph at the bottom of page 12".
  • The bbox was stored but thrown away. document_blocks.bbox held the geometry, but the provenance join never selected it, so RetrievedEvidence could not support the highlight in [Feat] Citation click → open source at page/block with highlight #72.
  • Retrieval was inline and provenance-free — and a strategy could not be added or measured without editing the service.
  • The new reindex would downgrade legacy documents. The migration only adds documents.structure; re-indexing a pre-upgrade PDF would rebuild flat paragraphs and drop the page/bbox provenance it used to have.

Related issue

Fixes #109
Fixes #110
Fixes #74

Part of the v1.4 — Trusted Research Loop epic: #82

What changed?

  • Source identity (fix). Split indexing into indexDocument(documentId, content, structure, options) (derived chain only, never touches documents) and clearDerivedIndex(documentId). addDocument/addDocumentFromFile create the source then index; reindexDocument clears then indexes the same id. deleteDocument reuses the same teardown.
  • Persisted parse structure. New documents.structure JSON column (migration 0016), so a re-index rebuilds the same blocks without re-parsing a file that may have changed on disk.
  • Lazy legacy recovery. reindexDocument() recovers structure from localFilePath before clearing the old index, and only when the re-parsed content is byte-identical to documents.content.
  • Paragraph-level PDF blocks (feat). New pure pdfTextLayout module groups TextItems into lines and paragraphs, each with its own normalized bbox. buildPageBlocks() consumes page.blocks, falling back to the old shape.
  • bbox through the whole chain (fix). ChunkProvenanceBlock now carries bbox; the provenance join selects document_blocks.bbox and it reaches RetrievedEvidence.locator.blocks[].
  • Retriever seam (refactor). New src/main/services/retrieval/: Retriever, RetrievedEvidence, DenseRetriever, hydrateEvidence(). KnowledgeService.search() keeps its signature and delegates; the embedding-space guard stays in the service. includeContent was removed from RetrieveOptions (it is a legacy mapping concern).
  • Batch provenance. resolveChunksProvenance(db, chunkIds) resolves any number of chunks with one join; no N+1 at topK.
  • Docs. docs/architecture.md records the provenance chain and the Retrieval seam.

How was this tested?

  • npm test — 125 pass (added test/pdfTextLayout.test.ts and test/retrievalEvidence.test.ts, the latter now asserting the paragraph bbox survives).
  • npm run typecheck — clean.
  • npm run lint — 0 errors (pre-existing warnings, unchanged).
  • npm run check:design — no violations.
  • npm run build — clean.
  • npm run build:unpack && npm run smoke:packaged — 18/18 checks pass, including three new packaged-app checks:
    • re-index rebuilds the derived index in place and keeps the source identity,
    • DenseRetriever returns retrieved evidence with a page/block locator (asserts page, bbox),
    • re-index recovers structure for documents imported before the column existed.

Follow-ups (deliberately not in this PR)

Checklist

  • I have reviewed my own changes.
  • npm run typecheck passes.
  • npm run build passes.
  • I have tested the affected user workflow.
  • I have not included unrelated changes.
  • I have updated documentation when necessary.

Desktop / build changes

  • Not applicable
  • npm run build:unpack passes.
  • npm run smoke:packaged passes.

reindexDocument() used to delete the old derived index and then call
addDocument(), which minted a new documentId, dropped localFilePath and
left the old row stuck in 'processing'. Once citations persist a
documentId, that silently repoints every historical answer at a different
source.

Split the indexing path so a document is a stable source identity:

  addDocument / addDocumentFromFile -> create source -> indexDocument(id)
  reindexDocument                   -> clearDerivedIndex(id) -> indexDocument(id)

indexDocument() owns only the derived chain (blocks -> chunks ->
embeddings -> vectors -> status) and never inserts or deletes documents.
clearDerivedIndex() is the symmetric teardown, shared with deleteDocument.

Persist the parse structure on the source row (documents.structure) so a
re-index rebuilds the same page/paragraph blocks without re-parsing a file
that may have changed on disk.

Fixes #109
PdfLoader joined every TextItem on a page with a space and emitted one
page-level block, so a citation could say "page 12" but never "the
paragraph at the bottom of page 12".

Add a pure pdfTextLayout module that groups items into lines by baseline
and lines into paragraphs by vertical gap and first-line indent, each with
its own normalized bbox. PdfLoader builds the canonical content from those
paragraphs, so content.slice(startOffset, endOffset) === block.text keeps
holding; buildPageBlocks() consumes page.blocks when present and falls back
to the old one-block-per-page shape otherwise.

Geometry is unit-tested with synthetic items; the existing single-run
fixtures still produce one block per page.

Fixes #110
Retrieval was inline in KnowledgeService.search(): verify the embedding
space, embed the query, query the vector store, then join chunks and
documents. A new strategy could not be added without editing the service,
and the eval harness (#75) had no stable entry point.

Introduce src/main/services/retrieval/:

  interface Retriever {
    search(notebookId, query, opts): Promise<RetrievedEvidence[]>
  }

RetrievedEvidence carries provenance as a first-class field (source
title/type, page range, ordered block spans), so the citation layer (#69)
never queries the database again. DenseRetriever is the current strategy;
the embedding-space guard stays in KnowledgeService because it is a
precondition, not a strategy concern.

Provenance hydration is now batch: resolveChunksProvenance() resolves any
number of chunks with one join instead of one query per chunk, avoiding the
N+1 that topK=5 would otherwise create. hydrateEvidence() wires the batch
queries into RetrievedEvidence[].

KnowledgeService.search() keeps its signature and delegates; SearchResult
gains a locator field and is otherwise byte-for-byte the same.

Fixes #74
@github-actions github-actions Bot added the enhancement New feature or request label Sep 25, 2026
PdfLoader produces per-paragraph bboxes and document_blocks stores them,
but ChunkProvenanceBlock never read document_blocks.bbox, so the value was
dropped before RetrievedEvidence. #69 could build a citation on the current
shape and #72 would then only know the page and blockId, not where on the
page the paragraph sits.

Select document_blocks.bbox in the provenance join, project it into
ChunkProvenanceBlock and assert it survives into
RetrievedEvidence.locator.blocks in the unit test, the evidence test and
the packaged smoke check.

Also drop includeContent from RetrieveOptions: which fields the legacy
SearchResult exposes is a mapping concern, not a retrieval strategy one,
and DenseRetriever never read it.
The migration only adds documents.structure; documents imported before it
have NULL there. The new reindexDocument() clears the derived index first,
so re-indexing a pre-upgrade PDF would rebuild flat paragraphs and silently
drop the page/bbox provenance it used to have.

Recover the structure from localFilePath before clearing anything: re-parse
the local copy, keep the result only when the content is byte-identical to
the canonical documents.content, and persist it. If recovery fails the old
index is still cleared, but we never destroy existing provenance to attempt
recovery.

Verified in the packaged smoke test by nulling structure on an imported PDF,
re-indexing, and asserting the structure is backfilled and the blocks keep
their page number.
@mrsibe
mrsibe merged commit 0545f6c into main Sep 25, 2026
4 checks passed
@mrsibe
mrsibe deleted the feat/v1.4-pre-citation-provenance branch September 25, 2026 17:10
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request

Projects

None yet

1 participant