Repository navigation
feat(provenance): persist chunk↔block mapping and page range - #108
Merged
Merged
Conversation
A retrieval result only carries a chunk id. Without a persisted mapping it cannot be resolved back to a page, a heading or a character range without re-parsing the document — which is what #68 (and the second half of #67) requires. - `chunk_blocks` (chunk_id, block_id, start_in_block, end_in_block, PK on chunk+block) plus denormalized `chunks.page_start`/`page_end`. Additive migration 0015; existing documents keep working at document-level. - `chunkProvenance.ts`: `insertChunkBlocks()` writes the mapping, and `resolveChunkProvenance()` resolves a chunk id with a single join through `chunk_blocks` into `document_blocks`, returning the ordered block spans and the page range. `projectChunkProvenance()` (pure) does the ordering and projection. - `KnowledgeService.saveChunks()` writes each chunk and its mapping in one transaction, so a partially indexed document never has chunks without their mapping. `getChunkProvenance()` exposes the resolver. `deleteDocument` and `reindexDocument` delete the mapping explicitly rather than relying on a foreign-key cascade the connection does not enable. - The packaged smoke test now round-trips a fixture through the real database: document + blocks + chunk + mapping, then resolves the chunk and checks the block spans and page range. Verification: `test/chunkBlocks.test.ts` gains two tests — the ordered-span projection and the page-range invariant (page_start/page_end equal the min/max page of the mapped blocks) — 110 tests pass. `npm run typecheck`, `npm run build` and `npx eslint` are clean. The packaged smoke check runs in CI (no xvfb locally). Fixes #68 Refs #67
This was referenced Sep 25, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What does this PR do?
Persists the chunk → block mapping and the denormalized page range, so a retrieval result can be resolved back to its ordered blocks and its page without re-parsing the document.
This is the #68 half of #67, now rebased directly onto
mainafter #106 (the chunking half) merged. It supersedes #107, which GitHub closed when its stacked base branch was deleted.Why?
A retrieval result only carries a chunk id. Even with correct canonical offsets, provenance cannot survive in the database without the mapping — "chunking without a persisted block mapping is not a partial feature, it is an unverifiable one" (#68, epic #82).
Related issue
Fixes #68. Completes the persistence half of #67.
What changed?
0015_colossal_solo.sqlchunk_blocks(chunk_id, block_id, start_in_block, end_in_block)with a composite primary key and an index onblock_id.chunks.page_start/chunks.page_end(bothNULLfor non-paged sources), so a citation does not need aMIN/MAXjoin on the hot path.src/main/services/chunkProvenance.tsinsertChunkBlocks()writes the mapping.resolveChunkProvenance()resolves a chunk id with a single join (chunks → chunk_blocks → document_blocks) and returns the ordered block spans plus the page range.projectChunkProvenance()is the pure ordering/projection step and is unit-tested.KnowledgeServicesaveChunks()writes each chunk and its mapping in one transaction, so a partially indexed document never has a chunk without its mapping, and storespage_start/page_end.getChunkProvenance()exposes the resolver.deleteDocument()andreindexDocument()delete mappings explicitly, because the connection does not enable foreign-key cascades.How was this tested?
npm test— 110 tests pass.test/chunkBlocks.test.tsgains the ordered-span projection and the page-range invariant (page_start/page_endequal the min/max page of the mapped blocks).better-sqlite3(the Node unit-test runner has an Electron-ABI build).npm run typecheck— clean.npm run build— clean.npx eslint— clean.Screenshots / recordings
Not applicable — no UI change.
Checklist
npm run typecheckpasses.npm run buildpasses.Desktop / build changes
npm run build:unpackpasses. (delegated to CI)npm run smoke:packagedpasses. (delegated to CI; includes the provenance round trip)