Repository navigation
feat(chunking): block-aware chunks with canonical offsets - #106
Merged
Merged
Conversation
`ChunkingService.chunk()` ran `preprocessText()` first and computed `startOffset`/`endOffset` against the cleaned string, so `chunks` did not address `documents.content` at all — the root cause behind "we show a source title but cannot highlight the source". Rewrite the chunker around the block sequence from #66: - `chunkBlocks(content, blocks)` splits at semantic unit boundaries (line → sentence → whitespace) inside each block, so a boundary can never fall inside a heading: a heading block is one atomic unit. It never re-cleans text, and every chunk's `content` is exactly `content.slice(startOffset, endOffset)`. - Chunks carry `blockSpans` (the ordered blocks covered, with the character range consumed in each) and `pageStart`/`pageEnd`. Paged documents break at page boundaries unless `allowSpanPages` is set. - Overlap is expressed in canonical offsets and aligned to unit starts, so two overlapping chunks share the exact same bytes. - The char-window fallback (`chunk()`) remains for block-less sources and now also produces exact canonical offsets; the old length-changing `preprocessText()` and the unused `chunkBySentence()` are gone. - `KnowledgeService` builds one batch of identified blocks, persists it, and hands the same batch to the chunker, so chunk spans can be persisted next (#68) without rebuilding anything. Persisting `chunk_blocks` and the `chunks.page_start/page_end` columns is the separate #68 step; this change already makes new chunks address `documents.content`. Verification: `test/chunkBlocks.test.ts` (8 tests) asserts the slice invariant, dense indexing, span resolution, heading atomicity, page ranges, cross-page opt-in, overlap byte-identity, and the fallback — 108 tests pass. `npm run typecheck` and `npm run build` are clean. Refs #67
5 of 9 tasks
This was referenced Sep 25, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What does this PR do?
Makes
chunks.start_offset/end_offsetaddressdocuments.contentagain, and makes the chunker structure-aware: it consumes thedocument_blockssequence from #66, never splits a heading, never crosses a PDF page by default, and records which blocks each chunk covers.This is the chunking half of #67. Persisting
chunk_blocksandchunks.page_start/page_endis the stacked #68 PR, which builds on this branch.Why?
ChunkingService.chunk()ranpreprocessText()first and then computed offsets against the cleaned string, whiledocuments.contentholds the loader's text. Sochunksdid not address the stored document at all — the root cause behind "we show a source title but cannot highlight the source" (#67, epic #82).Related issue
Refs #67 (this PR is the chunking half; #68 persists the mapping)
What changed?
ChunkingServicerewritten aroundchunkBlocks(content, blocks):contentis exactlycontent.slice(startOffset, endOffset), so the offset contract holds by construction.blockSpans(ordered blocks covered +startInBlock/endInBlock) andpageStart/pageEndfor each chunk. Paged documents break at page boundaries unlessallowSpanPages: true.chunk()remains as the char-window fallback for block-less sources and now also produces exact canonical offsets.preprocessText()and the unusedchunkBySentence().assignBlockIds(documentId, drafts)indocumentBlocks.tsgenerates block ids once.KnowledgeServicebuilds one identified batch, persists it, and passes the same batch to the chunker — so [Feat] Persist chunk↔block mapping and denormalized page range on chunks #68 can persist spans without rebuilding anything.KnowledgeServicecallschunkBlocks()in both ingestion paths.How was this tested?
npm test— 108 tests pass (8 new intest/chunkBlocks.test.ts): the exact-slice invariant, dense/ordered indexing, block spans resolving back into their blocks in document order, heading atomicity under a tinychunkSize, page ranges, cross-page opt-in, byte-identical overlap, and the char-window fallback.npm run typecheck— clean.npm run build— clean.npx eslinton every changed file — clean.Not run locally:
build:unpack+smoke:packaged(noxvfb). CI runs them on all three platforms.Screenshots / recordings
Not applicable — no UI change.
Checklist
npm run typecheckpasses.npm run buildpasses.Desktop / build changes
npm run build:unpackpasses. (delegated to CI)npm run smoke:packagedpasses. (delegated to CI)