Repository navigation
Conversation
|
Implementation details for this first version:
Repository files now use the same helper before embedding. The repository allow-list also accepts
Built-in vault files now use the same extension-based strategy as uploads and repository files, so the three ingestion paths produce consistent Markdown chunks.
Added coverage for:
Validation completed:
This is intentionally a small first version for #247. Could you please confirm whether this Markdown structure-aware chunking approach is acceptable? In particular, should the next iteration add:
I would appreciate feedback on whether this implementation direction is suitable and what changes should be made before expanding the scope. |
Summary
.markdownto the repository knowledge-file allow-list..txtand.rstfiles.Why
Markdown headings provide useful document structure for retrieval. Previously, repository and built-in vault sync always used fixed-window chunking, while uploaded Markdown used the structure-aware splitter. This made the same Markdown content behave differently depending on its ingestion path.
This first version centralizes that decision while keeping the existing behavior for plain text and reStructuredText.
Scope
This PR does not change the search API, embedding model, Qdrant schema, point ID rules, or retrieval ranking logic.
Validation
go test -race ./internal/manager/biz/knowledge -count=1go vet ./...go build ./...git diff --checkThe full
go test -race ./...run was attempted. It was blocked by the current container environment: the knowledge package encountered an embed-filecannot allocate memoryfailure, and two existing root/container-sensitive tests failed inaiops/chatruntimeandserver/edge.Risk and rollback
This is an ingestion-only change with no API, schema, or deployment changes. Reverting commit
68a7875restores the previous chunking behavior and repository extension allow-list.Closes #247
Author confirmation