Skip to content

fix(knowledge): unify markdown chunking across ingestion paths - #485

Open
014-code wants to merge 1 commit into
ongridio:mainfrom
014-code:fix/knowledge-markdown-chunking
Open

014-code wants to merge 1 commit into
ongridio:mainfrom
014-code:fix/knowledge-markdown-chunking

Conversation

@014-code

@014-code 014-code commented Oct 9, 2026

Copy link
Copy Markdown
Contributor

Summary

  • Use the structure-aware Markdown splitter consistently for uploads, repository sync, and built-in vault sync.
  • Add .markdown to the repository knowledge-file allow-list.
  • Keep the existing fixed-window splitter for .txt and .rst files.
  • Add regression coverage for extension-based splitter selection and repository scanning.

Why

Markdown headings provide useful document structure for retrieval. Previously, repository and built-in vault sync always used fixed-window chunking, while uploaded Markdown used the structure-aware splitter. This made the same Markdown content behave differently depending on its ingestion path.

This first version centralizes that decision while keeping the existing behavior for plain text and reStructuredText.

Scope

This PR does not change the search API, embedding model, Qdrant schema, point ID rules, or retrieval ranking logic.

Validation

  • go test -race ./internal/manager/biz/knowledge -count=1
  • go vet ./...
  • go build ./...
  • git diff --check

The full go test -race ./... run was attempted. It was blocked by the current container environment: the knowledge package encountered an embed-file cannot allocate memory failure, and two existing root/container-sensitive tests failed in aiops/chatruntime and server/edge.

Risk and rollback

This is an ingestion-only change with no API, schema, or deployment changes. Reverting commit 68a7875 restores the previous chunking behavior and repository extension allow-list.

Closes #247

Author confirmation

@014-code
014-code requested a review from singchia as a code owner October 9, 2026 06:26
@014-code

014-code commented Oct 9, 2026

Copy link
Copy Markdown
Contributor Author

Implementation details for this first version:

  1. Upload ingestion

ingestUpload now delegates chunking to splitKnowledgeContent. Markdown, .markdown, DOCX, and PDF inputs use the existing structure-aware Markdown splitter. DOCX/PDF content has already been converted to Markdown by the upload extraction path.

  1. Repository sync

Repository files now use the same helper before embedding. The repository allow-list also accepts .markdown. The existing exclusions for YAML, TOML, JSON, and other configuration/data formats remain unchanged.

  1. Built-in vault sync

Built-in vault files now use the same extension-based strategy as uploads and repository files, so the three ingestion paths produce consistent Markdown chunks.

  1. Compatibility

.txt and .rst continue to use the existing fixed-window splitter. No search API, embedding model, Qdrant schema, point ID rule, or ranking logic was changed.

  1. Tests

Added coverage for:

  • Markdown heading-aware splitting
  • Case-insensitive .MARKDOWN handling
  • Fixed-window fallback for .txt and .rst
  • Repository discovery of .md, .markdown, .txt, and .rst
  • Continued exclusion of .yaml

Validation completed:

  • go test -race ./internal/manager/biz/knowledge -count=1
  • go vet ./...
  • go build ./...
  • git diff --check

This is intentionally a small first version for #247. Could you please confirm whether this Markdown structure-aware chunking approach is acceptable?

In particular, should the next iteration add:

  • real retrieval-quality evaluation data or a small RAG benchmark;
  • support for additional Markdown-like extensions;
  • different chunk-size/overlap or embedding-cost limits;
  • separate policies for Markdown, plain text, and other document types?

I would appreciate feedback on whether this implementation direction is suitable and what changes should be made before expanding the scope.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

改进 RAG 检索质量 (Improve RAG retrieval quality)

1 participant