Skip to content

Latest commit

 

History

History
103 lines (70 loc) · 2.67 KB

File metadata and controls

103 lines (70 loc) · 2.67 KB

DocForge Chunking Guide

DocForge produces deterministic chunks: identical input + config always yields identical chunk IDs and content boundaries.

Strategies

headings

Splits on Markdown heading lines (# through ######).

Parameters:

Param Default Description
min_heading_level 1 Minimum heading level to split on
max_heading_level 3 Maximum heading level to split on
max_tokens 512 Sub-split oversized sections by size
include_heading_path true Include heading_path array in each chunk

When to use: Structured docs (manuals, wikis, reports with clear sections).

Example heading_path:

["Introduction", "Getting Started", "Installation"]

size

Splits by token count using tiktoken (cl100k_base by default).

Parameters:

Param Default Description
max_tokens 512 Maximum tokens per chunk (64–8192)
overlap_tokens 64 Overlap between consecutive chunks
token_model cl100k_base tiktoken encoding name

When to use: Unstructured text, transcripts, logs, or fixed-size embedding windows.

Chunk record schema

{
  "id": "a1b2c3d4e5f67890",
  "index": 0,
  "content": "chunk text…",
  "token_estimate": 142,
  "char_count": 580,
  "heading_path": ["Section", "Subsection"],
  "metadata": {
    "strategy": "headings",
    "section_index": 0
  }
}

Stable IDs

id = first 16 hex chars of SHA256(document_id:index:content).

Use id for idempotent vector upserts.

Token estimates

Computed via tiktoken encoding. Default model encoding: cl100k_base (OpenAI GPT-4 / text-embedding-ada-002 family).

Estimates are consistent within DocForge but may differ slightly from provider billing tokenizers.

RAG pipeline recommendations

Source type Strategy max_tokens
PDF reports headings 512
HTML pages headings 384
DOCX policies headings 512
Plain logs size 256–512
Code-adjacent docs headings 768

Determinism guarantees

  • Same document_id, markdown, and ChunkConfig → same chunks
  • Re-ingesting the same file produces a new document_id unless you pass markdown to /chunk with a fixed document_id
  • For reproducible IDs across runs, call /chunk with an explicit document_id

Example: headings vs size

Input:

# Product Guide

## Overview
Short intro.

## Details
Long paragraph… (5000 tokens)

headings: 2+ chunks (Overview separate; Details may sub-split by size)

size: N chunks of ~512 tokens with 64-token overlap regardless of headings