DocForge produces deterministic chunks: identical input + config always yields identical chunk IDs and content boundaries.
Splits on Markdown heading lines (# through ######).
Parameters:
| Param | Default | Description |
|---|---|---|
min_heading_level |
1 | Minimum heading level to split on |
max_heading_level |
3 | Maximum heading level to split on |
max_tokens |
512 | Sub-split oversized sections by size |
include_heading_path |
true | Include heading_path array in each chunk |
When to use: Structured docs (manuals, wikis, reports with clear sections).
Example heading_path:
["Introduction", "Getting Started", "Installation"]Splits by token count using tiktoken (cl100k_base by default).
Parameters:
| Param | Default | Description |
|---|---|---|
max_tokens |
512 | Maximum tokens per chunk (64–8192) |
overlap_tokens |
64 | Overlap between consecutive chunks |
token_model |
cl100k_base | tiktoken encoding name |
When to use: Unstructured text, transcripts, logs, or fixed-size embedding windows.
{
"id": "a1b2c3d4e5f67890",
"index": 0,
"content": "chunk text…",
"token_estimate": 142,
"char_count": 580,
"heading_path": ["Section", "Subsection"],
"metadata": {
"strategy": "headings",
"section_index": 0
}
}id = first 16 hex chars of SHA256(document_id:index:content).
Use id for idempotent vector upserts.
Computed via tiktoken encoding. Default model encoding: cl100k_base (OpenAI GPT-4 / text-embedding-ada-002 family).
Estimates are consistent within DocForge but may differ slightly from provider billing tokenizers.
| Source type | Strategy | max_tokens |
|---|---|---|
| PDF reports | headings | 512 |
| HTML pages | headings | 384 |
| DOCX policies | headings | 512 |
| Plain logs | size | 256–512 |
| Code-adjacent docs | headings | 768 |
- Same
document_id, markdown, andChunkConfig→ same chunks - Re-ingesting the same file produces a new
document_idunless you pass markdown to/chunkwith a fixeddocument_id - For reproducible IDs across runs, call
/chunkwith an explicitdocument_id
Input:
# Product Guide
## Overview
Short intro.
## Details
Long paragraph… (5000 tokens)headings: 2+ chunks (Overview separate; Details may sub-split by size)
size: N chunks of ~512 tokens with 64-token overlap regardless of headings