Repository navigation
Read cost and OOM on fragment/history-heavy branches #566
Description
Activity
- addedacceptedTriaged and validated; open for a PRTriaged and validated; open for a PRP-highHigh priorityHigh priority
on Aug 28, 2026 Controlled fragment measurements for the read-performance portion of this issue:
Two synthetic graphs contained the same 512 Event IDs and values (sum 130,816). Both had 514 live internal-manifest fragments; one had 512 Event data fragments and the other one Event fragment. Fresh 0.10.0 CLI processes read them through local MinIO with 17 ms additional delay per physical HTTP request, preserving request concurrency.
Layout Physical requests Response bytes Read latency 512 Event fragments, 514 manifest fragments 2,084 4.04 MB 1.821 s 1 Event fragment, 514 manifest fragments 1,573 3.65 MB 1.503 s After optimize: 1 Event and 1 manifest fragment 46 0.98 MB 0.691 / 0.701 s across the two graphs The paired pre-maintenance control isolates a 511-request difference attributable to data-fragment layout. Much of the remaining cold work is internal metadata. Optimize changes both layouts and builds indexes, so its full effect is not a pure index experiment.
Warm-server sum queries on the same small fixtures used only 6–7 requests, with roughly 0.15–0.16 s medians before and after optimize. Repeated warm reads can hide the cold amplification. A benchmark should therefore report initial open/read, steady state, eviction/restart, and intervening-write behavior separately, including request bytes and RSS.
These results establish fragment-related read amplification. The tiny synthetic rows do not reproduce the large-payload OOM or establish a cloud-provider latency percentile. The proxy buffered bodies and did not reproduce provider jitter/service limits.
Merge-specific metadata/base costs and independent-target queueing are now separately tracked in #641 and #643; their mechanisms should not be inferred from warm point-read timings.
- changed the title
[-]Production papercuts: large-branch read OOM (fragment-count scaling), no stored-query hot-reload, order-by-aggregate unsupported[/-][+]Read cost and OOM on fragment/history-heavy branches[/+]on Oct 7, 2026 Narrowing this issue to the remaining fragment/history-heavy read problem. Live stored-query activation is implemented by #861/#878, and aggregate ordering shipped in 0.12. The earlier controlled measurements establish cold-read amplification but do not reproduce the reported large-payload OOM or prove its resolution on the current engine; that qualification remains open.
Remaining problem
The original 0.8.1 report described a plain read OOM-killing a single-node server on S3-compatible storage after a staged branch accumulated roughly 3,000 commits. Reading the branch for review became progressively more expensive. This issue now tracks that read-path problem only; its original query-deployment and aggregate-ordering requests are resolved below.
Controlled 0.10.0 measurements established fragment-related cold-read amplification: equal 512-row graphs used 2,084 requests with 512 data fragments versus 1,573 with one data fragment; both retained 514 manifest fragments. After optimize, each used 46 requests. Warm reads hid much of the cold cost. Those small fixtures did not reproduce the original large-payload OOM, and optimize changed several physical properties, so its result does not isolate one cause.
Later history, query-engine and server work may change these costs. The original OOM has not been qualified as fixed on the current release; neither the old measurements nor a warm-read benchmark establish a current regression or a memory bound.
Completion criteria
General process memory admission remains #724; per-query timeout/cancellation remains #565. Merge-specific costs remain #641 and #643.
Resolved parts of the original report
cluster apply --server.order { count($d) desc }example.