Skip to content

Order and build the aggregation rows in the group table - #766

Merged
tamnd merged 3 commits into
mainfrom
px-sinks
Aug 25, 2026
Merged

tamnd merged 3 commits into
mainfrom
px-sinks

Conversation

@tamnd

@tamnd tamnd commented Aug 25, 2026

Copy link
Copy Markdown
Owner

The 100k group query on the groupby bench scales about 1.7x from one worker to eight, and the twelve group query beside it scales 9.4x. A phase split of the two halves says why: at eight threads the parallel drive is 35 to 49 ms and the serial finish behind it is 24 to 64, so the tail is as big as the scan it is waiting on. Inside the finish the four parts read merge 8 to 18, drain 2 to 12, sort 7 to 19 and row build 7 to 17 ms.

This takes the drain, the sort and the row build. The sink used to drain the folded table into a Vec<(Vec<Value>, Vec<Acc>)>, which is two small vectors per group before anything is known about where any of them goes, sort that by comparing the values the groups decoded to, and then walk the result a third time to put the keys and the aggregates back in the order the RETURN clause named them. A hundred thousand groups is three hundred thousand allocations made and thrown away in the tail of a query.

The table has the packed key words, so it can do all three at once. GroupTable::rows orders an index vector with a compare that reads those words, then decodes each group once, straight into its finished row. The word order is the value order for every kind a key part can be: an integer and a temporal lane order by the word, a node by its table and then its offset, a string by its bytes. A date is the one lane stored narrower than the word it rides in, so it compares through the same narrowing the decode does, and there is a test for a date before the epoch. The equivalence itself has a test too, on a key holding one part of every kind with both string forms in it, against the old drain and sort.

The hundred thousand group query went 136.7 ms to 111.9 at one worker and 80.6 to 64.9 at eight, medians of six paired alternating runs of two separately built binaries on the local M series, and the branch was faster in every pair. The twelve group shape did not move and neither did the string key shape over a thousand groups, which is the check that this was the tail and not the scan. The floors in budgets.toml are unchanged and the prose beside them records the measurement.

drain stays behind cfg(test), since the tests read groups in insertion order and that is still the thing worth asserting about the probe.

Part of P2, zu#75. The merge is still serial and the finish is still the reason that query does not scale, so the parallel sink checkbox stays open.

tamnd added 2 commits August 25, 2026 13:56
The keyed aggregation sink drained its folded table into a vector of
(keys, states) pairs, sorted that by comparing the decoded Values, and
then walked it again to interleave keys and aggregates into clause
order. For a hundred thousand groups that is two small vectors per
group before anything is known about where they go, a sort that chases
a pointer per compare, and a third pass that throws the pairs away.

The table has the packed key words, so it can do all three at once:
sort an index vector with a compare that reads the words, then decode
each group once, straight into its finished row. The word order is the
value order for every part kind a key can hold, which is what the new
test holds it to against the old drain and sort.

drain stays for the tests, which read groups in insertion order.
Two decimal entities from #765, which merged without them.
@tamnd
tamnd merged commit 6cc92ea into main Aug 25, 2026
40 checks passed
@tamnd
tamnd deleted the px-sinks branch August 25, 2026 07:39
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant