Skip to content

Size a bounded list's per-row count to its declared bound - #770

Merged
tamnd merged 1 commit into
mainfrom
bounded-list-offsets
Aug 25, 2026
Merged

tamnd merged 1 commit into
mainfrom
bounded-list-offsets

Conversation

@tamnd

@tamnd tamnd commented Aug 25, 2026

Copy link
Copy Markdown
Owner

S2 item 1's second half, on #700. FixedList(n) landed earlier; this is BoundedList(n), which the issue describes as sizing the offsets to ceil(log2(n)) bits instead of 32.

A LIST<T>(n) column whose rows are all at the bound already drops the count out of every row and keeps it in the directory. That is what got a 768 dimension embedding row down to 3072 bytes and it is the case S6 cares about. A column whose rows are ragged cannot do that: a row of 700 in a (768) column still has to say it is 700. What it does not have to do is say so in four bytes. The count is a number between nought and the declared bound, check_declared has already said so for every row, and a number that cannot pass 255 does not need four bytes.

So count_width picks the head width from the bound, one byte to 255 and two to 65535, and ListRows replaces the Option<usize> that list_elements and encode_list_row took. Fixed(n) is the directory holding the count and Counted(w) is the row holding it in w bytes. Neither of those is tellable from the bytes, since a bounded LIST<INT32>(768) row of 767 counted elements is exactly as long as one of 768 uncounted ones, so the column entry says and the reader takes the answer. That was already true of the fixed case and this makes it true of the width too.

The one design decision worth reviewing is that the width is carried in the column entry rather than worked out from the declared bound. A fold rewrites the directory at the current PROPS_VERSION while leaving the row blobs where they are, so a column written at version 11 comes back through encode with a version 12 header over rows that still spend four bytes on a count a one byte field would hold. Deriving the width from the bound would read those rows as a tiny count followed by an enormous number of elements, which is to say it would report damage to a file that has none. fold.rs copies count_width off the old entry the way it already copies fixed_len, and there is a test that pins exactly that combination: a narrow bound and a wide count.

The saving is real where the elements are small and the bound is small with them, which is where the count was the larger part of the row. A LIST<INT8>(4) row of four elements was four bytes of payload and four more to say so, and is now four and one. On an embedding column the same change would be two bytes in three thousand, and there it does not arise at all, because a column written whole at its bound carries no count. A list with no declared bound keeps its four bytes, having no number to be smaller than.

Directory version 12. A version 11 directory decodes to a width of four and its rows read as they always did. A width byte that is not one, two or four is refused as corrupt, because that byte decides every element offset below it and a length read at a width nobody wrote it at is a row of nonsense.

Local gate green: fmt, clippy with -D warnings across the workspace and all targets, xtask terms, xtask api-map with no diff, and 89 test result: ok lines across the six suites.

A `LIST<T>(n)` column whose rows are all at the bound already drops the
count out of every row and keeps it in the directory, which is what got
a 768 dimension embedding row down to 3072 bytes. A column whose rows
are ragged cannot do that: a row of 700 in a `(768)` column still has
to say it is 700. But it does not have to say so in four bytes. The
count is a number between nought and the bound, `check_declared` has
just said so for every row, and a number that cannot pass 255 fits in
one byte.

So `count_width` picks the head from the bound, one byte to 255 and two
to 65535, and `ListRows` replaces the `Option<usize>` that `list_elements`
and `encode_list_row` took: `Fixed(n)` is the directory holding the
count and `Counted(w)` is the row holding it in `w` bytes. Neither is
tellable from the bytes, which is why it is a field and not an
inference, and that was already true of the fixed case.

The width is carried in the column entry rather than worked out from the
declared bound, which is the one design decision here. A fold rewrites
the directory at the current version while leaving the row blobs where
they are, so a column written at version 11 comes back through `encode`
with a version 12 header over rows that still spend four bytes on their
counts. Deriving the width from the bound would read those rows as a
tiny count and an enormous number of elements, which is to say it would
report damage to a file that has none.

The saving is real where the elements are small and the bound is small
with them, which is where the count was the larger part of the row: a
`LIST<INT8>(4)` row of four elements was four bytes of payload and four
more to say so. On an embedding column it would be two bytes in three
thousand, and there it does not arise at all, because those rows carry
no count. A list with no declared bound keeps its four bytes, having no
number to be smaller than.
@tamnd
tamnd merged commit 31bb59b into main Aug 25, 2026
39 of 40 checks passed
@tamnd
tamnd deleted the bounded-list-offsets branch August 25, 2026 07:37
tamnd added a commit that referenced this pull request Aug 25, 2026
Bounded list entities from #770, which merged without them.
tamnd added a commit that referenced this pull request Aug 25, 2026
* exec: fold the group tables up a tree

Every worker sees keys from all over the scan, so each partial ends up
about as wide as the answer and folding eight of them into the first is
seven full width merges one after another, at the point in the query
where nothing else is running. On the hundred thousand group bench that
is 8 to 18 ms of a 65 ms query.

Same seven merges, three rounds, each round on the pool the scan has
just finished with. An odd table carries to the next round rather than
joining a pair, so every merge in a round is the same size, and this
thread takes one pair itself the way worker zero does on the scan.

* exec: build the answer rows on the pool

The rows of a wide aggregation were built one after another on the
thread that finished the run, and that is 7 to 17 ms of a 65 ms query
at eight workers. Decoding is per group and the groups are settled once
the order is, so slices of the order are independent work and the pool
the scan just finished with is sitting idle.

One hand per worker that ran, not one per core: a query asked for on
one thread is answered on one thread, tail included, or the per core
numbers in budgets.toml stop meaning what they say. Under four thousand
groups a hand it stays on this thread, since a latch and a lock per
hand cost more than the decoding they would split.

* docs: regenerate the api model

Bounded list entities from #770, which merged without them.

* bench: write down what the parallel finish is worth
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant