Skip to content

Store a decimal column as unscaled units on the scalar lane - #765

Merged
tamnd merged 1 commit into
mainfrom
decimal-column
Aug 25, 2026
Merged

tamnd merged 1 commit into
mainfrom
decimal-column

Conversation

@tamnd

@tamnd tamnd commented Aug 25, 2026

Copy link
Copy Markdown
Owner

Second half of the exact decimal work. #761 gave the engine a decimal value; this gives it a decimal column, which is the box S2 asks for.

The design

The lane stores the unscaled integer and the catalog stores the scale. A column of DECIMAL(12,2) is a column of pence, and the declared type is what turns a word of a hundred and twenty of them back into 1.20. It is the declared versus encoding split of schema/06 section 2 at its plainest, and it is the first column type whose declaration is needed to read a row back rather than only to check one. Every other lane column reads back the same value whether or not you know what it was declared as; this one does not, which is why the props version goes to 11.

Where the frontier falls, and why it is not a number picked here

LogicalType::physical already maps a decimal by IntBits::for_digits, so lane_width already answered for a decimal before this change: four bytes to nine digits, eight to eighteen, and nothing above. That gives the rule for free. Eighteen digits of unscaled units ride a 64 bit word and nineteen do not, so DECIMAL(18,2) is storable and DECIMAL(38,2) stays declarable in the grammar and refused by the catalog. That is the same shape INT128 and FLOAT16 have on the S2 list, and it is why every test that used DECIMAL(12,2) as its example of an unstorable type now uses DECIMAL(38,2).

The bound lives in extended_bytes, which both the declared form and the column form go through, rather than one caller up in column_type_bytes. That keeps the two forms one encoding for this type: a decimal a graph type may name is a decimal a column can hold, so the refusal lands at the declaration where the user wrote it rather than at the first insert. The alternative, letting the catalog record any precision and refusing at table time the way a nested list is refused, is a larger change and belongs with whatever gives a wide decimal a column.

Exactness is enforced and not repaired

A value the column cannot write exactly at its own scale is refused with 22003. CAST('1.234' AS DECIMAL(6,3)) into a DECIMAL(12,2) column raises rather than rounding to 1.23. Rounding a price on the way into a ledger is the mistake this whole type exists to stop, and a caller who wants it rounded has ROUND to say so with. A value wider than the declared precision is refused the same way. Both checks are on the statement path in insert.rs::cell, which SET also goes through, and the precision one is repeated at the column writer in check_declared, which is where a bulk caller meets it.

Going in, an integer is taken as well as a decimal, since 7 is a value of DECIMAL(12,2) and asking a user to write CAST(7 AS DECIMAL(12,2)) would be pedantry. Coming out, the value carries the column's scale, so a 0.5 written into a two place column reads back as 0.50.

Two hazards closed on purpose

bounded_list now refuses a decimal element. It would otherwise have said yes, because it asks lane_width and a decimal answers, and the element reader has no decimal arm, so a row would have gone in as a list of prices and come back as a list of integers of pence. Refusing at the writer is what keeps this from being a file the reader misreads rather than one it refuses.

check_col in snapshot.rs now refuses a decimal column alongside a zoned one. The vector read path has no decimal arm either, and a caller reaching it another way would have been handed unscaled units dressed as integers. Both of these are wrong answers rather than missing ones, which is the line this codebase draws between a refusal and a gap.

What is pinned

graph_type.rs moves DECIMAL(12,2) into the storable block and adds DECIMAL(38,2) beside INT128, so the frontier table now has one type on both sides of it and says which argument decides. props.rs gains the same split at the encoder, asserting that column_type_bytes and declared_type_bytes agree for every precision. declare.rs gains the round trip: three values in at three different scales, all three back at the column's scale, and both refusals.

The conformance corpus gains the positive case beside the negative one, and the negative one now asks about a decimal too wide for a word rather than about decimals in general.

Left open

Decimals as list elements. The bulk append path, where ColBuf::for_type returns nothing for a decimal, which is the same state Str(n,n), Bytes(n), bounded lists and zoned columns are in. The vector path. And precisions above eighteen digits, which want a plane of their own the way a zoned column has one, or a 128 bit lane.

Gate

Local run is green: fmt, clippy with -D warnings, xtask terms, xtask api-map, and 89 test result: ok lines across the six test targets.

Closes nothing on its own; ticks the Decimal(p,s) box on #700.

A decimal value has been exact since the last change; a decimal column
has not existed. This adds it, for the precisions a lane word holds.

The lane stores the unscaled integer and the declared type stores the
scale, so a column of DECIMAL(12,2) is a column of pence and the type
is what makes a word of them into one pound twenty. That is the
declared versus encoding split of schema/06 section 2 at its plainest,
and it is the first column type whose declaration is needed to read a
row back rather than only to check one.

Where the frontier falls follows from the lane rather than from a
number picked here. LogicalType::physical already maps a decimal by
IntBits::for_digits, so eighteen digits of unscaled units ride a 64
bit word and nineteen do not. DECIMAL(18,2) is storable and
DECIMAL(38,2) stays declarable in the grammar and refused by the
catalog, which is the same shape INT128 and FLOAT16 have on the S2
list.

A value the column cannot hold exactly at its own scale is refused
with 22003 rather than rounded. Rounding a price on the way into a
ledger is the mistake this type exists to stop, and a caller who wants
it rounded has ROUND to ask with. A value wider than the declared
precision is refused the same way, on the statement path and again at
the column writer, which is where a bulk caller meets it.

Two hazards are closed on purpose. A bounded list of decimals would
have passed the lane width test and produced a file whose elements
read back as integers of units, so bounded_list refuses a decimal
element. The vector read path has no decimal arm, so check_col refuses
a decimal column there rather than handing back unscaled units dressed
as integers.

The declared form and the column form stay one encoding: the lane
bound lives in extended_bytes, which both go through, so a decimal a
graph type may name is a decimal a column can hold and there is no gap
between them for a declaration to fall into.

Left open: decimals as list elements, the bulk append path, the vector
path, and precisions above eighteen digits, which want a plane of
their own the way a zoned column has one.
@tamnd
tamnd merged commit 09e6a01 into main Aug 25, 2026
39 of 40 checks passed
@tamnd
tamnd deleted the decimal-column branch August 25, 2026 06:46
tamnd added a commit that referenced this pull request Aug 25, 2026
Two decimal entities from #765, which merged without them.
tamnd added a commit that referenced this pull request Aug 25, 2026
Two decimal entities from #765, which merged without them.
tamnd added a commit that referenced this pull request Aug 25, 2026
* exec: order and build agg rows in the group table

The keyed aggregation sink drained its folded table into a vector of
(keys, states) pairs, sorted that by comparing the decoded Values, and
then walked it again to interleave keys and aggregates into clause
order. For a hundred thousand groups that is two small vectors per
group before anything is known about where they go, a sort that chases
a pointer per compare, and a third pass that throws the pairs away.

The table has the packed key words, so it can do all three at once:
sort an index vector with a compare that reads the words, then decode
each group once, straight into its finished row. The word order is the
value order for every part kind a key can hold, which is what the new
test holds it to against the old drain and sort.

drain stays for the tests, which read groups in insertion order.

* bench: write down what the group row build cost

* docs: regenerate the api model

Two decimal entities from #765, which merged without them.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant