You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Milestone S2 of the schema program (Spec/2064g/schema/00-overview.md). Specs: schema/05, schema/06 §6, schema/02 §5. Depends on S1. Unblocks S3 and S6.
This is the core of the program. Everything else is arranged around it, and its first item is deliberately the one S6 needs.
A closed graph type, feature GG02, is the only construct in GQL that licenses a storage optimisation, because it is the only promise that a property nobody declared will not turn up. zu claims GG02 and currently spends it on one check, promised at crates/zu/src/declare.rs:301. Neo4j's graph types are open only, so Neo4j cannot spend it at all. That is the whole advantage and S2 is where it gets collected.
Two halves. The stored encodings that a declaration should buy, and the enforcement that makes the declaration true.
The encodings
Ordered by value, which is the order they get built. The rule they all serve is in schema/06 §2: the catalog holds what the user declared, each block holds the narrowest encoding that losslessly represents its own values, and the reader promotes on the way out. Half of this exists already, in the split between TYPE_CODES at crates/zu-zu1/src/props.rs:76 and extended_bytes at props.rs:261. S2 generalises it.
FixedList(n) and BoundedList(n) over a fixed width element. LIST<T>(n) bounds the maximum, not the length, so a list of 700 is legal in a (768) column and a fixed stride is not guaranteed by the declaration alone. Two encodings, and the block directory says which. A batch of embeddings always lands in FixedList: no outer mask because of the outer NOT NULL, no child mask because of the element NOT NULL, no offsets because every list is exactly n, so the column is n times 3072 contiguous bytes that can be mmapped and handed to a SIMD dot product as a flat &[f32]. BoundedList keeps offsets but sizes them to ceil(log2(n)) bits instead of 32.
Fixed width Bytes(n). BINARY(16) NOT NULL is the standard GQL spelling of a UUID and it should be sixteen bytes at a fixed stride with no mask and no offsets. Also hashes, IPv6 addresses, and binary quantised vectors.
Fixed width Str(n,n), which is CHAR(n). Stride n bytes, UTF-8 length checked on write.
ZonedDatetime and ZonedTime column codes. Both are declarable today and neither has a column code, so the declarable/storable asymmetry is not only about exotic types.
Decimal(p,s). i64 mantissa where p is at most 18, i128 above, scale in the catalog rather than per value. This is the money story.
Int128, Int256, UInt128. Fixed stride, large ids and hashes as numbers.
Float16. Two bytes, quantised ML scalars.
Shredded Record. One child column per field, recursively. Structs, geospatial points, and JSON with a stable shape. The largest item here, and the one schema/06 §7 says may need a format bump. If child column addressing does not fit the existing table of columns shape in the directory, this item moves to S6 and S2 ships without it. That is a decision to make early, not late.
Items 1 through 7 ride the existing format. New column type codes are extensible by construction and a reader meeting an unknown code already refuses cleanly, which is a rule props.rs already states and this milestone does not weaken.
The enforcement
Plan time, before a row moves: an undeclared property name with a suggestion, a literal whose type is not assignable to the declared type, a null literal into a NOT NULL property, an edge pattern no element type can type. These are 42002, 22G03, 22004 and 42002 respectively, and none of them should cost a row read.
Write time, vectorised, once per batch and never once per value:
null into NOT NULL: a popcount over the batch validity mask, 22004
string longer than the declared bound: a max over the length array, 22001 for the error case and 01004 with a warning where truncation is the defined behaviour
numeric out of the declared range: a min/max compare over the batch, 22003
list longer than the declared bound: a max over the offset deltas, 22G0B list data right truncation, which is exactly the embedding dimension check
null element in a LIST<T NOT NULL>: a popcount over the child mask, 22G0C
array truncation: 2202F
record shape mismatch: 22G0U, 22G0X, 22G0Y
graph type violation on the element as a whole: G2000
Endpoint typing on insert: free in the single element type case because the plan already knows which type it is, and one vectorised pass in the GG24 case where more than one type could hold the labels.
Every condition above gets one corpus case, and the message shapes get golden files.
The gate
Enforcement cannot be turned off. There is no session flag, because a schema that is sometimes true is worse than no schema, and because every optimisation in schema/06 §3 is unsound the moment it can be. That means the cost has to be small enough that nobody wants it off, so the gate is ingest throughput with enforcement on against enforcement off on perf-path-100k, and the ceiling is 3%. Every check above is batch level for that reason. If one of them turns out not to be, it gets redesigned rather than made optional.
Checklist
FixedList(n) over a fixed width element
BoundedList(n) over a fixed width element, which is the narrow offsets
Shredded Record, or the decision that it moves to S6
The declared versus encoding split in the block directory, with the promote edge (the integer tower is done, the rest is not)
Plan time checks
Write time checks, a row at a time on the statement path
Write time checks, vectorised, one pass per batch per column
Endpoint typing on insert
The condition mapping of schema/05 §3, one corpus case per row
Message shapes of schema/05 §4, golden file tested
The 3% ingest gate on perf-path-100k, published to docs/gql-performance.md
Bytes per row published for the 768 dim FLOAT32 column and the BINARY(16) column against Neo4j and Ladybug
The S0 round trip test un-ignored for every type above
Exit
A closed graph type is enforced on every write. The declarable/storable invariant holds for everything except the four types S6 owns. Ingest costs under 3%.
Milestone S2 of the schema program (Spec/2064g/schema/00-overview.md). Specs: schema/05, schema/06 §6, schema/02 §5. Depends on S1. Unblocks S3 and S6.
This is the core of the program. Everything else is arranged around it, and its first item is deliberately the one S6 needs.
A closed graph type, feature GG02, is the only construct in GQL that licenses a storage optimisation, because it is the only promise that a property nobody declared will not turn up. zu claims GG02 and currently spends it on one check,
promisedat crates/zu/src/declare.rs:301. Neo4j's graph types are open only, so Neo4j cannot spend it at all. That is the whole advantage and S2 is where it gets collected.Two halves. The stored encodings that a declaration should buy, and the enforcement that makes the declaration true.
The encodings
Ordered by value, which is the order they get built. The rule they all serve is in schema/06 §2: the catalog holds what the user declared, each block holds the narrowest encoding that losslessly represents its own values, and the reader promotes on the way out. Half of this exists already, in the split between
TYPE_CODESat crates/zu-zu1/src/props.rs:76 andextended_bytesat props.rs:261. S2 generalises it.FixedList(n)andBoundedList(n)over a fixed width element.LIST<T>(n)bounds the maximum, not the length, so a list of 700 is legal in a(768)column and a fixed stride is not guaranteed by the declaration alone. Two encodings, and the block directory says which. A batch of embeddings always lands in FixedList: no outer mask because of the outer NOT NULL, no child mask because of the element NOT NULL, no offsets because every list is exactly n, so the column is n times 3072 contiguous bytes that can be mmapped and handed to a SIMD dot product as a flat &[f32]. BoundedList keeps offsets but sizes them to ceil(log2(n)) bits instead of 32.Bytes(n).BINARY(16) NOT NULLis the standard GQL spelling of a UUID and it should be sixteen bytes at a fixed stride with no mask and no offsets. Also hashes, IPv6 addresses, and binary quantised vectors.Str(n,n), which isCHAR(n). Stride n bytes, UTF-8 length checked on write.ZonedDatetimeandZonedTimecolumn codes. Both are declarable today and neither has a column code, so the declarable/storable asymmetry is not only about exotic types.Decimal(p,s). i64 mantissa where p is at most 18, i128 above, scale in the catalog rather than per value. This is the money story.Int128,Int256,UInt128. Fixed stride, large ids and hashes as numbers.Float16. Two bytes, quantised ML scalars.Record. One child column per field, recursively. Structs, geospatial points, and JSON with a stable shape. The largest item here, and the one schema/06 §7 says may need a format bump. If child column addressing does not fit the existing table of columns shape in the directory, this item moves to S6 and S2 ships without it. That is a decision to make early, not late.Items 1 through 7 ride the existing format. New column type codes are extensible by construction and a reader meeting an unknown code already refuses cleanly, which is a rule props.rs already states and this milestone does not weaken.
The enforcement
Plan time, before a row moves: an undeclared property name with a suggestion, a literal whose type is not assignable to the declared type, a null literal into a NOT NULL property, an edge pattern no element type can type. These are 42002, 22G03, 22004 and 42002 respectively, and none of them should cost a row read.
Write time, vectorised, once per batch and never once per value:
LIST<T NOT NULL>: a popcount over the child mask, 22G0CEndpoint typing on insert: free in the single element type case because the plan already knows which type it is, and one vectorised pass in the GG24 case where more than one type could hold the labels.
Every condition above gets one corpus case, and the message shapes get golden files.
The gate
Enforcement cannot be turned off. There is no session flag, because a schema that is sometimes true is worse than no schema, and because every optimisation in schema/06 §3 is unsound the moment it can be. That means the cost has to be small enough that nobody wants it off, so the gate is ingest throughput with enforcement on against enforcement off on perf-path-100k, and the ceiling is 3%. Every check above is batch level for that reason. If one of them turns out not to be, it gets redesigned rather than made optional.
Checklist
FixedList(n)over a fixed width elementBoundedList(n)over a fixed width element, which is the narrow offsetsBytes(n)Str(n,n)ZonedDatetimeandZonedTimecolumn codesDecimal(p,s), both planes: a lane word of unscaled units to p at most 18 (Store a decimal column as unscaled units on the scalar lane #765), sixteen bytes of them above that (Store the wide decimal on the sixteen byte plane #785), scale in the catalog rather than per valueInt128,Int256,UInt128.Int128landed in Store an INT128 column as sixteen bytes at a fixed stride #777; the other two are refused at the declaration because nothing here carries a value of themFloat16Record, or the decision that it moves to S6Exit
A closed graph type is enforced on every write. The declarable/storable invariant holds for everything except the four types S6 owns. Ingest costs under 3%.