You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Transport / compression: evaluate CBOR (record structures / Packed) for large tabular instances #127
An OO-LD instance can be large - a multi-column table or time series is an array of row records. This issue evaluates CBOR and its record/packing extensions as a binary transport that keeps the JSON / JSON-LD data model (see the guide's "Large and bulk data") and answers: (a) size and total encode/decode time vs other formats, (b) whether a record format for tabular data competes with generic compression, (c) whether partial read/write and streaming are possible (cf HDF5).
Example data model
OO-LD schema (structure + @context in one document) for the tabular case:
The benchmark generates N such rows (deterministic pseudo-random floats at 3 decimals, plus the string and bool).
Size and total time (end-to-end)
Each row is a (format x compressor) combination. Ship B is the transported size; compressor = none is the uncompressed size. Total enc = build payload + compress; Total dec = decompress + parse. cbor-x 1.6.5, Node 25.6, gzip level 9, brotli with explicit quality + size hint; single-run timings. Script: specs/tooling-tests/bench-matrix.mjs.
N = 1,000 rows
Format
Compressor
Ship B
Total enc ms
Total dec ms
JSON
none
79,896
0.37
0.57
JSON
gzip-9
17,156
9.35
0.94
JSON
brotli-q5
17,307
2.91
0.96
JSON
brotli-q11
12,956
127.67
1.00
CBOR maps
none
67,987
0.72
1.38
CBOR maps
gzip-9
17,730
3.44
0.93
CBOR maps
brotli-q5
16,267
3.52
0.99
CBOR maps
brotli-q11
14,393
100.31
0.86
CBOR records
none
40,030
0.49
0.68
CBOR records
gzip-9
16,284
2.02
0.33
CBOR records
brotli-q5
14,869
2.55
0.34
CBOR records
brotli-q11
13,473
57.70
0.38
CBOR columnar
none
33,044
0.35
0.29
CBOR columnar
brotli-q5
11,797
1.84
0.37
N = 100,000 rows
Format
Compressor
Ship B
Total enc ms
Total dec ms
JSON
none
8,174,316
43.23
75.93
JSON
gzip-9
1,666,390
1092.78
93.33
JSON
brotli-q5
330,556
159.85
94.19
JSON
brotli-q11
262,522
22603.04
121.05
CBOR maps
none
6,987,964
63.95
76.08
CBOR maps
gzip-9
1,670,234
331.30
90.13
CBOR maps
brotli-q5
233,871
157.24
98.99
CBOR maps
brotli-q11
207,751
17017.28
80.63
CBOR records
none
4,188,007
35.64
13.90
CBOR records
gzip-9
1,558,171
194.72
27.79
CBOR records
brotli-q5
191,613
111.52
23.20
CBOR records
brotli-q11
198,310
8059.13
22.62
CBOR columnar
none
3,489,054
10.19
5.59
CBOR columnar
brotli-q5
94,198
47.48
18.14
CBOR columnar
brotli-q11
63,899
1377.53
14.72
Encoding of each variant
CBOR maps - new Encoder({ useRecords: false }): one map per row, keys repeated.
CBOR records - new Encoder({ useRecords: true }): cbor-x's record-structures extension; the shared shape is written once, rows carry only values. Functional equivalent of the draft's tabular-record idea.
CBOR columnar + typed arrays - struct-of-arrays (Float64Array per numeric column); does not preserve the row-entity model.
A separate pack: true run (not shown above) verified on the wire that cbor-x's packer emits tag 51 and only deduplicates repeated strings (draft-ietf-cbor-packed-03 vintage), not the draft-19 tag-114 tabular-record transform; no packer implementing tag-114 was available, so that mechanism is not benchmarked.
Observations
Uncompressed: CBOR records is 0.51x JSON (4.19 vs 8.17 MB) and decodes about 4x faster with no decompression step (13.9 vs 75.9 ms at 100k).
Compressor choice dominates size and encode time. gzip level 9 has a 32 KB window and cannot deduplicate across a multi-MB file: JSON+gzip-9 = 1.67 MB vs JSON+brotli-q5 = 331 KB (5x smaller) at ~1/7 the encode time (160 vs 1093 ms). A large-window compressor (brotli, or zstd) is required for large tabular data.
brotli quality 11 is impractical at scale: 22.6 s (JSON) and 8.1 s (records) per 100k rows, for no gain over q5 on records (q5 191,613 B is smaller than q11 198,310 B and ~70x faster). brotli-q5 is the usable high-ratio setting.
Smallest structure-preserving: CBOR records + brotli-q5 = 191,613 B (0.58x of JSON+brotli-q5, 0.82x of CBOR maps+brotli-q5), total enc 112 ms, total dec 23 ms. Its decode stays ~4x faster than JSON/maps after compression (23 vs 94/99 ms).
Answer to (b): the record layout and generic compression are complementary, not competing. Under gzip they look redundant (records+gzip 1.56 MB vs JSON+gzip 1.67 MB), but that is a gzip-window artifact; under brotli the record layout is still ~40% smaller than compressed JSON and decodes ~4x faster. The record-only advantages compression cannot provide are fast decode, small size without a decompression step, and streamability.
Compression is opaque: a gzip/brotli payload must be fully decompressed before any row is read, forfeiting the streaming/early-stop/random-access properties in the next section. The trade is smallest-at-rest (records+brotli) vs streamable (raw records / sequence).
Incremental read: decodeMultiple with a callback returning false stops after K items (verified: fired 10/10 for K=10); the remainder is not decoded.
Random access: with an external byte-offset index, one row decodes from its own 66-70 B slice (rows 0 / 500 / 99,999), no other bytes read.
Append: a 3-row tail encodes to 210 B and concatenates to the existing buffer; all 100,000 items still decode. Append cost is O(appended), no rewrite.
Record stream (useRecords: true, sequential: true; 4,187,896 B, keys once):
In-order decodeMultiple resolves field names for all rows.
An isolated tail item does not resolve field names (returns a positional array under tag 57344). Record items are not self-contained: reading item N requires reading from the start (to accumulate structures) or a shared structure catalog.
encodeAsIterable at N=100,000 emitted a single 4,187,947 B chunk (chunk boundary is buffer-size driven, not reached at ~4 MB); incremental emission for larger inputs uses EncoderStream.
Summary: streaming read, early-stop, and append work directly on an uncompressed CBOR sequence; random access to row N needs an external offset index; a record stream halves the size but items are not self-contained, so random access additionally needs read-from-start or a shared structure catalog; and any generic compression forfeits all of these. HDF5/Parquet provide chunk+offset addressing natively (#109).
CBOR Packed (draft-ietf-cbor-packed-19)
Standards-Track IETF draft (draft-ietf-cbor-packed-19, Bormann/Guetschow): transforms CBOR into a packed form so "a separate decompression step is often unnecessary"; it standardizes the format, not a packer algorithm. Shared-item + argument tables; a "record" construct combines an array of keys with per-row arrays of values into maps (the tabular case). Positioned as a lightweight alternative to DEFLATE. cbor-x implements the string-table style (tag 51, draft-03 vintage), not this record construct.
OO-LD stays transport-agnostic: a Packed/record document unpacks to the same JSON, so the data model is unchanged; this is guidance, not a normative format.
Scope / relation
Transport only. Feeds the "Large and bulk data" guide and #109 (dataset storage/distributions); the encoder would live in the oold-python bundler/serialization (OO-LD/oold-python#30, OO-LD/oold-python#32).
Acceptance criteria
Benchmark the draft-19 tag-114 record packer (find or build a conformant packer).
Re-run with realistic domain data (high-precision doubles, real strings) and add zstd.
Confirm the recommendation: record CBOR + brotli-q5 vs external container.
Decide whether OO-LD references a binary transport in guidance or stays agnostic.
Document the streaming/append vs random-access vs smallest-at-rest trade-off in the "Large and bulk data" guide.
Motivation
An OO-LD instance can be large - a multi-column table or time series is an array of row records. This issue evaluates CBOR and its record/packing extensions as a binary transport that keeps the JSON / JSON-LD data model (see the guide's "Large and bulk data") and answers: (a) size and total encode/decode time vs other formats, (b) whether a record format for tabular data competes with generic compression, (c) whether partial read/write and streaming are possible (cf HDF5).
Example data model
OO-LD schema (structure +
@contextin one document) for the tabular case:{ "$schema": "https://json-schema.org/draft/2020-12/schema", "$id": "https://example.org/Measurement.schema.json", "title": "Measurement", "@context": { "ex": "https://example.org/vocab#", "qudt": "http://qudt.org/schema/qudt/", "metadata": "ex:metadata", "unit": "qudt:unit", "data": "ex:hasRow", "name": "ex:label", "value1": "ex:value1", "value2": "ex:value2", "value3": "ex:value3", "ok": "ex:ok" }, "type": "object", "properties": { "metadata": { "type": "object", "properties": { "id": { "type": "string" }, "unit": { "type": "string" } } }, "data": { "type": "array", "items": { "type": "object", "properties": { "name": { "type": "string" }, "value1": { "type": "number" }, "value2": { "type": "number" }, "value3": { "type": "number" }, "ok": { "type": "boolean" } }, "required": ["name"] } } } }Instance excerpt (a valid OO-LD / JSON-LD document; row values illustrative):
{ "@context": "https://example.org/Measurement.schema.json", "$schema": "https://example.org/Measurement.schema.json", "metadata": { "id": "run-1", "unit": "V" }, "data": [ { "name": "test0", "value1": 1.0, "value2": 2.0, "value3": 3.0, "ok": true }, { "name": "test1", "value1": 4.0, "value2": 5.0, "value3": 6.0, "ok": false } ] }The benchmark generates N such rows (deterministic pseudo-random floats at 3 decimals, plus the string and bool).
Size and total time (end-to-end)
Each row is a (format x compressor) combination.
Ship Bis the transported size;compressor = noneis the uncompressed size.Total enc= build payload + compress;Total dec= decompress + parse. cbor-x 1.6.5, Node 25.6, gzip level 9, brotli with explicit quality + size hint; single-run timings. Script:specs/tooling-tests/bench-matrix.mjs.N = 1,000 rows
N = 100,000 rows
Encoding of each variant
new Encoder({ useRecords: false }): one map per row, keys repeated.new Encoder({ useRecords: true }): cbor-x's record-structures extension; the shared shape is written once, rows carry only values. Functional equivalent of the draft's tabular-record idea.Float64Arrayper numeric column); does not preserve the row-entity model.pack: truerun (not shown above) verified on the wire that cbor-x's packer emits tag 51 and only deduplicates repeated strings (draft-ietf-cbor-packed-03 vintage), not the draft-19 tag-114 tabular-record transform; no packer implementing tag-114 was available, so that mechanism is not benchmarked.Observations
Downsides of record structures
Streaming and partial access
Measured on the 100,000-row dataset, cbor-x 1.6.5, on uncompressed payloads. Script:
specs/tooling-tests/bench-streaming.mjs.CBOR sequence (each row an independent CBOR item, RFC 8742; 6,987,863 B, keys repeated):
decodeMultiplewith a callback returningfalsestops after K items (verified: fired 10/10 for K=10); the remainder is not decoded.Record stream (
useRecords: true, sequential: true; 4,187,896 B, keys once):decodeMultipleresolves field names for all rows.encodeAsIterableat N=100,000 emitted a single 4,187,947 B chunk (chunk boundary is buffer-size driven, not reached at ~4 MB); incremental emission for larger inputs usesEncoderStream.Summary: streaming read, early-stop, and append work directly on an uncompressed CBOR sequence; random access to row N needs an external offset index; a record stream halves the size but items are not self-contained, so random access additionally needs read-from-start or a shared structure catalog; and any generic compression forfeits all of these. HDF5/Parquet provide chunk+offset addressing natively (#109).
CBOR Packed (draft-ietf-cbor-packed-19)
Standards-Track IETF draft (draft-ietf-cbor-packed-19, Bormann/Guetschow): transforms CBOR into a packed form so "a separate decompression step is often unnecessary"; it standardizes the format, not a packer algorithm. Shared-item + argument tables; a "record" construct combines an array of keys with per-row arrays of values into maps (the tabular case). Positioned as a lightweight alternative to DEFLATE. cbor-x implements the string-table style (tag 51, draft-03 vintage), not this record construct.
cbor-x
github.com/kriszyp/cbor-x: CBOR (RFC 8949), typed arrays (RFC 8746), sequences (RFC 8742), string-table packing (draft-ietf-cbor-packed), and a proposed record-structures extension (structure once, referenced by same-shape objects; on by default), with streaming (
EncoderStream/DecoderStream,encodeAsIterable,decodeMultiple) and an externalstructurescatalog.Recommendation (to confirm)
Scope / relation
Transport only. Feeds the "Large and bulk data" guide and #109 (dataset storage/distributions); the encoder would live in the oold-python bundler/serialization (OO-LD/oold-python#30, OO-LD/oold-python#32).
Acceptance criteria
References
specs/tooling-tests/bench-matrix.mjs(size + total enc/dec),specs/tooling-tests/bench-streaming.mjs(streaming/partial access),specs/tooling-tests/bench-transport.mjs(initial raw+gzip+brotli)