Skip to content

Transport / compression: evaluate CBOR (record structures / Packed) for large tabular instances #127

Description

@simontaurus

Motivation

An OO-LD instance can be large - a multi-column table or time series is an array of row records. This issue evaluates CBOR and its record/packing extensions as a binary transport that keeps the JSON / JSON-LD data model (see the guide's "Large and bulk data") and answers: (a) size and total encode/decode time vs other formats, (b) whether a record format for tabular data competes with generic compression, (c) whether partial read/write and streaming are possible (cf HDF5).

Example data model

OO-LD schema (structure + @context in one document) for the tabular case:

{
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "$id": "https://example.org/Measurement.schema.json",
  "title": "Measurement",
  "@context": {
    "ex": "https://example.org/vocab#",
    "qudt": "http://qudt.org/schema/qudt/",
    "metadata": "ex:metadata",
    "unit": "qudt:unit",
    "data": "ex:hasRow",
    "name": "ex:label",
    "value1": "ex:value1",
    "value2": "ex:value2",
    "value3": "ex:value3",
    "ok": "ex:ok"
  },
  "type": "object",
  "properties": {
    "metadata": {
      "type": "object",
      "properties": { "id": { "type": "string" }, "unit": { "type": "string" } }
    },
    "data": {
      "type": "array",
      "items": {
        "type": "object",
        "properties": {
          "name":   { "type": "string" },
          "value1": { "type": "number" },
          "value2": { "type": "number" },
          "value3": { "type": "number" },
          "ok":     { "type": "boolean" }
        },
        "required": ["name"]
      }
    }
  }
}

Instance excerpt (a valid OO-LD / JSON-LD document; row values illustrative):

{
  "@context": "https://example.org/Measurement.schema.json",
  "$schema": "https://example.org/Measurement.schema.json",
  "metadata": { "id": "run-1", "unit": "V" },
  "data": [
    { "name": "test0", "value1": 1.0, "value2": 2.0, "value3": 3.0, "ok": true  },
    { "name": "test1", "value1": 4.0, "value2": 5.0, "value3": 6.0, "ok": false }
  ]
}

The benchmark generates N such rows (deterministic pseudo-random floats at 3 decimals, plus the string and bool).

Size and total time (end-to-end)

Each row is a (format x compressor) combination. Ship B is the transported size; compressor = none is the uncompressed size. Total enc = build payload + compress; Total dec = decompress + parse. cbor-x 1.6.5, Node 25.6, gzip level 9, brotli with explicit quality + size hint; single-run timings. Script: specs/tooling-tests/bench-matrix.mjs.

N = 1,000 rows

Format Compressor Ship B Total enc ms Total dec ms
JSON none 79,896 0.37 0.57
JSON gzip-9 17,156 9.35 0.94
JSON brotli-q5 17,307 2.91 0.96
JSON brotli-q11 12,956 127.67 1.00
CBOR maps none 67,987 0.72 1.38
CBOR maps gzip-9 17,730 3.44 0.93
CBOR maps brotli-q5 16,267 3.52 0.99
CBOR maps brotli-q11 14,393 100.31 0.86
CBOR records none 40,030 0.49 0.68
CBOR records gzip-9 16,284 2.02 0.33
CBOR records brotli-q5 14,869 2.55 0.34
CBOR records brotli-q11 13,473 57.70 0.38
CBOR columnar none 33,044 0.35 0.29
CBOR columnar brotli-q5 11,797 1.84 0.37

N = 100,000 rows

Format Compressor Ship B Total enc ms Total dec ms
JSON none 8,174,316 43.23 75.93
JSON gzip-9 1,666,390 1092.78 93.33
JSON brotli-q5 330,556 159.85 94.19
JSON brotli-q11 262,522 22603.04 121.05
CBOR maps none 6,987,964 63.95 76.08
CBOR maps gzip-9 1,670,234 331.30 90.13
CBOR maps brotli-q5 233,871 157.24 98.99
CBOR maps brotli-q11 207,751 17017.28 80.63
CBOR records none 4,188,007 35.64 13.90
CBOR records gzip-9 1,558,171 194.72 27.79
CBOR records brotli-q5 191,613 111.52 23.20
CBOR records brotli-q11 198,310 8059.13 22.62
CBOR columnar none 3,489,054 10.19 5.59
CBOR columnar brotli-q5 94,198 47.48 18.14
CBOR columnar brotli-q11 63,899 1377.53 14.72

Encoding of each variant

  • CBOR maps - new Encoder({ useRecords: false }): one map per row, keys repeated.
  • CBOR records - new Encoder({ useRecords: true }): cbor-x's record-structures extension; the shared shape is written once, rows carry only values. Functional equivalent of the draft's tabular-record idea.
  • CBOR columnar + typed arrays - struct-of-arrays (Float64Array per numeric column); does not preserve the row-entity model.
  • A separate pack: true run (not shown above) verified on the wire that cbor-x's packer emits tag 51 and only deduplicates repeated strings (draft-ietf-cbor-packed-03 vintage), not the draft-19 tag-114 tabular-record transform; no packer implementing tag-114 was available, so that mechanism is not benchmarked.

Observations

  • Uncompressed: CBOR records is 0.51x JSON (4.19 vs 8.17 MB) and decodes about 4x faster with no decompression step (13.9 vs 75.9 ms at 100k).
  • Compressor choice dominates size and encode time. gzip level 9 has a 32 KB window and cannot deduplicate across a multi-MB file: JSON+gzip-9 = 1.67 MB vs JSON+brotli-q5 = 331 KB (5x smaller) at ~1/7 the encode time (160 vs 1093 ms). A large-window compressor (brotli, or zstd) is required for large tabular data.
  • brotli quality 11 is impractical at scale: 22.6 s (JSON) and 8.1 s (records) per 100k rows, for no gain over q5 on records (q5 191,613 B is smaller than q11 198,310 B and ~70x faster). brotli-q5 is the usable high-ratio setting.
  • Smallest structure-preserving: CBOR records + brotli-q5 = 191,613 B (0.58x of JSON+brotli-q5, 0.82x of CBOR maps+brotli-q5), total enc 112 ms, total dec 23 ms. Its decode stays ~4x faster than JSON/maps after compression (23 vs 94/99 ms).
  • Answer to (b): the record layout and generic compression are complementary, not competing. Under gzip they look redundant (records+gzip 1.56 MB vs JSON+gzip 1.67 MB), but that is a gzip-window artifact; under brotli the record layout is still ~40% smaller than compressed JSON and decodes ~4x faster. The record-only advantages compression cannot provide are fast decode, small size without a decompression step, and streamability.
  • Compression is opaque: a gzip/brotli payload must be fully decompressed before any row is read, forfeiting the streaming/early-stop/random-access properties in the next section. The trade is smallest-at-rest (records+brotli) vs streamable (raw records / sequence).
  • Smallest overall is columnar + brotli (94 KB at q5, 64 KB at q11), but it is not structure-preserving (no row-entity model, no nested/ragged rows) and corresponds to the analytics/external-container shape (Parquet/HDF5, Scientific profile: dataset storage model (datasets, distributions, storage modes); core ask: instance-level data references (data bundling) #109).

Downsides of record structures

  • cbor-x record structures are a proposed extension (tag 57344 range); a generic CBOR decoder returns positional arrays, not keyed objects.
  • The size win assumes rows share a shape; heterogeneous/ragged rows create multiple structures and erode it.
  • When compressed the payload is opaque (no streaming/partial access).
  • The size edge over CBOR maps after compression is modest (0.82x); the durable, compression-proof advantage is decode speed.

Streaming and partial access

Measured on the 100,000-row dataset, cbor-x 1.6.5, on uncompressed payloads. Script: specs/tooling-tests/bench-streaming.mjs.

CBOR sequence (each row an independent CBOR item, RFC 8742; 6,987,863 B, keys repeated):

  • Incremental read: decodeMultiple with a callback returning false stops after K items (verified: fired 10/10 for K=10); the remainder is not decoded.
  • Random access: with an external byte-offset index, one row decodes from its own 66-70 B slice (rows 0 / 500 / 99,999), no other bytes read.
  • Append: a 3-row tail encodes to 210 B and concatenates to the existing buffer; all 100,000 items still decode. Append cost is O(appended), no rewrite.

Record stream (useRecords: true, sequential: true; 4,187,896 B, keys once):

  • In-order decodeMultiple resolves field names for all rows.
  • An isolated tail item does not resolve field names (returns a positional array under tag 57344). Record items are not self-contained: reading item N requires reading from the start (to accumulate structures) or a shared structure catalog.

encodeAsIterable at N=100,000 emitted a single 4,187,947 B chunk (chunk boundary is buffer-size driven, not reached at ~4 MB); incremental emission for larger inputs uses EncoderStream.

Summary: streaming read, early-stop, and append work directly on an uncompressed CBOR sequence; random access to row N needs an external offset index; a record stream halves the size but items are not self-contained, so random access additionally needs read-from-start or a shared structure catalog; and any generic compression forfeits all of these. HDF5/Parquet provide chunk+offset addressing natively (#109).

CBOR Packed (draft-ietf-cbor-packed-19)

Standards-Track IETF draft (draft-ietf-cbor-packed-19, Bormann/Guetschow): transforms CBOR into a packed form so "a separate decompression step is often unnecessary"; it standardizes the format, not a packer algorithm. Shared-item + argument tables; a "record" construct combines an array of keys with per-row arrays of values into maps (the tabular case). Positioned as a lightweight alternative to DEFLATE. cbor-x implements the string-table style (tag 51, draft-03 vintage), not this record construct.

cbor-x

github.com/kriszyp/cbor-x: CBOR (RFC 8949), typed arrays (RFC 8746), sequences (RFC 8742), string-table packing (draft-ietf-cbor-packed), and a proposed record-structures extension (structure once, referenced by same-shape objects; on by default), with streaming (EncoderStream/DecoderStream, encodeAsIterable, decodeMultiple) and an external structures catalog.

Recommendation (to confirm)

  • Smallest structure-preserving with fast read and write: CBOR records + brotli-q5.
  • Write-heavy, streamable, or partial-access: raw CBOR records or a CBOR sequence (optionally gzip for a cheap reduction, losing streaming).
  • Avoid gzip level 9 and brotli-q11 for large tabular payloads (small window / seconds-scale encode).
  • Maximum compression or random access to bulk numeric data: columnar + typed arrays, or an external container (Parquet/HDF5 + descriptor, Scientific profile: dataset storage model (datasets, distributions, storage modes); core ask: instance-level data references (data bundling) #109).
  • OO-LD stays transport-agnostic: a Packed/record document unpacks to the same JSON, so the data model is unchanged; this is guidance, not a normative format.

Scope / relation

Transport only. Feeds the "Large and bulk data" guide and #109 (dataset storage/distributions); the encoder would live in the oold-python bundler/serialization (OO-LD/oold-python#30, OO-LD/oold-python#32).

Acceptance criteria

  • Benchmark the draft-19 tag-114 record packer (find or build a conformant packer).
  • Re-run with realistic domain data (high-precision doubles, real strings) and add zstd.
  • Confirm the recommendation: record CBOR + brotli-q5 vs external container.
  • Decide whether OO-LD references a binary transport in guidance or stays agnostic.
  • Document the streaming/append vs random-access vs smallest-at-rest trade-off in the "Large and bulk data" guide.

References

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    documentationImprovements or additions to documentationenhancementNew feature or requesttransportBinary transport, compression, serialization

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions