Skip to content

Repository files navigation

parquet-engine

A small, fast JSON ⇄ Parquet encode/decode engine written in Rust (Arrow + Parquet via PyO3), distributed as a portable abi3 wheel — no PyArrow / pandas dependency.

It exists so projects can read and write Parquet (with column-aware encoding for float-heavy data) without pulling the large PyArrow wheel.

Install

pip install parquet-engine

The published wheels are abi3 (built once per platform, work on CPython ≥ 3.10).

Usage

The engine speaks newline-delimited JSON (one JSON object per line):

import json
import parquet_engine as pe

rows = [{"x": 0.1, "y": 1.5, "g": 0}, {"x": 0.2, "y": 1.6, "g": 1}]
ndjson = ("\n".join(json.dumps(r) for r in rows) + "\n").encode()

parquet_bytes = pe.encode(ndjson)            # NDJSON  -> Parquet (bytes)
ndjson_back   = pe.decode(parquet_bytes)     # Parquet -> NDJSON (str)

API

encode(data: bytes, encodings: dict[str, str] | None = None,
       compression_level: int = 3) -> bytes
decode(data: bytes) -> str
  • encode infers a schema from the NDJSON, writes Parquet with ZSTD compression (compression_level, default 3), and chooses a per-column encoding.
  • decode reads Parquet bytes back into newline-delimited JSON.

Numeric columns, without the JSON

For blocks that are wholly numeric, the JSON round trip is the expensive part: formatting and reparsing every value costs far more than the Parquet write itself, and the intermediate text runs several times the size of the data. The binary gateway skips it.

import numpy as np
import parquet_engine as pe

x = np.linspace(0.0, 1.0, 1_000_000)
g = np.arange(1_000_000, dtype=np.int64)

blob = pe.encode_columns({"x": ("f8", x.tobytes()), "g": ("i8", g.tobytes())})

columns = pe.decode_columns(blob)
dtype, raw = columns["x"]
back = np.frombuffer(raw, dtype=dtype)      # bit-for-bit identical to `x`
encode_columns(columns: dict[str, tuple[str, bytes]],
               encodings: dict[str, str] | None = None,
               compression_level: int = 3) -> bytes
decode_columns(data: bytes) -> dict[str, tuple[str, bytes]]

A column is a (dtype, data) pair: dtype is "f8" (float64) or "i8" (int64), and data is that column's raw native-endian bytes. Columns are written in iteration order and must share a length. decode_columns hands back pairs in the same shape, so its output feeds straight back into encode_columns.

Both gateways write ordinary Parquet, so they interoperate: decode_columns reads anything encode wrote whose columns are all float64 or int64, and vice versa. They differ on non-finite floats. encode stores NaN and ±Inf as null (JSON has no spelling for them), so ±Inf comes back as NaN; encode_columns stores the values themselves and round-trips them bit for bit. Reading a JSON-written file through decode_columns maps those nulls back to NaN.

decode_columns raises on a column that is neither float64 nor int64, and on an integer column containing nulls, since neither has a raw representation. Use decode for those.

Column encoding

By default the encoding is inferred per column:

Column type Encoding
floating point (f16/f32/f64) BYTE_STREAM_SPLIT (dictionary off)
everything else dictionary

This matches the empirical result that byte-stream-split beats dictionary on full-precision floats, while dictionary wins on low-cardinality columns.

Override per column with the encodings map:

pe.encode(ndjson, {"x": "bss", "y": "byte_stream_split",
                   "g": "dictionary", "id": "plain"})

Accepted values: "bss" / "byte_stream_split", "dictionary" / "dict", "plain". Unknown values raise ValueError. Columns absent from the map fall back to inference.

Building from source

maturin develop          # build + install into the active venv
maturin build --release  # produce a wheel

Windows + MSVC note: if your shell sets CC/CXX to MinGW (e.g. msys2), the bundled zstd C build will mismatch the MSVC linker. Build with those unset so the cc crate auto-detects cl.exe: env -u CC -u CXX maturin develop.

License

MIT

About

Encoder/Decoder for parquet files.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages