Skip to content

[WIP][POC] Pfor encoding - #50088

Draft
prtkgaur wants to merge 6 commits into
apache:mainfrom
prtkgaur:pfor-encoding
Draft

prtkgaur wants to merge 6 commits into
apache:mainfrom
prtkgaur:pfor-encoding

Conversation

@prtkgaur

@prtkgaur prtkgaur commented Jun 3, 2026 •

Copy link
Copy Markdown

Doc : https://docs.google.com/document/d/1ZZOtxmq6K8pNU0npijfSglTJVkspXL5GLDKPSGj9HlA/edit?tab=t.0

Thanks for opening a pull request!

If this is your first pull request you can find detailed information on how to contribute here:

Please remove this line and the above text before creating your pull request.

Rationale for this change

What changes are included in this PR?

Are these changes tested?

Are there any user-facing changes?

This PR includes breaking changes to public APIs. (If there are any breaking changes to public APIs, please explain which changes are breaking. If not, you can remove this.)

This PR contains a "Critical Fix". (If the changes fix either (a) a security vulnerability, (b) a bug that caused incorrect or invalid data to be produced, or (c) a bug that causes a crash (even when the API contract is upheld), please provide explanation. If not, you can remove this.)

@github-actions

github-actions Bot commented Jun 3, 2026

Copy link
Copy Markdown

Thanks for opening a pull request!

If this is not a minor PR. Could you open an issue for this pull request on GitHub? https://github.com/apache/arrow/issues/new/choose

Opening GitHub issues ahead of time contributes to the Openness of the Apache Arrow project.

Then could you also rename the pull request title in the following format?

GH-${GITHUB_ISSUE_ID}: [${COMPONENT}] ${SUMMARY}

or

MINOR: [${COMPONENT}] ${SUMMARY}

See also:

A frame-of-reference decoder adds the frame to every value right after
unpacking it. Doing that in a second pass reads and writes the whole
output again, so the scalar and SIMD unpackers now take an optional bias
and add it before each store. With no bias the existing paths are
unchanged, including the plain memcpy used when the bit width equals the
output width. The tests run the biased form across widths, offsets,
epilogues and every dispatched instruction set.
PFOR packs each vector of 1024 values at a width chosen for most of them
and stores the rest as patched exceptions. This adds frame and width
selection, the exception section, an endian-independent page layout, and
decoding that checks every length against the page before reading it.
int32 and int64 are supported. The codec builds under both CMake and
Meson, with unit tests and a microbenchmark.
Registers PFOR for INT32 and INT64 and wires it into the encoder and
decoder factories. Reads work for dense and nullable columns,, and the
decoder can also return one vector at a time.
The tests write and read whole files, including nulls, batched reads,
malformed pages and the factory paths. The benchmark compares PFOR with
the existing integer encodings on synthetic columns modeled on common
analytic data. It skips any compression codec the build does not include
rather than aborting.
The PFOR templates are instantiated in libarrow and used from libparquet
and its tests. Without export markers this only links on ELF platforms,
where symbols are visible by default.
Lists the supported integer types and the Preview status. Writers never
pick PFOR on their own; a column has to ask for it.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants