Antalya 25.8 - Backport of #99521 Parquet UUID support - #2446
arthurpassos wants to merge 1 commit into
Conversation
- Implements C++ reader/writer logic for UUIDs in Parquet/Arrow - Adds stateless test coverage using clickhouse-local to prevent CI collisions Backport to `antalya-25.8` of ClickHouse#99521.
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 80463d6110
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
|
||
| struct ConverterUUID | ||
| { | ||
| using Statistics = StatisticsFixedStringRef; |
There was a problem hiding this comment.
Copy UUID statistics instead of retaining batch pointers
When a UUID page or column chunk is encoded from multiple batches, StatisticsFixedStringRef retains pointers into swapped_buf, but every subsequent getBatch() resizes and overwrites that buffer. Earlier extrema are therefore mutated or left dangling before page/column statistics are serialized, producing incorrect min/max values that can cause predicate pushdown to skip rows containing matches. Use owning fixed-size statistics storage for UUIDs rather than the reference-based implementation.
Useful? React with 👍 / 👎.
| else | ||
| { | ||
| UUID res; | ||
| std::memcpy(&res, fixed_binary_array.GetValue(i), 16); |
There was a problem hiding this comment.
Validate UUID value buffers before reading them
For a malformed Arrow FixedSizeBinary chunk whose declared length exceeds its data buffer, the byte-width check still succeeds and GetValue(i) returns an address beyond the allocation, which this memcpy then reads. This newly added path also reserves from the unvalidated length first, unlike the existing fixed-binary readers; validate every chunk with the fixed-size buffer helper before reserving or accessing values so truncated Arrow input is rejected as INCORRECT_DATA rather than causing an out-of-bounds read.
Useful? React with 👍 / 👎.
|
Grok:
|
Backport to
antalya-25.8of ClickHouse#99521.Changelog category (leave one):
Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md):
Adds native support for importing and exporting UUID data types in Arrow and Parquet formats. Users can now directly query and transfer UUID data between ClickHouse and other data tools without requiring manual string conversions or workarounds. Automated logical inference for top-level UUIDs, and support for explicit schema hint for nested UUIDs.
Documentation entry for user-facing changes
...
CI/CD Options
Exclude tests:
Regression jobs to run: