Conversation
|
Synced with latest main (now 13 commits ahead of the original base, including the pyarrow 24→25 upgrade). pyarrow 25 deprecated the SortOptions-level Verified locally: unit test, docker-based REST catalog integration test ( |
eeef9cd to
3660095
Compare
|
This pull request has been marked as stale due to 30 days of inactivity. It will be closed in 1 week if no further activity occurs. If you think that's incorrect or this pull request requires a review, please simply write any comment. If closed, you can revive the PR at any time and @mention a reviewer or discuss it on the dev@iceberg.apache.org list. Thank you for your contributions. |
Related: #271 and #3848
Rationale for this change
PyIceberg accepts table sort metadata but currently writes unsorted files and hard-codes
sort_order_id=None. This prevents readers from safely using sort-order-aware pruning and leaves manifest metadata inconsistent with users' write intent.What changes?
Honors table sort orders for materialized
pyarrow.Tablewrites when every sort field uses an identity transform and one consistent null placement:WriteTaskDataFile.sort_order_idUnsupported transforms, nested/missing fields, mixed null placement, and streaming
RecordBatchReaderwrites preserve current behavior: data is not claimed as sorted and the file sort-order ID remains null. A warning explains why.Are these changes tested?
sort_order_idagainst a REST catalogPYTHONPATH=. uv run pytest tests/io/test_pyarrow.py -k "sort_table_for_identity_sort_order or write_sorted_data_files_per_partition" -q— passedruff checkon changed files — passedgit diff --check— passedAre there any user-facing changes?
Yes. Tables with an identity-transform sort order now write physically sorted data files carrying the truthful
sort_order_id; previously all files were written unsorted with a null sort-order ID. Unsupported sort orders keep the previous behavior. Changelog label requested.Tooling note: developed with assistance from DS v4 Pro. I reviewed and verified the changes.