feat: add parquet_file_metadata and parquet_page_index to datafusion-cli - #25640
Conversation
880b6dd to
3df7c6f
Compare
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #25640 +/- ##
==========================================
- Coverage 82.51% 82.51% -0.01%
==========================================
Files 1141 1141
Lines 439773 439773
Branches 439773 439773
==========================================
- Hits 362881 362872 -9
- Misses 54950 54956 +6
- Partials 21942 21945 +3 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
31ee9e1 to
b76fba2
Compare
|
@kumarUjjawal can you please review this one? or should I split this one to two PRs to make it easier to review? |
kumarUjjawal
left a comment
There was a problem hiding this comment.
Thank you @jay-dee7
| let rbs = df.collect().await?; | ||
|
|
||
| assert_snapshot!(batches_to_string(&rbs), @r" | ||
| +-------------------------------------------------------+--------------+-----------+--------------+-----------------+--------+----------------------+-------------+------------+------------+ |
There was a problem hiding this comment.
this is very cool -- thank you
alamb
left a comment
There was a problem hiding this comment.
Thanks @jay-dee7 nd @kumarUjjawal -- this looks very useful
| _ => format!("{val:?}"), | ||
| }; | ||
|
|
||
| match index { |
There was a problem hiding this comment.
As a follow on, we could also potentially use this:
https://docs.rs/parquet/latest/parquet/arrow/arrow_reader/statistics/struct.StatisticsConverter.html#method.row_group_mins
To convert the min/max values and then call the arrow cast kernel to turn them into strings
There is similar code for data page mins here:
https://docs.rs/parquet/latest/parquet/arrow/arrow_reader/statistics/struct.StatisticsConverter.html#method.data_page_mins
That would handle things like min/max dates better (I think this is just going to show the raw values rather than formatted as a date)
There was a problem hiding this comment.
@alamb thanks for sharing this, should I include it here? or in a follow-up PR?
There was a problem hiding this comment.
I think a follow up would be better.
There was a problem hiding this comment.
sure, will do, and thanks for being pro-active :)
Adds two datafusion-cli table functions for inspecting parquet files: - `parquet_file_metadata(path)`: one row per file with the writer, format version, row and row group counts, key-value metadata (as a map) and footer length. - `parquet_page_index(path)`: one row per data page in the page index, with the page location and the column index min, max and null count. Part of apache#25499 Signed-off-by: jay-dee7 <me@jsdp.dev>
b76fba2 to
1703b67
Compare
Which issue does this PR close?
Rationale for this change
datafusion-clican't show a parquet file's footer metadata (writer, key-value metadata, footer size) or its page index. Checking bloom filter sizing, page index cost or writer metadata currently means using DuckDB or a custom tool. This PR covers items 2 and 3 of #25499; item 1 is in #25638.What changes are included in this PR?
created_by,version,num_rows,num_row_groups,key_value_metadata(a map, sokey_value_metadata['ARROW:schema']works) andfooter_length.page_ordinal,first_row_index,offset,compressed_page_size) and the column indexmin_value,max_valueandnull_count. Files without a page index return no rows.docs/source/user-guide/cli/functions.md.What is the testing strategy for this PR?
Snapshot tests
test_parquet_file_metadata_worksandtest_parquet_page_index_worksindatafusion-cli/src/main.rs, run againstparquet-testingfiles. They cover an all-null page (NULL min/max) and UTF8 min/max shown as strings. The page values match the parquet-cli output inint32_with_null_pages.md, andfooter_lengthmatches the file's raw footer bytes.Are there any user-facing changes?
Yes: two new
datafusion-clitable functions, documented in the CLI functions guide.parquet_metadataoutput is unchanged and there are no breaking changes.