Skip to content

feat: add parquet_file_metadata and parquet_page_index to datafusion-cli - #25640

Merged
kumarUjjawal merged 1 commit into
apache:mainfrom
o11y-one:feat/cli-parquet-file-metadata-page-index
Oct 5, 2026
Merged

kumarUjjawal merged 1 commit into
apache:mainfrom
o11y-one:feat/cli-parquet-file-metadata-page-index

Conversation

@jay-dee7

Copy link
Copy Markdown
Contributor

Which issue does this PR close?

Rationale for this change

datafusion-cli can't show a parquet file's footer metadata (writer, key-value metadata, footer size) or its page index. Checking bloom filter sizing, page index cost or writer metadata currently means using DuckDB or a custom tool. This PR covers items 2 and 3 of #25499; item 1 is in #25638.

What changes are included in this PR?

  • New parquet_file_metadata(path) table function, one row per file: created_by, version, num_rows, num_row_groups, key_value_metadata (a map, so key_value_metadata['ARROW:schema'] works) and footer_length.
  • New parquet_page_index(path) table function, one row per data page: the page location from the offset index (page_ordinal, first_row_index, offset, compressed_page_size) and the column index min_value, max_value and null_count. Files without a page index return no rows.
  • The path argument parsing is now shared by all three parquet functions.
  • Docs for both functions in docs/source/user-guide/cli/functions.md.

What is the testing strategy for this PR?

Snapshot tests test_parquet_file_metadata_works and test_parquet_page_index_works in datafusion-cli/src/main.rs, run against parquet-testing files. They cover an all-null page (NULL min/max) and UTF8 min/max shown as strings. The page values match the parquet-cli output in int32_with_null_pages.md, and footer_length matches the file's raw footer bytes.

Are there any user-facing changes?

Yes: two new datafusion-cli table functions, documented in the CLI functions guide. parquet_metadata output is unchanged and there are no breaking changes.

@github-actions github-actions Bot added the documentation Improvements or additions to documentation label Sep 23, 2026
@jay-dee7
jay-dee7 force-pushed the feat/cli-parquet-file-metadata-page-index branch from 880b6dd to 3df7c6f Compare September 23, 2026 02:19
@codecov-commenter

codecov-commenter commented Sep 23, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 82.51%. Comparing base (cee15b7) to head (b76fba2).
⚠️ Report is 62 commits behind head on main.

Additional details and impacted files
@@            Coverage Diff             @@
##             main   #25640      +/-   ##
==========================================
- Coverage   82.51%   82.51%   -0.01%     
==========================================
  Files        1141     1141              
  Lines      439773   439773              
  Branches   439773   439773              
==========================================
- Hits       362881   362872       -9     
- Misses      54950    54956       +6     
- Partials    21942    21945       +3     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@jay-dee7
jay-dee7 force-pushed the feat/cli-parquet-file-metadata-page-index branch 3 times, most recently from 31ee9e1 to b76fba2 Compare September 27, 2026 09:32
@jay-dee7

jay-dee7 commented Oct 1, 2026

Copy link
Copy Markdown
Contributor Author

@kumarUjjawal can you please review this one? or should I split this one to two PRs to make it easier to review?

@kumarUjjawal kumarUjjawal left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you @jay-dee7

let rbs = df.collect().await?;

assert_snapshot!(batches_to_string(&rbs), @r"
+-------------------------------------------------------+--------------+-----------+--------------+-----------------+--------+----------------------+-------------+------------+------------+

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

this is very cool -- thank you

@alamb alamb left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @jay-dee7 nd @kumarUjjawal -- this looks very useful

_ => format!("{val:?}"),
};

match index {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

As a follow on, we could also potentially use this:
https://docs.rs/parquet/latest/parquet/arrow/arrow_reader/statistics/struct.StatisticsConverter.html#method.row_group_mins

To convert the min/max values and then call the arrow cast kernel to turn them into strings

There is similar code for data page mins here:
https://docs.rs/parquet/latest/parquet/arrow/arrow_reader/statistics/struct.StatisticsConverter.html#method.data_page_mins

That would handle things like min/max dates better (I think this is just going to show the raw values rather than formatted as a date)

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@alamb thanks for sharing this, should I include it here? or in a follow-up PR?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think a follow up would be better.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

sure, will do, and thanks for being pro-active :)

Adds two datafusion-cli table functions for inspecting parquet files:

- `parquet_file_metadata(path)`: one row per file with the writer, format
  version, row and row group counts, key-value metadata (as a map) and
  footer length.
- `parquet_page_index(path)`: one row per data page in the page index,
  with the page location and the column index min, max and null count.

Part of apache#25499

Signed-off-by: jay-dee7 <me@jsdp.dev>
@jay-dee7
jay-dee7 force-pushed the feat/cli-parquet-file-metadata-page-index branch from b76fba2 to 1703b67 Compare October 5, 2026 06:54
@kumarUjjawal
kumarUjjawal enabled auto-merge October 5, 2026 06:57
@kumarUjjawal
kumarUjjawal added this pull request to the merge queue Oct 5, 2026
Merged via the queue into apache:main with commit e5fdc9d Oct 5, 2026
43 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation v56.0.0

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants