Skip to content

GH-ISSUE_NUMBER: [Python][Parquet] Fix ambiguous dotted column projection - #53010

Open
Mohit-25-tech wants to merge 1 commit into
apache:mainfrom
Mohit-25-tech:fix/parquet-dotted-column-projection
Open

Mohit-25-tech wants to merge 1 commit into
apache:mainfrom
Mohit-25-tech:fix/parquet-dotted-column-projection

Conversation

@Mohit-25-tech

Copy link
Copy Markdown

What changes are included in this PR?

Fixes #52997

This PR fixes an inconsistent column-projection behavior in pyarrow.parquet.ParquetFile when a literal top-level column name collides with a nested field path.

Previously, requesting columns=["a.b"] could unexpectedly return both the literal column a.b and a struct column a containing nested field b.

The high-level pq.read_table() API returns only the literal column in this scenario.

Root cause

ParquetFile._build_nested_paths() builds projection lookup keys by joining physical path components with ".".

This makes two distinct paths share the same lookup key:

Literal top-level field: ["a.b"]    -> "a.b"
Nested field:           ["a", "b"] -> "a.b"

As a result, _get_column_indices() can select physical columns belonging to both fields.

Changes

  • Track exact top-level column names separately from nested-path prefixes.
  • Prioritize an exact top-level column match when resolving requested projections.
  • Preserve existing nested-field selection when no exact top-level match exists.
  • Update the relevant ParquetFile method documentation to explain the precedence rule.
  • Add regression tests covering:
    • Literal dotted names colliding with nested paths.
    • read(), read_row_group(), and iter_batches().
    • Deeper nested prefixes.
    • Non-conflicting dotted names.
    • Duplicate column selections.
    • Pandas index metadata handling.

Expected behavior

For a file containing a literal field a.b and a struct a with child b:

file.read(columns=["a.b"])

Now selects the literal top-level field a.b.

The nested field remains accessible by selecting its parent a.

This intentionally clarifies the precedence of ambiguous dotted names and aligns the result with pq.read_table() in the reproduced case.

Testing

Original bug reproduction

  • Independently reproduced with PyArrow 26.0.0 in Kaggle/Linux.
  • Also observed with locally installed PyArrow 16.1.0 and 24.0.0.

Proposed fix validation

  • 8 independent Kaggle checks passed using an adapted Python monkeypatch on PyArrow 26.0.0.
  • 4 focused pytest tests passed using the modified Python methods and locally installed PyArrow 24.0.0 binary reader.
  • Flake8 passed.
  • autopep8 diff check passed.
  • Python syntax compilation passed.
  • git diff --check passed.

Limitations: The full repository test suite could not run locally because the checkout does not contain a built pyarrow.lib. Pre-commit was unavailable. Full checkout and cross-platform validation remain pending in CI.

Compatibility considerations

This change intentionally modifies the behavior of ambiguous projections where a dotted top-level column name collides with a nested path.

Existing callers relying on both fields being selected for one ambiguous name may observe a different output schema.

Maintainer feedback on the proposed exact-name precedence is welcome.

AI assistance

OpenAI Codex assisted with investigation, implementation, regression-test development, and code review.

I independently reproduced the original issue in Kaggle, validated the proposed resolution using additional checks, and reviewed the changes and test results.

Related issue

Fixes #ISSUE_NUMBER

Copilot AI balanced review requested due to automatic review settings October 10, 2026 15:34

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@github-actions

Copy link
Copy Markdown

Thanks for opening a pull request!

This pull request has been automatically converted to a draft because its title doesn't match Arrow's required format.

If this is not a minor PR, could you open an issue for this pull request on GitHub? https://github.com/apache/arrow/issues/new/choose

Opening GitHub issues ahead of time contributes to the Openness of the Apache Arrow project.

Then could you also rename the pull request title in the following format?

GH-${GITHUB_ISSUE_ID}: [${COMPONENT}] ${SUMMARY}

or

MINOR: [${COMPONENT}] ${SUMMARY}

After updating the title, you can mark the pull request as ready for review.

See also:

@github-actions
github-actions Bot marked this pull request as draft October 10, 2026 15:34
@Mohit-25-tech
Mohit-25-tech marked this pull request as ready for review October 10, 2026 15:36
@Mohit-25-tech Mohit-25-tech changed the title GH-ISSUE_NUMBER: [Python][Parquet] Prefer exact top-level column names in ParquetFile projection GH-ISSUE_NUMBER: [Python][Parquet] Fix ambiguous dotted column projection Oct 10, 2026
@github-actions
github-actions Bot marked this pull request as draft October 10, 2026 15:40
@Mohit-25-tech
Mohit-25-tech marked this pull request as ready for review October 10, 2026 15:41
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Python][Parquet] ParquetFile projection includes extra column when dotted field name matches nested path

2 participants