Skip to content

GH-50842: [Python] Convert dictionaries to objects without decoding when the dense array would overflow - #52547

Open
pearu wants to merge 1 commit into
apache:mainfrom
pearu:pearu/dict-to-object-on-overflow
Open

pearu wants to merge 1 commit into
apache:mainfrom
pearu:pearu/dict-to-object-on-overflow

Conversation

@pearu

@pearu pearu commented Oct 9, 2026

Copy link
Copy Markdown
Contributor

Rationale for this change

Converting a DictionaryArray with string/binary values to a NumPy object array (to_numpy, np.asarray), or to pandas as the child of a nested type, decodes the dictionary into a dense array first. When the decoded data exceeds the 32-bit offset limit this fails with Take operation overflowed binary array capacity (a segfault or garbage before #50841), although the result, a NumPy object array, has no such limit.

What changes are included in this PR?

  • DecodingWouldOverflow predicts from the dictionary offsets and the indices whether decoding a chunk would exceed the limit the take kernel enforces. Two cheap bounds settle the common case without reading the rows.
  • Only when it would overflow, the dictionary is left undecoded and ObjectWriterVisitor::Visit(const DictionaryType&) converts each dictionary value to a Python object once and references it per index (the deduplicate_objects idea, without the memo table since the dictionary already holds the unique values).
  • Every other conversion takes the unchanged decode path. Measured on 10⁶–10⁷ rows with dictionaries of 10²–10⁶ entries: time and memory are within noise of the current code.

Not covered: map keys and items decode before reaching this check (ConvertMap), so a map with an overflowing dictionary child still raises.

Are these changes tested?

Yes: test_dictionary_to_numpy_without_decoding (plus null index, null dictionary value and binary values) and test_list_of_dictionary_without_decoding (the nested path, also with several chunks). The dense form in the tests is 2.1 GB, built from 16385 references to one 128 KiB value, so the tests themselves need under 1 MB.

Are there any user-facing changes?

Conversions that raised now succeed. In that regime, repeated values in the result refer to one Python object, as to_pandas(deduplicate_objects=True) already does.

AI usage disclosure

Developed with AI assistance (Claude Code): the approach was worked out and the implementation, tests and benchmarks were generated under my direction and reviewed line by line by me. I own and can debug these changes.

@github-actions

github-actions Bot commented Oct 9, 2026

Copy link
Copy Markdown

⚠️ GitHub issue #50842 has been automatically assigned in GitHub to PR creator.

@github-actions

github-actions Bot commented Oct 9, 2026

Copy link
Copy Markdown

⚠️ GitHub issue #50842 has no components, please add labels for components.

…ding when the dense array would overflow

Converting a string/binary dictionary array to a NumPy object array, or to
pandas as the child of a nested type, decodes the dictionary into a dense
array first. When the decoded data exceeds the 32-bit offset limit this
fails although the object array has no such limit.

Predict the overflow from the dictionary offsets and the indices, and only
then convert the dictionary values to Python objects once and reference
them per index instead of decoding. All other conversions are unchanged.

Drafted by Claude Code (an AI agent) and reviewed & approved by pearu.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant