Repository navigation
Conversation
|
|
|
|
…ding when the dense array would overflow Converting a string/binary dictionary array to a NumPy object array, or to pandas as the child of a nested type, decodes the dictionary into a dense array first. When the decoded data exceeds the 32-bit offset limit this fails although the object array has no such limit. Predict the overflow from the dictionary offsets and the indices, and only then convert the dictionary values to Python objects once and reference them per index instead of decoding. All other conversions are unchanged. Drafted by Claude Code (an AI agent) and reviewed & approved by pearu.
pearu
force-pushed
the
pearu/dict-to-object-on-overflow
branch
from
October 9, 2026 13:58
320d429 to
7ceca05
Compare
pearu
marked this pull request as ready for review
October 9, 2026 14:41
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Rationale for this change
Converting a
DictionaryArraywithstring/binaryvalues to a NumPy object array (to_numpy,np.asarray), or to pandas as the child of a nested type, decodes the dictionary into a dense array first. When the decoded data exceeds the 32-bit offset limit this fails withTake operation overflowed binary array capacity(a segfault or garbage before #50841), although the result, a NumPy object array, has no such limit.What changes are included in this PR?
DecodingWouldOverflowpredicts from the dictionary offsets and the indices whether decoding a chunk would exceed the limit thetakekernel enforces. Two cheap bounds settle the common case without reading the rows.ObjectWriterVisitor::Visit(const DictionaryType&)converts each dictionary value to a Python object once and references it per index (thededuplicate_objectsidea, without the memo table since the dictionary already holds the unique values).Not covered: map keys and items decode before reaching this check (
ConvertMap), so a map with an overflowing dictionary child still raises.Are these changes tested?
Yes:
test_dictionary_to_numpy_without_decoding(plus null index, null dictionary value andbinaryvalues) andtest_list_of_dictionary_without_decoding(the nested path, also with several chunks). The dense form in the tests is 2.1 GB, built from 16385 references to one 128 KiB value, so the tests themselves need under 1 MB.Are there any user-facing changes?
Conversions that raised now succeed. In that regime, repeated values in the result refer to one Python object, as
to_pandas(deduplicate_objects=True)already does.AI usage disclosure
Developed with AI assistance (Claude Code): the approach was worked out and the implementation, tests and benchmarks were generated under my direction and reviewed line by line by me. I own and can debug these changes.