Repository navigation
[Python] Support DictionaryArray -> numpy/pandas conversion when the decoded data exceeds the 32-bit offset limit #50842
Description
Activity
@pearu ...I’d be interested in contributing to this. The direct dictionary → Python object approach (b) makes sense to me, especially since it avoids the large intermediate allocation. If you’re not already planning to implement it, I’d be happy to work on the implementation and regression tests.
Thanks for the offer, and good to have a second voice for (b) — that's the direction I'd settled on too.
I am planning to implement it, so I'll take this one. The prerequisite fix (#50841) has landed, so the path is clear.
If you'd like something adjacent to pick up: (a) is still potentially worth doing as a complement, for code paths that genuinely need a dense array and so aren't covered by (b) —
strings_to_categoricalbeing the obvious one. That's separable from this issue. Otherwise, review on the PR once it's up would be very welcome.Worth noting no maintainer has commented on the direction yet, so (b) is still two contributors' preference rather than a settled decision — input from an Arrow maintainer would be welcome before I get too far in.
🤖 Drafted by Claude Code (an AI agent) and reviewed & approved by pearu.
@rok @jorisvandenbossche following the suggestion at the dev sync, I looked at whether the
deduplicate_objectsmachinery could solve this. It can't as it stands, for two reasons:- It runs too late.
ConvertChunkedArrayToPandasfirst decodes the dictionary into a dense Arrow array (DecodeDictionaries→compute::Cast(dictionary → value_type), which is atakeof the dictionary by the indices). Only that densestringarray then reaches the object writer,ConvertAsPyObjects, wherededuplicate_objectslives. The decode is exactly the step that overflows the 32-bit offsets, so deduplication never gets to run. - On the
np.asarray/to_numpypath it isn't enabled at all.to_numpysetsdecode_dictionaries,zero_copy_onlyandto_numpy, and leavesdeduplicate_objectsat its C++ default offalse. (to_pandasenables it, but reason 1 still applies.)
Approach: let the object writer consume the dictionary array directly instead of a decoded copy: convert each dictionary value to a Python object once, then fill the output by
Py_INCREF-ing the object for each index. This is theunique_values+Py_INCREFpart ofdeduplicate_objects, minus the memo table, because the dictionary already is the set of unique values.- Reason 1 goes away because there is no dense intermediate any more; nothing before the object stage can overflow.
- Reason 2 goes away because the dictionary path doesn't depend on
deduplicate_objectsbeing enabled; it applies whenever a binary-like dictionary is decoded into an object array, which coversto_numpyas is.
Plan: one PR in
arrow_to_pandas.cc, with regression tests for the two entry points that hit the decode:to_numpy/np.asarray, andto_pandasof nested dictionaries (e.g.list<dictionary<string>>).I'll prepare the PR unless there are objections to the direction.
- It runs too late.
Yes, that sounds good!
- added a commit that references this issue
on Oct 9, 2026 Follow-up on the plan above: the PR (#52547) narrows it. Benchmarking the always-on variant (dictionary path for every
string/binarydictionary) against the current decode showed that it is faster and far lighter for ordinary dictionaries (10²–10⁴ entries: 2–4× faster, ~100× less memory), but worse in three situations: dictionaries beyond ~10⁵ entries (the table of Python objects no longer fits in cache, up to 3× slower), many chunks sharing one large dictionary (the dictionary is converted once per chunk; 1000 chunks × 10³ rows over a 10⁵-entry dictionary took 10 s instead of ~50 ms), and small slices of arrays with large dictionaries.So the PR takes the dictionary path only when decoding would actually overflow: a cheap check on the dictionary offsets and indices predicts whether
Takewould reject the decode, and everything else goes through the existing code unchanged (measured at parity on 10⁶–10⁷ rows with 10²–10⁶-entry dictionaries). Using the dictionary path more widely, where it is the better choice, can be a separate change.
🤖 Drafted by Claude Code (an AI agent) and reviewed & approved by pearu.
Describe the enhancement requested
Once #50840 is fixed (#50841), converting a dictionary array whose decoded form exceeds
INT32_MAXbytes raises cleanly instead of crashing:That is a strict improvement over a segfault, but the conversion still fails on data that is entirely representable in the output. The result of
np.asarrayhere is a numpy object array of Pythonstr— a format with no 2 GiB limit. The failure comes purely from an intermediate step.Why it fails
Array.to_numpy(zero_copy_only=False)setsdecode_dictionaries=True. Inarrow_to_pandas.cc,ConvertChunkedArrayToPandas()then decodes by casting to the dictionary's value type:So a 50M × 50-byte decode has to fit in a 32-bit-offset
stringarray. It cannot — hence the error. But that dense array is a pure implementation artifact; nothing downstream needs it to be astringarray specifically.Approach (a): widen the intermediate type
Use
large_string/large_binaryas the decode target instead ofstring/binary.Pros
dense_typeto_numpyandto_pandaspaths togetherCons
strobjects for 50M indices — so peak memory stays very large even though the input holds a single distinct valueApproach (b): skip the dense intermediate entirely
For the object-output path, convert each dictionary value to a
PyObjectonce, then walk the indices andPy_INCREFthe corresponding object into the output array.Pros
Cons
DecodeDictionariesstrings_to_categorical) would still want (a)Recommendation
(b), because it addresses the actual redundancy rather than raising a ceiling. Dictionary-encoded data is used precisely when values repeat, so materializing one Python object per index rather than per dictionary entry is wasted work in every case — the overflow is just where it becomes fatal. It also improves the common path, not only the pathological one.
(a) remains worth keeping in reserve for any remaining code path that must produce a dense array; the two are complementary rather than exclusive.
I'm happy to implement (b), pending agreement on the direction.
Component(s)
Python
🤖 Drafted by Claude Code (an AI agent) and reviewed & approved by pearu.