Skip to content

fix: raise InvalidOperationError for mismatched list.contains items on pyarrow-backed backends - #3915

Open
jonasdedden wants to merge 4 commits into
narwhals-dev:mainfrom
jonasdedden:upstream/list-contains
Open

jonasdedden wants to merge 4 commits into
narwhals-dev:mainfrom
jonasdedden:upstream/list-contains

Conversation

@jonasdedden

@jonasdedden jonasdedden commented Sep 6, 2026 •

Copy link
Copy Markdown
Contributor

Description

Now that #4001 added list.contains for PyArrow, pandas and Dask, an item whose type can't be compared with the inner dtype (e.g. a bool or str item on an int list) leaked pyarrow's ArrowNotImplementedError. It now raises InvalidOperationError, matching Polars and ArrowSeries.__contains__.

Also adds tests (follow-up from narwhals-daft differential testing) for:

  • Numeric coercion: 2.0 matches ints, 1.5 does not, 300 overflows Int8, and 1 matches a Float64 list. Mixing numeric kinds is unspecified (Decide on a cross-backend coercion for is_in & __contains__ #3900), so these pin today's behaviour rather than guarantee it. Polars 2.0 requires an explicit cast, hence the xfail there.
  • Mismatched items raise: bool, str and precision-mismatched datetime items raise InvalidOperationError, as in Polars. SQL backends coerce the item instead, and pyarrow-backed backends cast the datetime precision, hence the xfails.
  • All-null inner list: [None, None] contains nothing. Skipped on Ibis and PySpark, which cannot infer the type of an all-null column.

What type of PR is this? (check all applicable)

  • 💾 Refactor
  • ✨ Feature
  • 🐛 Bug Fix
  • 🔧 Optimization
  • 📝 Documentation
  • ✅ Test
  • 🐳 Other

Related issues

AI assistance

  • No AI tools were used for this PR.
  • AI tools were used.

Checklist

  • Code follows style guide (ruff)

  • Tests added

  • Documented the changes (N/A, error type only)

  • If this is your first PR to narwhals, attach a screenshot of pytest passing locally (not CI):

    PYTEST_ADDOPTS="--numprocesses=logical" \
    make run-ci DEPS="--extra pandas --extra dask --group core-tests --group sklearn --group plugins" \
    CMD="pytest tests --cov=src --cov=tests --runslow --constructors=pandas,pandas[nullable],pandas[pyarrow],pyarrow,polars[eager],polars[lazy],dask,duckdb,sqlframe"

Comment thread tests/expr_and_series/list/contains_test.py Outdated
@jonasdedden
jonasdedden marked this pull request as ready for review September 10, 2026 09:33

@FBruzzesi FBruzzesi left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @jonasdedden , just a few minor comments 🙌🏼

Comment thread tests/expr_and_series/list/contains_test.py Outdated
Comment thread tests/expr_and_series/list/contains_test.py Outdated
Comment thread tests/expr_and_series/list/contains_test.py Outdated
Comment thread tests/expr_and_series/list/contains_test.py Outdated
Comment thread tests/expr_and_series/list/contains_test.py Outdated
@jonasdedden

Copy link
Copy Markdown
Contributor Author

Hey @FBruzzesi, just a question, this PR isn't blocked by anything, right? :)

1dc8a02e52075792e0a08dbee71a0d00.jpg

@FBruzzesi FBruzzesi changed the title test: cover list.contains numeric coercion and all-null inner test: cover list.contains numeric coercion and all-null inner Sep 25, 2026
@FBruzzesi

Copy link
Copy Markdown
Member

Hey @jonasdedden sorry for the slow feedback on this one. I was about to merge, then I got ghosts from pyspark. Running it locally got me 4 errors:

FAILED tests/expr_and_series/list/contains_test.py::test_contains_all_null_inner_expr[pyspark] - pyspark.errors.exceptions.base.PySparkValueError: [CANNOT_DETERMINE_TYPE] Some of types cannot be dete...
FAILED tests/expr_and_series/list/contains_test.py::test_contains_numeric_coercion_expr[pyspark-float_matches_int] - AssertionError: Mismatch at index 3, key a: None != False
FAILED tests/expr_and_series/list/contains_test.py::test_contains_numeric_coercion_expr[pyspark-overflow] - AssertionError: Mismatch at index 0, key a: None != False
FAILED tests/expr_and_series/list/contains_test.py::test_contains_numeric_coercion_expr[pyspark-non_integer] - AssertionError: Mismatch at index 0, key a: None != False

I am after a long day, so I won't investigate much further right now. Hopefully during the weekend I can dive deeper, yet feel free to provide a fix for these cases in either this or a separate PR 🙏🏼

@jonasdedden

Copy link
Copy Markdown
Contributor Author

@FBruzzesi thanks for noticing!! Converted this PR to draft for now, as it first requires a new 16LOC bugfix PR #3990 to be merged.

@jonasdedden
jonasdedden force-pushed the upstream/list-contains branch from ef43b60 to df56d9b Compare September 28, 2026 14:52
@jonasdedden
jonasdedden marked this pull request as ready for review September 28, 2026 15:02
@jonasdedden

Copy link
Copy Markdown
Contributor Author

Actually, let's keep this PR on hold again for #4001 to be merged first? What do you think @FBruzzesi ?

@FBruzzesi

Copy link
Copy Markdown
Member

Ok, I guess we are ready to come back to this one 🤣
There are some conflicts to resolve first, then I will do a final pass 🙏🏼

…ns` items

PyArrow, pandas and Dask leaked pyarrow's `ArrowNotImplementedError` when the item
cannot be compared with the inner dtype (e.g. a bool or str item on an int list).
Raise `InvalidOperationError` instead, as Polars and `ArrowSeries.__contains__` do.
…-null inner

Pin how each backend handles mixed numeric kinds (unspecified, see narwhals-dev#3900), check that
mismatched bool/str/datetime-precision items raise where Polars does, and that an
all-null inner list contains nothing. SQL backends coerce mismatched items instead,
and pyarrow-backed ones coerce the datetime precision, hence the xfails.
@jonasdedden
jonasdedden force-pushed the upstream/list-contains branch from df56d9b to 9e34ef1 Compare October 3, 2026 11:07
@jonasdedden jonasdedden changed the title test: cover list.contains numeric coercion and all-null inner fix: raise InvalidOperationError for mismatched list.contains items on pyarrow-backed backends Oct 3, 2026
@jonasdedden

Copy link
Copy Markdown
Contributor Author

@FBruzzesi I seem to not have covered exact exception handling in edge cases in #4001, so this PR unfortunately became a fix: PR again. But this is now isolated to commit 2c64548.

@FBruzzesi FBruzzesi left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks a lot @jonasdedden, and thanks for your patience with the back and forth on #3990 and #4001 🙌🏼

The fix is right. While testing it I found a few close cousins that still let pyarrow errors through:

  • a NaN item on a non-float list
  • a timezone-mismatched datetime
  • a column with zero rows

The inline comments explain each one. They all come down to the same fix: check the item's type once, before any data is processed. I've attached the full diff with that change, plus tests for the new cases and a couple of small cleanups in the test file.

There is probably one more case I didn't touch which is the OverflowError for items of 2**63 or more, since it comes from #4001 and needs a different fix. I'm happy to open a follow-up issue for it.

Feel free to push back on any of it! Once it's in, I think this is good to go 🚀

Comment thread src/narwhals/_arrow/utils.py Outdated
Comment on lines +623 to +627
try:
matches = pc.equal(values, lit(item))
except pa.ArrowNotImplementedError as exc:
msg = f"Unable to compare item of type {type(item)} with list of type {block.type}."
raise InvalidOperationError(msg) from exc

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for tracking this down! While testing it I found a few more ways pyarrow errors still get through, all close cousins of the one you fixed:

  • NaN item on a non-float list: item=float("nan") takes the pc.is_nan branch, which is outside the try. On a List(String) or List(Boolean) it raises a raw ArrowNotImplementedError, while Polars raises InvalidOperationError.
  • Timezone mismatch: a tz-aware datetime item on a naive List(Datetime) (or the reverse) raises ArrowInvalid, which the except doesn't catch. ArrowSeries.__contains__ catches (ArrowInvalid, ArrowNotImplementedError, ArrowTypeError), so we can match that.
  • Zero rows: with no rows, _list_blocks yields nothing, so pc.equal never runs and a mismatched item quietly returns an empty result. Polars raises here too, because it checks the dtype, not the data.

All three points can be address with a single check in list_contains, before any blocks, against an empty array of the inner type. pyarrow chooses its comparison function from the types alone, so an empty array is enough to trigger the error, and it costs about 3.5µs per call. In the attached diff:

  • The null / NaN / equal branching moves into a small _item_matches helper, so the check and the per-block code can't drift apart.
  • list_contains calls _item_matches(pa.array([], value_type), item) once and maps all three pyarrow errors to InvalidOperationError.
  • _list_block_contains no longer needs a try. Once the type check has passed, any pyarrow error during the real computation depends on the data, so we shouldn't relabel it as "can't compare".
diff --git a/src/narwhals/_arrow/utils.py b/src/narwhals/_arrow/utils.py
index 2014943e6..4b1606785 100644
--- a/src/narwhals/_arrow/utils.py
+++ b/src/narwhals/_arrow/utils.py
@@ -575,10 +575,27 @@ def list_contains(array: ChunkedArrayAny, item: NonNestedLiteral) -> ChunkedArra
     A running count of matches over the flattened values changes within a list iff the
     list holds a match. That keeps this linear, where a group-by per list would sort.
     """
+    list_type = cast("pa.ListType[Any] | pa.LargeListType[Any]", array.type)
+    try:
+        # Probe an empty array, so that a mismatched `item` raises even without rows.
+        _item_matches(pa.array([], list_type.value_type), item)
+    except (pa.ArrowInvalid, pa.ArrowNotImplementedError, pa.ArrowTypeError) as exc:
+        msg = (
+            f"Unable to compare item of type {type(item)} with list of type {list_type}."
+        )
+        raise InvalidOperationError(msg) from exc
     blocks = [_list_block_contains(block, item) for block in _list_blocks(array)]
     return pa.chunked_array(blocks, pa.bool_())
 
 
+def _item_matches(values: ArrayAny, item: NonNestedLiteral) -> pa.BooleanArray:
+    if item is None:
+        return pc.is_null(values)
+    if isinstance(item, float) and math.isnan(item):
+        return pc.is_nan(values)  # NaN matches NaN, as in Polars.
+    return pc.equal(values, lit(item))
+
+
 def _list_blocks(array: ChunkedArrayAny) -> Iterator[ListArrayAny]:
     """Split `array` into runs of whole lists, of up to about `_LIST_BLOCK_VALUES` values.
 
@@ -615,16 +632,7 @@ def _list_block_contains(block: ListArrayAny, item: NonNestedLiteral) -> pa.Bool
     offsets = pc.subtract(offsets, lit(first, offsets.type))  # type: ignore[arg-type]
     ends, starts = offsets.slice(1), offsets.slice(0, len(block))
 
-    if item is None:
-        matches = pc.is_null(values)
-    elif isinstance(item, float) and math.isnan(item):
-        matches = pc.is_nan(values)  # NaN matches NaN, as in Polars.
-    else:
-        try:
-            matches = pc.equal(values, lit(item))
-        except pa.ArrowNotImplementedError as exc:
-            msg = f"Unable to compare item of type {type(item)} with list of type {block.type}."
-            raise InvalidOperationError(msg) from exc
+    matches = _item_matches(values, item)
 
     # Counting modulo 2^k stays exact within lists shorter than 2^k, so the narrowest
     # type that fits the longest list is enough.

Let me know what you think of this

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Applied, with one change: the check now runs against pa.nulls(1, value_type) instead of an empty array. pyarrow skips its timezone check on empty input, so with the empty array the tz mismatch still leaked ArrowInvalid.

One open question: Polars doesn't raise on a tz mismatch, it coerces (True on every version from 0.20.4 to 1.44). We now raise InvalidOperationError, like ArrowSeries.__contains__. Is that fine, or should we follow Polars? We could also add a test for it or leave it for the follow-up.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

timezones are a minefield territory. Happy to keep as follow up to be honest, unless the change is quite straightforward

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Also great catch on the empty array, thanks!

Comment thread tests/expr_and_series/list/contains_test.py
jonasdedden and others added 2 commits October 3, 2026 18:10
… data

Probe the item against a single null of the inner type, so that a NaN item
on a non-float list, a timezone mismatch and empty input all raise
`InvalidOperationError` instead of leaking pyarrow errors or returning
silently. An empty probe isn't enough, as pyarrow skips its timezone check
on empty arrays.

Tests: add a NaN item and a zero-row case, select xfails by item value
instead of param id, and xfail the NaN case on polars<1.28, which doesn't
raise there.

Co-authored-by: Francesco Bruzzesi <42817048+FBruzzesi@users.noreply.github.com>
… coverage skips it

The full-coverage job runs polars>=1.28, so the branch never ran there.
@jonasdedden

Copy link
Copy Markdown
Contributor Author

Thanks a lot @jonasdedden, and thanks for your patience with the back and forth on #3990 and #4001 🙌🏼

All good! I anyways was always extremely disappointed or annoyed by the very lackluster or inconsistent handling of list-types in various dataframe libraries (with examples such as pyarrow,pandas and Dask just outright missing important stuff), so if narwhals can be a library that now unifies these behaviors finally, that is very nice! ❤️

@FBruzzesi FBruzzesi added the fix label Oct 4, 2026

@FBruzzesi FBruzzesi left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @jonasdedden 🙏🏼

@MarcoGorelli

Copy link
Copy Markdown
Member

nice one, thanks both! 🙏

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants