Repository navigation
fix: align list.contains null handling with Polars on lazy backends - #3996
Conversation
Relands narwhals-dev#3990 (reverted in narwhals-dev#3992). On PySpark, `contains(None)` uses `exists`: `array_intersect` is only faster for integer lists, and `sort_array` is O(n log n). DuckDB uses `array_position(..., NULL)` only from 1.4, since 1.3 misses nulls among nested elements.
PySpark: `array_intersect` is only fast for integer lists, and `exists` stops at the first null, so pick per dtype with `typeof`, which Spark folds when planning. DuckDB: `array_position(..., NULL)` is 20-30x slower for struct/list elements and misses their nulls before 1.4, so use `len > list_count` on every version.
Polars 2.0 handles them lazily but panics eagerly, while 1.28-1.x raise, so the outcome is Polars' own and varies by version.
|
@FBruzzesi I'm reasonably sure that this implementation now is close to ideal for Spark, DuckDB and co.; Have a look at above benchmarks to see the full reasoning. Currently working on pandas, pyarrow, Dask and co. compatibility, as this isn't working at all yet. |
list.contains null handling with Polars on lazy backends|
EDIT: Sorry for the continuous back and forth. This PR now again is fix-only, and I use #4001 to implement |
0a85094 to
c9d6432
Compare
list.contains null handling with Polars on lazy backends
FBruzzesi
left a comment
There was a problem hiding this comment.
Thanks @jonasdedden, the benchmarks are thorough and I was able to reproduce them. Two issues though:
-
On PySpark, the
typeofdispatch breakscontains(None)for non-orderable inner types. Spark type-checks the deadarray_intersectbranch before folding it, soarray<map<...>>, structs containing maps, andarray<variant>fail withDATATYPE_MISMATCH.INVALID_ORDERING_TYPE:>from pyspark.sql import SparkSession import narwhals as nw spark = SparkSession.builder.getOrCreate() df = nw.from_native(spark.sql("SELECT array(map('k', 1), NULL) AS a")) df.select(nw.col("a").list.contains(None)).to_native().show()
pyspark.errors.exceptions.captured.AnalysisException: [DATATYPE_MISMATCH.INVALID_ORDERING_TYPE] Cannot resolve "array_intersect(a, array(NULL))" due to data type mismatch: The `array_intersect` does not support ordering on type "MAP<STRING, INT>". SQLSTATE: 42K09; 'Project [CASE WHEN typeof(a#0) IN (array<tinyint>,array<smallint>,array<int>,array<bigint>) THEN '`>`('array_size(array_intersect(a#0, cast(array(null) as array<map<string,int>>))), 0) ELSE exists(a#0, lambdafunction(isnull(lambda x_1#2), lambda x_1#2, false)) END AS a#1]Plain
existsworks for all of them, and returnstruehere. I'd go withexistsonly: integer lists pay ~40x in the worst case, but a failing query is worse than a slow one. -
On Polars 1.28 and 1.29,
[].list.contains(None)returnsTrue(fixed in 1.30), so the new tests fail there. Extending thelen > len(drop_nulls)workaround to< 1.30fixes it. It also lets nested inner dtypes work on 1.28 and 1.29, so the test skip can start at 1.30.
…1.28-1.29 Spark type-checks the dead `array_intersect` branch of the `typeof` dispatch, so non-orderable inner types (map, structs holding maps, variant, calendar interval) failed analysis. Polars 1.28-1.29 return `True` for a single empty list, so use `len > len(drop_nulls)` for `contains(None)` before 1.30. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ba9JftUrcct4rBVY5qsB1C
|
@FBruzzesi oh man, thank you so much for your investigation again! Sorry, I'm arguably only working with limited availability on this, so although there are thorough investigations there is a lot that slips through 🥲 Your findings were reproduced (not in its entirety, but mostly) in another session, I'll link a response. Haven't had a possibility to further investigate on my own yet, please take the new commit with caution!
Failing CI is because of some fluke I think: cause: Failed to fetch: |
FBruzzesi
left a comment
There was a problem hiding this comment.
@jonasdedden no worries, nothing to be sorry about! We are doing this in a stretch of time! Left couple more comments for the tests, there should be nothing else to add in the main codebase 🙏🏼
Maybe update the PR description if you fancy to align with the latest implementation status
…rder-independent Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ba9JftUrcct4rBVY5qsB1C
|
Changed PR description accordingly. This was a lot of running around in circles, hehe 😄 |
FBruzzesi
left a comment
There was a problem hiding this comment.
Thanks @jonasdedden - happy we could figure out a solution to this issue 🙏🏼
…arwhals-dev#3996 `test_contains_none_expr` from narwhals-dev#3996 checks the same rows (and more) now that SQL backends and Polars<1.24 match Polars for a null item. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ba9JftUrcct4rBVY5qsB1C

Description
Relands #3990 (reverted in #3992).
list.containsnow matches Polars around nulls on every lazy backend:[1, None],.contains(2)Falsenull[1, None],.contains(None)Truenull; PySpark: raises.contains(None)on Polars < 1.24null[]as the only row,.contains(None), Polars 1.28-1.29FalseTrueFor a non-null item, PySpark keeps
array_containsand turns itsnullintofalsefor non-null lists:coalesce(array_contains(a, item), when(a IS NOT NULL, false)).contains(None)needs a "does the list hold a null" check. Which check is fastest depends on the inner dtype and on where the null is.PySpark:
existsonlyexists(a, x -> x IS NULL)works for every inner type and stops at the first null.size(array_intersect(a, array(NULL))) > 0is faster on integer lists (3-5x end to end from parquet), but it rejects non-orderable inner types (map, variant, calendar interval, structs holding maps). Picking it viatypeof(a)doesn't help either: Spark type-checks the unused branch before optimising it away, so the query still fails. A schema lookup to choose per dtype would cost a round trip, which no spark-like expression does today.DuckDB, Ibis, SQLFrame
DuckDB 1.5, which also runs the SQL that Ibis and SQLFrame generate in our tests. Same setup, lists without nulls:
array_positionlen > list_countarray_compact(SQLFrame)list_sortlen(a) > list_count(a): about 1 ns for every inner dtype, and correct on every DuckDB version.array_position(a, NULL)saves under 1 ns on flat inner types, but is 20-30x slower for structs and lists, and misses their nulls before DuckDB 1.4.len(filter(a, x -> x IS NULL)) > 0: Ibis has no null-aware search or count.size(a) > size(array_compact(a)): SQLFrame has noexists, and itsarray_intersectdrops nulls on DuckDB.len(a) > len(drop_nulls(a)):list.contains(None)returns a singlenullbefore 1.24, and on 1.28-1.29 returnsTruefor[]when it is the only row (a length-1 Series,pl.lit([]), or a frame filtered down to one row). As a side effect, nested inner dtypes also work on 1.28-1.29; from 1.30 Polars itself raises for them, so that test is skipped there.What type of PR is this?
Related issues
list.containsnull handling with Polars on lazy backends #3990 (reverted in Revert "fix: alignlist.containsnull handling with Polars on lazy backends" #3992)InvalidOperationErrorfor mismatchedlist.containsitems on pyarrow-backed backends #3915AI assistance
Checklist