Skip to content

fix: align list.contains null handling with Polars on lazy backends - #3996

Merged
FBruzzesi merged 5 commits into
narwhals-dev:mainfrom
jonasdedden:upstream/list-contains-nulls
Sep 28, 2026
Merged

FBruzzesi merged 5 commits into
narwhals-dev:mainfrom
jonasdedden:upstream/list-contains-nulls

Conversation

@jonasdedden

@jonasdedden jonasdedden commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

Description

Relands #3990 (reverted in #3992). list.contains now matches Polars around nulls on every lazy backend:

case Polars before
[1, None], .contains(2) False PySpark: null
[1, None], .contains(None) True DuckDB, Ibis, SQLFrame: null; PySpark: raises
.contains(None) on Polars < 1.24 one value per row a single null
[] as the only row, .contains(None), Polars 1.28-1.29 False True

For a non-null item, PySpark keeps array_contains and turns its null into false for non-null lists:
coalesce(array_contains(a, item), when(a IS NOT NULL, false)).

contains(None) needs a "does the list hold a null" check. Which check is fastest depends on the inner dtype and on where the null is.

PySpark: exists only

exists(a, x -> x IS NULL) works for every inner type and stops at the first null. size(array_intersect(a, array(NULL))) > 0 is faster on integer lists (3-5x end to end from parquet), but it rejects non-orderable inner types (map, variant, calendar interval, structs holding maps). Picking it via typeof(a) doesn't help either: Spark type-checks the unused branch before optimising it away, so the query still fails. A schema lookup to choose per dtype would cost a round trip, which no spark-like expression does today.

DuckDB, Ibis, SQLFrame

DuckDB 1.5, which also runs the SQL that Ibis and SQLFrame generate in our tests. Same setup, lists without nulls:

inner dtype array_position len > list_count filter nulls (Ibis) array_compact (SQLFrame) list_sort
Int64, Float64, Decimal, String, Date 0.1-0.3 0.9 1.6-1.7 2.5-4.4 48-80
Struct 22 1.1 1.9 5.8 150
List 30 0.9 1.7 6.3 176
  • DuckDB → len(a) > list_count(a): about 1 ns for every inner dtype, and correct on every DuckDB version. array_position(a, NULL) saves under 1 ns on flat inner types, but is 20-30x slower for structs and lists, and misses their nulls before DuckDB 1.4.
  • Ibis → len(filter(a, x -> x IS NULL)) > 0: Ibis has no null-aware search or count.
  • SQLFrame → size(a) > size(array_compact(a)): SQLFrame has no exists, and its array_intersect drops nulls on DuckDB.
  • Polars < 1.30 → len(a) > len(drop_nulls(a)): list.contains(None) returns a single null before 1.24, and on 1.28-1.29 returns True for [] when it is the only row (a length-1 Series, pl.lit([]), or a frame filtered down to one row). As a side effect, nested inner dtypes also work on 1.28-1.29; from 1.30 Polars itself raises for them, so that test is skipped there.

What type of PR is this?

  • 💾 Refactor
  • ✨ Feature
  • 🐛 Bug Fix
  • 🔧 Optimization

Related issues

AI assistance

  • No AI tools were used for this PR.
  • AI tools were used.

Checklist

  • Code follows style guide (ruff)
  • Tests added
  • Documented the changes (N/A, behaviour now matches Polars)

Relands narwhals-dev#3990 (reverted in narwhals-dev#3992). On PySpark, `contains(None)` uses
`exists`: `array_intersect` is only faster for integer lists, and
`sort_array` is O(n log n). DuckDB uses `array_position(..., NULL)`
only from 1.4, since 1.3 misses nulls among nested elements.
PySpark: `array_intersect` is only fast for integer lists, and `exists`
stops at the first null, so pick per dtype with `typeof`, which Spark folds
when planning. DuckDB: `array_position(..., NULL)` is 20-30x slower for
struct/list elements and misses their nulls before 1.4, so use
`len > list_count` on every version.
@jonasdedden
jonasdedden marked this pull request as draft September 27, 2026 12:39
Polars 2.0 handles them lazily but panics eagerly, while 1.28-1.x raise, so the outcome is Polars' own and varies by version.
@jonasdedden

Copy link
Copy Markdown
Contributor Author

@FBruzzesi I'm reasonably sure that this implementation now is close to ideal for Spark, DuckDB and co.; Have a look at above benchmarks to see the full reasoning.

Currently working on pandas, pyarrow, Dask and co. compatibility, as this isn't working at all yet.

@jonasdedden jonasdedden changed the title fix: align list.contains null handling with Polars on lazy backends feat: support list.contains on PyArrow and pandas, align its null handling with Polars Sep 27, 2026
@jonasdedden
jonasdedden marked this pull request as ready for review September 27, 2026 13:23
@jonasdedden

jonasdedden commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor Author

@FBruzzesi @MarcoGorelli okay, now this PR increased a bit in scope (revertible with 0a85094 if too much):

  • list.contains is now implemented for PyArrow and pandas & co.
  • its null handling generally is aligned with Polars

Reason for that increased scope primarily is that the new tests required a lot of xfail markers for PyArrow and co., and actually implementing it wasn't so much code.

Dask is still xfail, but I think it potentially could work there too.

EDIT: Sorry for the continuous back and forth. This PR now again is fix-only, and I use #4001 to implement list.contains for actually all of PyArrow, pandas and Dask. Split them up into multiple to keep scope smaller.

@jonasdedden
jonasdedden force-pushed the upstream/list-contains-nulls branch from 0a85094 to c9d6432 Compare September 27, 2026 18:31
@jonasdedden jonasdedden changed the title feat: support list.contains on PyArrow and pandas, align its null handling with Polars fix: align list.contains null handling with Polars on lazy backends Sep 27, 2026
@FBruzzesi FBruzzesi added fix nested data `list`, `struct`, etc labels Sep 28, 2026

@FBruzzesi FBruzzesi left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @jonasdedden, the benchmarks are thorough and I was able to reproduce them. Two issues though:

  1. On PySpark, the typeof dispatch breaks contains(None) for non-orderable inner types. Spark type-checks the dead array_intersect branch before folding it, so array<map<...>>, structs containing maps, and array<variant> fail with DATATYPE_MISMATCH.INVALID_ORDERING_TYPE:>

    from pyspark.sql import SparkSession
    import narwhals as nw
    
    spark = SparkSession.builder.getOrCreate()
    df = nw.from_native(spark.sql("SELECT array(map('k', 1), NULL) AS a"))
    df.select(nw.col("a").list.contains(None)).to_native().show()
    pyspark.errors.exceptions.captured.AnalysisException: [DATATYPE_MISMATCH.INVALID_ORDERING_TYPE]
    Cannot resolve "array_intersect(a, array(NULL))" due to data type mismatch:
    The `array_intersect` does not support ordering on type "MAP<STRING, INT>". SQLSTATE: 42K09;
    'Project [CASE WHEN typeof(a#0) IN (array<tinyint>,array<smallint>,array<int>,array<bigint>)
              THEN '`>`('array_size(array_intersect(a#0, cast(array(null) as array<map<string,int>>))), 0)
              ELSE exists(a#0, lambdafunction(isnull(lambda x_1#2), lambda x_1#2, false)) END AS a#1]
    

    Plain exists works for all of them, and returns true here. I'd go with exists only: integer lists pay ~40x in the worst case, but a failing query is worse than a slow one.

  2. On Polars 1.28 and 1.29, [].list.contains(None) returns True (fixed in 1.30), so the new tests fail there. Extending the len > len(drop_nulls) workaround to < 1.30 fixes it. It also lets nested inner dtypes work on 1.28 and 1.29, so the test skip can start at 1.30.

Comment thread src/narwhals/_polars/expr.py Outdated
Comment thread src/narwhals/_spark_like/expr_list.py
…1.28-1.29

Spark type-checks the dead `array_intersect` branch of the `typeof` dispatch,
so non-orderable inner types (map, structs holding maps, variant, calendar
interval) failed analysis. Polars 1.28-1.29 return `True` for a single empty
list, so use `len > len(drop_nulls)` for `contains(None)` before 1.30.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ba9JftUrcct4rBVY5qsB1C
@jonasdedden

Copy link
Copy Markdown
Contributor Author

@FBruzzesi oh man, thank you so much for your investigation again! Sorry, I'm arguably only working with limited availability on this, so although there are thorough investigations there is a lot that slips through 🥲

Your findings were reproduced (not in its entirety, but mostly) in another session, I'll link a response. Haven't had a possibility to further investigate on my own yet, please take the new commit with caution!

Thanks, both confirmed and applied.

PySpark: reproduced INVALID_ORDERING_TYPE for map, struct-with-map and variant, and calendar intervals (make_interval) fail the same way. Since analysis type-checks the dead branch before optimisation removes it, there's no safe SQL-level dispatch, so it's exists only now, plus a slow-marked PySpark regression test with array. On cost: I measured exists at 3.4–5.5× slower than array_intersect end to end from parquet (6–16× on cached data), so ~40× holds for the operator itself but a real query pays less.

Polars: applied your restructure. One nuance: on 1.28–1.29 the wrong True only appears for single-row inputs (length-1 Series, pl.lit([]), or a frame filtered to one row). With multi-row frames the empty list gives False, which is why the existing tests pass on 1.28.1/1.29.0 for me. I added a test that filters down to one empty list: it fails on 1.28.1/1.29.0 before the fix and passes on 0.20.4 through 1.31 after.

Failing CI is because of some fluke I think:

cause: Failed to fetch: https://pypi.org/simple/pytest-cov/
cause: HTTP status server error (503 Service Unavailable) for url

@FBruzzesi FBruzzesi left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@jonasdedden no worries, nothing to be sorry about! We are doing this in a stretch of time! Left couple more comments for the tests, there should be nothing else to add in the main codebase 🙏🏼

Maybe update the PR description if you fancy to align with the latest implementation status

Comment thread tests/expr_and_series/list/contains_test.py
Comment thread tests/expr_and_series/list/contains_test.py Outdated
…rder-independent

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ba9JftUrcct4rBVY5qsB1C
@jonasdedden

Copy link
Copy Markdown
Contributor Author

Changed PR description accordingly. This was a lot of running around in circles, hehe 😄

@FBruzzesi FBruzzesi left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @jonasdedden - happy we could figure out a solution to this issue 🙏🏼

@FBruzzesi
FBruzzesi merged commit b4b3b3f into narwhals-dev:main Sep 28, 2026
44 checks passed
jonasdedden pushed a commit to jonasdedden/narwhals that referenced this pull request Sep 28, 2026
…arwhals-dev#3996

`test_contains_none_expr` from narwhals-dev#3996 checks the same rows (and more) now that
SQL backends and Polars<1.24 match Polars for a null item.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ba9JftUrcct4rBVY5qsB1C
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

fix nested data `list`, `struct`, etc

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants