Skip to content

Merge upstream DataFusion 55.1.0 into spiceai-55 - #243

Merged
phillipleblanc merged 901 commits into
spiceai-55from
spiceai-55-patches-3
Oct 2, 2026
Merged

phillipleblanc merged 901 commits into
spiceai-55from
spiceai-55-patches-3

Conversation

@krinart

@krinart krinart commented Oct 2, 2026

Copy link
Copy Markdown

Merges upstream DataFusion 55.1.0 into the Spice line (spiceai-55, cut from spiceai-54); conflicts are resolved in the merge commit and recorded in its message.

On top of the merge:

The metadata-column listing pruning from spiceai-54 (#229) is not carried on this line yet, because Spice has no test that fails without it.

Please merge with a merge commit, not squash, so each patch stays its own commit.

Pinned by spiceai/spiceai#14612 (tracking: spiceai/spiceai#13570). Supersedes #238.

kosiew and others added 30 commits July 25, 2026 10:39
…che#22980)

## Which issue does this PR close?

* Part of apache#20835

## Rationale for this change

`FixedSizeList` containing `Struct` values was not handled by the
existing recursive nested adaptation logic used for schema evolution. As
a result, planner-time compatibility checks, nested cast detection, and
runtime casting did not support additive struct evolution within
`FixedSizeList` containers.

This change adds `FixedSizeList` support and verifies planner/runtime
parity so that planning allows exactly the cases runtime can adapt while
continuing to reject incompatible schema changes.

## What changes are included in this PR?

* Extend `cast_column` to support recursive casting of `FixedSizeList`
values when source and target list sizes match.
* Add `FixedSizeList` handling to:

  * `requires_nested_struct_cast`
  * `validate_data_type_compatibility`
* Implement recursive casting of nested `Struct` values contained in
`FixedSizeList`.
* Preserve planner/runtime parity by validating child type compatibility
before runtime fallback logic is applied.
* Add handling for null-parent `FixedSizeList` entries by masking hidden
child values before retrying casts, avoiding failures caused by
semantically inaccessible child data.
* Refactor list and list-view casting helpers to use Arrow `AsArray`
accessors.

## Are these changes tested?

Yes.

The following tests were added:

* `test_cast_fixed_size_list_struct`
* `test_validate_fixed_size_list_struct_compatibility`
*
`test_validate_fixed_size_list_struct_missing_non_nullable_field_rejected`
* `test_validate_fixed_size_list_struct_size_mismatch_rejected`
* `test_cast_fixed_size_list_struct_all_null`
*
`test_fixed_size_list_struct_planner_runtime_parity_on_incompatible_type`
*
`test_cast_fixed_size_list_struct_missing_non_nullable_field_runtime_rejected`
*
`test_cast_fixed_size_list_struct_ignores_hidden_child_values_for_null_parent`

Existing coverage in `test_requires_nested_struct_cast` was also
extended to include `FixedSizeList` cases.

These tests cover:

* Additive nullable nested-field evolution
* All-null and partially null list cases
* Incompatible nested type changes
* Non-nullable field addition rejection
* Planner/runtime parity validation

## Are there any user-facing changes?

No user-facing changes. This is an internal enhancement to nested schema
adaptation and casting behavior for `FixedSizeList<Struct>` types.

## LLM-generated code disclosure

This PR includes LLM-generated code and comments. All LLM-generated
content has been manually reviewed.

---------

Co-authored-by: Andrew Lamb <andrew@nerdnetworks.org>
… spark (apache#23766)

## Which issue does this PR close?

N/A

## Rationale for this change

Hex encoding was implemented six times across the workspace, at two
different levels of optimization.

Three copies used a fast byte-pair lookup table:

- `datafusion/spark/src/function/math/hex.rs`
- `datafusion/functions/src/string/to_hex.rs`
- `datafusion/functions/src/encoding/inner.rs` (via the `hex` crate)

The other three used a slower nibble-at-a-time loop pushing one
character at a time:

- `datafusion/functions/src/crypto/md5.rs`
- `datafusion/spark/src/function/hash/sha1.rs`
- `datafusion/spark/src/function/hash/sha2.rs`

Beyond the duplication, this split meant the digest functions were
paying for a slower encoder
than the one already sitting elsewhere in the tree. Consolidating on a
single implementation
removes the duplication and moves `md5`, `sha1`, and `sha2` onto the
fast path.

## What changes are included in this PR?

A new `datafusion_common::utils::hex` module holds the only hex encoder
in the workspace:

```rust
pub enum HexCase { Lower, Upper }

pub fn encode_bytes_into(bytes: &[u8], case: HexCase, out: &mut Vec<u8>);
pub fn encode_bytes(bytes: &[u8], case: HexCase) -> String;
pub fn encode_bytes_to_slice(bytes: &[u8], case: HexCase, out: &mut [u8]);
pub fn encode_u64(v: u64, case: HexCase, buf: &mut [u8; 16]) -> &[u8];
```

Four entry points rather than one because the call sites have genuinely
different output needs,
and forcing them through a single shape would cost an allocation
somewhere: `to_hex` writes
straight into a `StringArray` values buffer, Spark's `hex` appends to a
reused scratch `Vec`,
`encode` writes into a pre-sized slice, and the digest functions want an
owned `String`.

Migrated call sites, all producing bit-identical output:

| File | Change |
| --- | --- |
| `functions/src/string/to_hex.rs` | local table and two write helpers
deleted; eight trait impls collapsed to two macros |
| `spark/src/function/math/hex.rs` | two nibble tables, two lookup
tables, `build_hex_lookup` and `hex_int64` deleted |
| `functions/src/crypto/md5.rs` | local table and `hex_encode` deleted |
| `spark/src/function/hash/sha1.rs` | local table and inline nibble loop
deleted |
| `spark/src/function/hash/sha2.rs` | local table and `hex_encode`
deleted; eight call sites migrated |
| `functions/src/encoding/inner.rs` | both encode sites moved off the
`hex` crate |

Deliberately left alone:

- `hex::decode` / `hex::decode_to_slice` in `encoding/inner.rs`, and
Spark's `unhex`. The decode
direction has different semantics — Spark's `unhex` left-pads odd-length
input, the `hex` crate
does not — so unifying it is a separate question. The `hex` dependency
stays for those.
- `ScalarValue`'s binary `Display` impl in `common/src/scalar/mod.rs`,
which writes to a
  `fmt::Formatter` rather than a byte buffer, and is a cold path.

Three details worth a reviewer's attention:

- The two ancestors of `encode_u64` disagreed on zero: `to_hex` wrote
`'0'` into the caller's
buffer, Spark's `hex_int64` returned a `'static` `b"0"` that never
touched it. The shared version
always writes into the buffer and returns a subslice of it, so the
lifetime is uniform. Both
  callers still produce `"0"`.
- Spark's `hex_encode_bytes` guards large binary input with
`checked_mul(2)` + `try_reserve`,
returning a `DataFusionError` rather than aborting on allocation
failure. That guard stays at the
call site; `encode_bytes_into` performs no reservation of its own, so
behaviour is unchanged.
- The encoders are `#[inline(always)]`, not `#[inline]`. They are called
once per row from two
other crates, and plain `#[inline]` left them out-of-line across the
crate boundary. Measured
cost of that: +3% on Spark's byte paths and up to +18% on `to_hex`'s i32
path, i.e. the
refactor was a net regression on those benchmarks until the attribute
changed.

## Are these changes tested?

Every migrated function keeps its existing unit and sqllogictest
coverage, which is what pins
bit-identical output — in particular `spark/hash/sha1.slt` and
`sha2.slt` assert concrete digest
strings for all four SHA-2 bit lengths across both the scalar and array
paths, and `expr.slt`
covers `to_hex` and `md5`.

New unit tests in `common/src/utils/hex.rs` cover zero, `u64::MAX`,
single-nibble values, the
odd/even digit-count boundary, two's complement of negative input, empty
input, all 256 byte
values in both cases, appending into a non-empty buffer, and that a
reused scratch buffer never
leaks stale digits between calls. Tests cross-check against
`format!("{:x}")` rather than
restating the implementation.

Two Spark tests changed. `test_hex_int64` now drives `hex_encode_int64`
instead of the deleted
private `hex_int64`, keeping all ten cases including `i64::MIN`, `-1`,
and the uppercase
expectations. `test_hex_lookup_table_covers_all_bytes` was deleted — it
only cross-checked the raw
lookup tables, and `encode_bytes_covers_every_byte_value` now does that
exhaustively through the
public API. A test was added for Spark's lowercase byte path, which
previously had no coverage.

### Benchmarks

Criterion, `apache/main` @ `eef101769` as baseline. Median of the
reported change interval.

`datafusion/functions/benches/to_hex.rs`:

| Benchmark | 1024 | 4096 | 8192 |
| --- | --- | --- | --- |
| `i32_random` | −13.8% | −12.8% | −10.2% |
| `i64_random` | −8.8% | −9.5% | −8.1% |
| `i64_large_values` | −9.9% | −9.8% | −7.9% |

`scalar_i32` −2.3%, `scalar_i64` −1.8%.

`datafusion/spark/benches/sha2.rs`:

| Benchmark | 1024 | 4096 | 8192 |
| --- | --- | --- | --- |
| `array_binary_256` | −20.7% | −18.2% | −19.1% |
| `array_scalar_binary_256` | −14.6% | −13.2% | −13.1% |

`scalar/size=1` −3.3%.

`datafusion/spark/benches/hex.rs`:

| Benchmark | 1024 | 4096 | 8192 |
| --- | --- | --- | --- |
| `hex_int64` | −2.9% | −3.6% | −3.5% |
| `hex_int64_dict` | −3.2% | −1.5% | −0.2% |
| `hex_utf8` | −0.2% | −0.3% | +0.7% |
| `hex_binary` | −0.4% | −0.6% | −0.3% |

The `hex_utf8` and `hex_binary` paths already used the byte-pair table
before this PR, so they are
expected to be flat; they are.

`datafusion/functions/benches/crypto.rs`: `md5_array` −4.0%,
`md5_scalar` −3.7%. The `sha224` /
`sha256` / `sha384` / `sha512` cases in that file range from −0.1% to
+2.2%, but they exercise
`crypto/basic.rs`, which this PR does not modify and which contains no
hex encoding — those numbers
are run-to-run variance, not an effect of this change.

No number is quoted for Spark `sha1`: it has no benchmark, and its
change is the same substitution
applied to `md5` and `sha2`.

## Are there any user-facing changes?

No behaviour change — all migrated functions produce byte-identical
output.

`datafusion_common::utils::hex` is new public API on
`datafusion-common`.

---------

Co-authored-by: Andrew Lamb <andrew@nerdnetworks.org>
Co-authored-by: Jeffrey Vo <jeffrey.vo.australia@gmail.com>
…t_batch` memory (apache#23873)

## Which issue does this PR close?

- related to apache#23716


## Rationale for this change

It adds test coverage for two gaps found while reviewing
apache#23716.


## What changes are included in this PR?

Tests only, no functional change.
- Dictionary inputs
- memory usage on retract (make sure memory is released)

Note I moved `array_agg` cases out `aggregate.slt` as it is already more
than 9k lines long

## Are these changes tested?
They are only tests

## Are there any user-facing changes?

No. Tests only, no public API changes.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
…3881)

## Which issue does this PR close?

N/A

## Rationale for this change

Two Spark functions allocated a `String` per row purely to render a
small,
bounded amount of text:

- `bin` called `format!("{value:b}")` for every row.
- `char` called `ch.to_string()` for every row — a heap allocation for a
single
  character.

In both cases the output has a known upper bound (64 binary digits for
an `i64`,
4 bytes for a UTF-8 character), so the rendering fits in a stack buffer
and the
result can be appended straight to a pre-sized builder.

## What changes are included in this PR?

`math/bin.rs`:

- `spark_bin` now writes digits right-aligned into a caller-supplied
`[u8; 64]`
and returns a `&str` borrowed from it, instead of returning an owned
`String`.
- The `collect::<StringArray>()` becomes an explicit loop over a
  `StringBuilder::with_capacity`, sized at 8 digits per row.
- Negative values still render as their two's-complement bit pattern,
matching
`{:b}`. The digit loop is a `loop`, not a `while`, so zero renders as
`"0"`
  rather than the empty string.

`string/char.rs`:

- `ch.to_string()` becomes `ch.encode_utf8(&mut encoded)` against a
`[u8; 4]`
  hoisted out of the loop.

Output is unchanged in both cases.

## Are these changes tested?

Existing coverage pins the behaviour. `spark/math/bin.slt` asserts
concrete
output for the cases the rewrite had to get right: zero, negative
values,
`i64::MIN` (`-9223372036854775808`), `i64::MAX`, and `-2147483648` /
`-32768` widened from narrower integer types. `spark/string/char.slt`
covers
the negative-input empty string, the null path, and characters on both
sides of
the ASCII boundary (`char(256)` and above wrap via `% 256`).

All 119 `spark/math` and `spark/string` sqllogictest files pass, along
with the
258 `datafusion-spark` unit tests.

The `bin` benchmark used below, `datafusion/spark/benches/bin.rs`, is
added
separately in apache#23882 so the baseline can be measured on `main` before
this
change lands. It covers 1024 and 8192 rows with 20% nulls over two value
distributions: small values that render to a handful of digits, and
full-range
values that render to the maximum 64. `char` already had
`datafusion/spark/benches/char.rs` on `main`, so no benchmark change is
needed
for it.

### Benchmarks

Criterion, `apache/main` @ `f1ab86dad` as baseline. Median of the
reported
change interval.

| Benchmark | 1024 | 8192 |
| --- | --- | --- |
| `bin/small` | −75.7% | −73.7% |
| `bin/wide` | −47.8% | −49.8% |

`char` (1024 rows): −76.1%.

## Are there any user-facing changes?

No. Both functions produce byte-identical output; this is purely an
allocation
change.
…ts enabled (apache#23848)

## Which issue does this PR close?

<!--
We generally require a GitHub issue to be filed for all bug fixes and
enhancements and this helps us generate change logs for our releases.
You can link an issue to this PR using the GitHub syntax. For example
`Closes #123` indicates that this PR will close issue #123.
-->

- Closes apache#23847 .

## Rationale for this change

Fixes a correctness bug, not sure when that config used/desirable.

This is another thing I ran into during
apache#21585


## What changes are included in this PR?

1. New SLT test
2. Fix for bug in `datafusion-optimizer`

## Are these changes tested?

Existing tests, and a new SLT test verifying the config change doesn't
affect the query result and basic unit test.

## Are there any user-facing changes?

None

---------

Signed-off-by: Adam Gutglick <adamgsal@gmail.com>
…lter (apache#23901)

## Which issue does this PR close?

- Closes apache#23900.

## Rationale for this change

`push_down_filter` infers equi-key predicates across a join's ON keys
and pushes them to the opposite side. For a null-aware join (the
`LeftAnti` join produced by `NOT IN` with a nullable subquery), an outer
predicate on the left key like `outer.id > 5` is rewritten to `sub.id >
5` and pushed onto the subquery input. Since the inferred predicate must
be null-rejecting to be pushed, this drops the subquery's NULL rows and
breaks the three-valued `NOT IN` semantics — a NULL in the subquery key
must reach the join so the result is empty.

Same class of bug as apache#23848, in a different rule.

## What changes are included in this PR?

- Skip predicate inference in `infer_join_predicates` when
`join.null_aware` is set (mirrors the apache#23848 guard on
`FilterNullJoinKeys`).
- A `push_down_filter` unit test asserting no predicate is inferred onto
the subquery side of a null-aware `LeftAnti` join.
- SLT coverage for the failing query, plus a `prefer_hash_join = false`
/ multi-partition variant.

## Are these changes tested?

Yes — new unit test (verified it fails without the guard) and SLT cases.
The full optimizer lib suite passes.

## Are there any user-facing changes?

No, aside from the correctness fix.
## Which issue does this PR close?

- Part of apache#19241.
- Stacked on [apache#23311](apache#23311).
- Next in stack: apache#23015.
- Extracted from apache#19390.

## Rationale for this change

For very small `IN` lists, building or probing a hash table can be more
work than just comparing the input value with each constant.

For example, for `x IN (10, 20, 30)`, the fast path can behave like:

```text
x == 10 OR x == 20 OR x == 30
```

Because the list is tiny, those comparisons are cheap. The
implementation stores the constants in a fixed-size array and checks
them with a compact comparison chain.

“Branchless” here means the comparisons are combined without stopping at
the first match. That can be faster for these small fixed-width lists
because the CPU gets a predictable sequence of simple operations instead
of hash-table setup and probe logic.

For primitive values that are not already plain unsigned integers, this
PR keeps the logical Arrow type explicit and uses a matching same-width
comparison representation only inside the branchless filter. For
example, `Float16` uses `UInt16` storage, `Float32` uses `UInt32`
storage, and `TimestampNanosecond` uses `UInt64` storage. `Decimal128`
and `IntervalMonthDayNano` use their own 16-byte native representation.
This preserves bit-pattern equality while relying on Arrow's native
primitive compatibility rules: timestamp timezone metadata and
Decimal128 precision/scale metadata may differ, while incompatible
primitive representations remain rejected.

## What changes are included in this PR?

- Adds a const-generic `BranchlessFilter` for small primitive `IN`
lists.
- Adds thresholds for when this path is used:
  - up to 16 values for 1-byte types
  - up to 8 values for 2-byte types
  - up to 32 values for 4-byte types
  - up to 16 values for 8-byte types
  - up to 4 values for 16-byte types
- Keeps dispatch concrete and explicit in `strategy.rs`.
- Maps each optimized logical type to the comparison representation used
by the branchless filter:
  - `Int8` -> `UInt8`
  - `Int16`, `Float16` -> `UInt16`
  - `Int32`, `Float32`, `Date32`, `Time32` -> `UInt32`
- `Int64`, `Float64`, `Date64`, `Time64`, `Timestamp`, `Duration` ->
`UInt64`
- `Decimal128`, `IntervalMonthDayNano` -> their native 16-byte
representation
- Leaves larger 1-byte and 2-byte lists on the existing bitmap filters.
- Leaves larger 4-byte and 8-byte lists on the existing hash/generic
paths.
- Leaves wider primitive types such as `Decimal256` and unsupported
complex types on the generic path.
- Keeps the same `IN` / `NOT IN` null behavior as the rest of the stack.
- Adds focused coverage for branchless null handling, signed boundary
values, slices, Float16/Float32/Float64 bit patterns, compatible
timestamp/Decimal128 metadata, incompatible timestamp units,
IntervalMonthDayNano values, and same-width wrong-type probe rejection.

## Are these changes tested?

Yes.

- `cargo fmt --all -- --check`
- `cargo test -p datafusion-physical-expr expressions::in_list --lib`
- `cargo test -p datafusion-physical-expr --bench in_list_strategy
--no-run`
- `cargo clippy --all-targets --all-features -- -D warnings`

## Are there any user-facing changes?

No. This is an internal performance optimization only.

## Local benchmark snapshot

Built and run with `release-nonlto`, filtered to the relevant small
primitive-list rows:

```bash
cargo bench -p datafusion-physical-expr --profile release-nonlto --bench in_list_strategy -- <filter> --save-baseline <baseline>
```

Filters used: `narrow_integer`, `primitive/i32/small_list`,
`primitive/i64/small_list`, `f32/small_list`, `timestamp_ns/small_list`,
and `interval_month_day_nano/small_list`.

Method: directly compared Criterion's raw sample minima (`min(time /
iterations)`) from `sample.json`. Lower is better; changes within +/-5%
are treated as noise.

Compared baselines:
[apache#23311](apache#23311) ->
[apache#23014](apache#23014)

Relevant scope: small primitive-list rows.

Summary: 39 relevant rows, 28 faster, 0 slower, 11 within +/-5%.

Largest relevant deltas:

| Benchmark | Before | After | Change |
|---|---:|---:|---:|
| `timestamp_ns/small_list/list=4/match=50%` | 46.55 us | 3.17 us |
-93.2% (14.69x faster) |
| `f32/small_list/list=4/match=50%` | 33.93 us | 3.04 us | -91.0%
(11.15x faster) |
| `primitive/i32/small_list/list=4/match=50%` | 32.63 us | 3.08 us |
-90.5% (10.58x faster) |
| `primitive/i64/small_list/list=4/match=50%` | 33.55 us | 3.18 us |
-90.5% (10.54x faster) |
| `timestamp_ns/small_list/list=4/match=0%` | 19.57 us | 3.18 us |
-83.8% (6.16x faster) |
| `f32/small_list/list=4/match=0%` | 18.14 us | 3.05 us | -83.2% (5.95x
faster) |
| `primitive/i32/small_list/list=4/match=0%` | 17.00 us | 3.04 us |
-82.1% (5.59x faster) |
| `primitive/i64/small_list/list=4/match=0%` | 17.12 us | 3.22 us |
-81.2% (5.31x faster) |
| `primitive/i32/small_list/list=16/match=50%/NOT_IN` | 31.98 us | 7.26
us | -77.3% (4.41x faster) |
| `nulls/primitive/i32/small_list/list=16/match=50%/nulls=20%` | 29.35
us | 7.32 us | -75.1% (4.01x faster) |
| `timestamp_ns/small_list/list=16/match=50%` | 45.32 us | 11.79 us |
-74.0% (3.84x faster) |
| `nulls/primitive/i32/small_list/list=16/match=50%/nulls=50%` | 25.89
us | 7.31 us | -71.8% (3.54x faster) |
| `nulls/primitive/i32/small_list/list=16/match=50%/nulls=20%/NOT_IN` |
26.05 us | 7.42 us | -71.5% (3.51x faster) |
| `interval_month_day_nano/small_list/list=4/match=50%` | 52.94 us |
15.52 us | -70.7% (3.41x faster) |
| `f32/small_list/list=32/match=50%` | 38.78 us | 13.27 us | -65.8%
(2.92x faster) |
| `primitive/i64/small_list/list=16/match=50%` | 29.46 us | 11.76 us |
-60.1% (2.50x faster) |

<details>
<summary>Full relevant table (39 rows)</summary>

| Benchmark | Before | After | Change |
|---|---:|---:|---:|
| `narrow_integer/u8/list=4/match=0%` | 3.86 us | 2.79 us | -27.8%
(1.38x faster) |
| `narrow_integer/u8/list=4/match=50%` | 3.84 us | 2.78 us | -27.7%
(1.38x faster) |
| `narrow_integer/u8/list=16/match=0%` | 3.88 us | 3.85 us | -0.8%
(within noise) |
| `narrow_integer/u8/list=16/match=50%` | 3.84 us | 3.86 us | +0.5%
(within noise) |
| `narrow_integer/i16/list=4/match=0%` | 3.93 us | 3.18 us | -19.1%
(1.24x faster) |
| `narrow_integer/i16/list=4/match=50%` | 3.92 us | 3.16 us | -19.5%
(1.24x faster) |
| `narrow_integer/i16/list=64/match=0%` | 3.96 us | 3.82 us | -3.5%
(within noise) |
| `narrow_integer/i16/list=64/match=50%` | 3.91 us | 3.80 us | -2.9%
(within noise) |
| `narrow_integer/i16/list=256/match=0%` | 3.90 us | 3.81 us | -2.5%
(within noise) |
| `narrow_integer/i16/list=256/match=50%` | 3.97 us | 3.81 us | -4.1%
(within noise) |
| `narrow_integer/f16/list=4/match=0%` | 3.87 us | 3.16 us | -18.5%
(1.23x faster) |
| `narrow_integer/f16/list=4/match=50%` | 3.94 us | 3.15 us | -20.2%
(1.25x faster) |
| `narrow_integer/f16/list=64/match=0%` | 3.87 us | 3.84 us | -0.6%
(within noise) |
| `narrow_integer/f16/list=64/match=50%` | 3.93 us | 3.85 us | -1.9%
(within noise) |
| `narrow_integer/f16/list=256/match=0%` | 3.90 us | 3.84 us | -1.5%
(within noise) |
| `narrow_integer/f16/list=256/match=50%` | 3.87 us | 3.91 us | +1.2%
(within noise) |
| `nulls/narrow_integer/u8/list=16/match=50%/nulls=20%` | 3.92 us | 4.02
us | +2.5% (within noise) |
| `primitive/i32/small_list/list=4/match=0%` | 17.00 us | 3.04 us |
-82.1% (5.59x faster) |
| `primitive/i32/small_list/list=4/match=50%` | 32.63 us | 3.08 us |
-90.5% (10.58x faster) |
| `primitive/i32/small_list/list=32/match=0%` | 16.34 us | 13.33 us |
-18.5% (1.23x faster) |
| `primitive/i32/small_list/list=32/match=50%` | 31.17 us | 13.31 us |
-57.3% (2.34x faster) |
| `primitive/i32/small_list/list=16/match=50%/NOT_IN` | 31.98 us | 7.26
us | -77.3% (4.41x faster) |
| `nulls/primitive/i32/small_list/list=16/match=50%/nulls=20%` | 29.35
us | 7.32 us | -75.1% (4.01x faster) |
| `nulls/primitive/i32/small_list/list=16/match=50%/nulls=20%/NOT_IN` |
26.05 us | 7.42 us | -71.5% (3.51x faster) |
| `nulls/primitive/i32/small_list/list=16/match=50%/nulls=50%` | 25.89
us | 7.31 us | -71.8% (3.54x faster) |
| `primitive/i64/small_list/list=4/match=0%` | 17.12 us | 3.22 us |
-81.2% (5.31x faster) |
| `primitive/i64/small_list/list=4/match=50%` | 33.55 us | 3.18 us |
-90.5% (10.54x faster) |
| `primitive/i64/small_list/list=16/match=0%` | 16.34 us | 11.93 us |
-27.0% (1.37x faster) |
| `primitive/i64/small_list/list=16/match=50%` | 29.46 us | 11.76 us |
-60.1% (2.50x faster) |
| `f32/small_list/list=4/match=0%` | 18.14 us | 3.05 us | -83.2% (5.95x
faster) |
| `f32/small_list/list=4/match=50%` | 33.93 us | 3.04 us | -91.0%
(11.15x faster) |
| `f32/small_list/list=32/match=0%` | 22.05 us | 13.43 us | -39.1%
(1.64x faster) |
| `f32/small_list/list=32/match=50%` | 38.78 us | 13.27 us | -65.8%
(2.92x faster) |
| `timestamp_ns/small_list/list=4/match=0%` | 19.57 us | 3.18 us |
-83.8% (6.16x faster) |
| `timestamp_ns/small_list/list=4/match=50%` | 46.55 us | 3.17 us |
-93.2% (14.69x faster) |
| `timestamp_ns/small_list/list=16/match=0%` | 19.73 us | 12.07 us |
-38.8% (1.63x faster) |
| `timestamp_ns/small_list/list=16/match=50%` | 45.32 us | 11.79 us |
-74.0% (3.84x faster) |
| `interval_month_day_nano/small_list/list=4/match=0%` | 20.12 us |
13.20 us | -34.4% (1.52x faster) |
| `interval_month_day_nano/small_list/list=4/match=50%` | 52.94 us |
15.52 us | -70.7% (3.41x faster) |

</details>

---------

Co-authored-by: Andrew Lamb <andrew@nerdnetworks.org>
## Which issue does this PR close?

- Closes apache#23770
- Part of apache#15914

## Rationale for this change

Spark provides [`hypot(expr1,
expr2)`](https://spark.apache.org/docs/latest/api/sql/#hypot), which
returns `sqrt(expr1^2 + expr2^2)` computed without intermediate overflow
or underflow. It was not yet implemented in `datafusion-spark` — only an
auto-generated test stub existed at `spark/math/hypot.slt` with its
query commented out.

## What changes are included in this PR?

- Add `SparkHypot` (implementing `ScalarUDFImpl`) in
`datafusion/spark/src/function/math/hypot.rs`, backed by Rust's
`f64::hypot` — the same overflow-safe algorithm as Java/Spark's
`Math.hypot`.
- Register it in `datafusion/spark/src/function/math/mod.rs`.
- Enable the `hypot.slt` sqllogictest.

The signature is `exact(Float64, Float64) -> Float64`, following the
`datafusion-spark` convention of only accepting types Spark supports.
Computation uses the Arrow `binary` kernel so NULL in either argument
propagates to a NULL result, matching Spark.

## Are these changes tested?

Yes — `datafusion/sqllogictest/test_files/spark/math/hypot.slt` covers:
- scalar Pythagorean triples (`hypot(3, 4)` → 5, `hypot(5, 12)` → 13),
- double inputs,
- NULL propagation when either argument is NULL,
- the array path (including a NULL row),
- overflow-safety: `hypot(3e200, 4e200)` stays finite, whereas a naive
`sqrt(a^2 + b^2)` would overflow to `Infinity`.

## Are there any user-facing changes?

Yes — adds the Spark-compatible `hypot` scalar function to
`datafusion-spark`. No breaking changes to public APIs.
## Which issue does this PR close?

<!--
We generally require a GitHub issue to be filed for all bug fixes and
enhancements and this helps us generate change logs for our releases.
You can link an issue to this PR using the GitHub syntax. For example
`Closes #123` indicates that this PR will close issue #123.
-->

- Part of apache#23393 .

   ## Rationale for this change
  
`SlidingMinAccumulator::size` and `SlidingMaxAccumulator::size` only
reported
the stack size of their `ScalarValue` field plus its heap payload,
ignoring the
memory held by the underlying `MovingMin` / `MovingMax` sliding-window
buffers.
For windowed `MIN`/`MAX` over string or list data, the two per-element
stacks
can hold megabytes of `ScalarValue` payload that the memory pool was
never
  told about, understating accumulator memory usage.

  ## What changes are included in this PR?

- Add a private `heap_size(elem_heap)` method to `MovingMin<T>` and
`MovingMax<T>`
that reports the two stack buffers' capacity in bytes plus each stored
    element's heap payload.
- Factor the shared implementation into a `moving_stacks_heap_size` free
helper
    so the two types stay in sync.
  - Include the buffer bytes in `SlidingMinAccumulator::size` and
    `SlidingMaxAccumulator::size` via the new method.

  ## Are these changes tested?

Yes. Two new unit tests in
`datafusion/functions-aggregate/src/min_max.rs`:

- `moving_min_max_heap_size_i32` — fixed-width `T`, verifies buffer-only
    accounting with and without pushed elements.
- `moving_min_max_heap_size_counts_elems` — `T = String`, verifies each
of the
two slots in a `(T, T)` pair contributes independently to the heap
payload
    (mirroring the two independent `Clone`s made by `push`).
## Which issue does this PR close?

- Related to
apache#21882 (comment).

## Rationale for this change

PR apache#21882 adds custom spill file support. 

Having an example showing how a downstream application can use this new
API will help make sure the API is good enough for our needs

## What changes are included in this PR?

This PR adds a small example showing how users can back spill files with
an `ObjectStore`, using a local object store for a runnable example
while keeping the implementation applicable to remote stores.

## Are these changes tested?

Yes by CI

## Are there any user-facing changes?

Yes. This adds a new example for configuring ObjectStore-backed spill
files.
## Which issue does this PR close?

<!--
We generally require a GitHub issue to be filed for all bug fixes and
enhancements and this helps us generate change logs for our releases.
You can link an issue to this PR using the GitHub syntax. For example
`Closes #123` indicates that this PR will close issue #123.
-->

- Closes #N/A.

## Rationale for this change

<!--
Why are you proposing this change? If this is already explained clearly
in the issue then this section is not needed.
Explaining clearly why changes are proposed helps reviewers understand
your changes and offer better suggestions for fixes.
-->

Since DataFusion doesn't typically guartantee wire format compatibility,
I remove the backward compatiblity shim that introduced in PR apache#23189

FYI:
apache#23189 (comment)


## What changes are included in this PR?

<!--
There is no need to duplicate the description in the issue here but it
is sometimes worth providing a summary of the individual changes in this
PR.
-->

DataFusion does not guarantee serialized plans across versions. Keeping
`partitioned_by_file_group` therefore leaves a dead schema field and
decoder path after `output_partitioning` became the source of truth.
Reserve the old field number and name to prevent future reuse.

## Are these changes tested?

<!--
We typically require tests for all PRs in order to:
1. Prevent the code from being accidentally broken by subsequent changes
2. Serve as another way to document the expected behavior of the code

If tests are not included in your PR, please explain why (for example,
are they covered by existing tests)?
-->

Yes

## Are there any user-facing changes?

<!--
If there are user-facing changes then we may require documentation to be
updated before approving the PR.
-->

<!--
If there are any breaking changes to public APIs, please add the `api
change` label.
-->
Yes

Signed-off-by: Jiawei Zhao <Phoenix500526@163.com>
Co-authored-by: Andrew Lamb <andrew@nerdnetworks.org>
…llable build key (apache#23173)

## Which issue does this close?

- Closes apache#23126.

## Rationale for this change

`x NOT IN (subquery)` plans to a null-aware `LeftAnti` hash join (build
= outer `x`, probe = subquery). Join dynamic filter pushdown pushes a
bounds + membership filter, built from the build keys, onto the probe
scan. That filter can prune every probe row. A null-aware `LeftAnti`
reads an empty probe as a genuinely-empty subquery, so it emits
build-side NULL rows that should drop: `NULL NOT IN (non-empty)` is
UNKNOWN, not TRUE.

The result is scan-dependent, so it's a silent correctness bug. A
`VALUES` scan ignores the pushed filter and stays correct; a parquet
scan applies it and is wrong.

apache#23103 (the probe-side NULL drop) is orthogonal; this is the build-side
NULL.

## What changes are included in this PR?

Skip join dynamic filter pushdown for a null-aware anti join when the
build key can be NULL. The build-side NULL emission depends on whether
the probe is truly empty, which the pushed filter can change by emptying
it. A NOT NULL build key has no such NULL, so it keeps the pushdown.

The check is static: a schema-nullable build key disables the pushdown
even when the data contains no NULLs. A runtime alternative (keep the
pushdown and neutralize the filter only when the build actually holds a
NULL key) would restore the optimization for those cases. I'd leave that
as a follow-up.

## Are these changes tested?

Yes. A `push_down_filter_parquet.slt` case reproduces it (build-side
NULL, a non-matching parquet probe) and asserts the single correct row.
Without the change it returns the extra NULL. In addition, unit tests
pin both directions of the guard: a nullable build key rejects the
pushdown and a NOT NULL build key keeps it.

## Are there any user-facing changes?

`NOT IN` over a parquet (or otherwise prunable) scan with a nullable
outer key now returns correct results. Such joins lose the dynamic
filter pushdown.
…cimal) (apache#23631)

## Which issue does this PR close?

<!--
We generally require a GitHub issue to be filed for all bug fixes and
enhancements and this helps us generate change logs for our releases.
You can link an issue to this PR using the GitHub syntax. For example
`Closes #123` indicates that this PR will close issue #123.
-->

N/A

## Rationale for this change

<!--
Why are you proposing this change? If this is already explained clearly
in the issue then this section is not needed.
Explaining clearly why changes are proposed helps reviewers understand
your changes and offer better suggestions for fixes.
-->

I noticed there were some subtle errors with how decimals were handled
in scalar values, and also opportunity to remove power calls in favour
of precomputed constant tables. Also filling out some other missing
support.

## What changes are included in this PR?

<!--
There is no need to duplicate the description in the issue here but it
is sometimes worth providing a summary of the individual changes in this
PR.
-->

I recommend looking at the commits as they are self contained with
detailed messages for each.

## Are these changes tested?

<!--
We typically require tests for all PRs in order to:
1. Prevent the code from being accidentally broken by subsequent changes
2. Serve as another way to document the expected behavior of the code

If tests are not included in your PR, please explain why (for example,
are they covered by existing tests)?
-->

Yes

## Are there any user-facing changes?

<!--
If there are user-facing changes then we may require documentation to be
updated before approving the PR.
-->

No

<!--
If there are any breaking changes to public APIs, please add the `api
change` label.
-->
…lter pushdown enabled (apache#23638)

## Which issue does this PR close?

- Closes apache#23531.

## Rationale for this change

`input_file_name()` is `FileSource` dependent like `file_row_index()`,
and therefore shouldn't be pushed down into a filter.

## What changes are included in this PR?

`PushdownChecker` now handles both UDFs consistently.

If we keep adding this sort of metadata functions, we might want a
better API to detect them, but for now I think this is a reasonable
approach that isn't very invasive.

## Are these changes tested?

Additional SLT test that verifies that both function behave correctly
with pushdown either enabled or disabled.

## Are there any user-facing changes?

None

---------

Signed-off-by: Adam Gutglick <adamgsal@gmail.com>
…xpr::check_bigger_cast (apache#23808) (apache#23809)

## Which issue does this PR close?

- Closes apache#23808.

## Rationale for this change
 ref. apache#23807

`CastExpr::check_bigger_cast` is used to determine whether a cast is a
widening (order-preserving) conversion. Currently, it classifies `Int32
-> Float32`, `UInt32 -> Float32`, `Int64 -> Float64`, and `UInt64 ->
Float64` as widening casts.

However, integer-to-float conversions for 32-bit and 64-bit integers
lose precision when values exceed the mantissa bit limit (24 bits for
`Float32`, 53 bits for `Float64`). For example:
`16_777_216_i32 as f32 == 16_777_217_i32 as f32` (both yield
16777216.0f32).

Because distinct integer inputs can collapse to the same float output,
these casts are not strictly 1-to-1 (injective) and can break suffix
sort key ordering in multi-column ordering analysis (e.g. `[CAST(a AS
Float32), b]`).

## What changes are included in this PR?

- Updated `CastExpr::check_bigger_cast` to exclude precision-losing
integer-to-float conversions (`Int32/UInt32 -> Float32` and
`Int64/UInt64 -> Float64`).
- Added unit test `test_check_bigger_cast_precision_loss` to verify
precision-losing casts return `false` while exact conversions (`Int16 ->
Float32`, `Int32 -> Float64`, etc.) continue to return `true`.

## Are these changes tested?

Yes, new unit test `test_check_bigger_cast_precision_loss` in `cast.rs`.

## Are there any user-facing changes?

No breaking API changes. Internal optimizer behavior fix.
…ossible (apache#23702)

## Which issue does this PR close?

N/A

## Rationale for this change

`SortPreservingMergeStream` is a little complex, so add some guiding
comments and make it as textbook-like as possible

## What changes are included in this PR?

Added comments, reorder code

While this was done this also fixed couple of bugs due to how it work:
1. leftover drain was not counted in the `elapsed_compute`
2. `limit(0)` returns 0 rows and not 1 

## Are these changes tested?
Existing tests

## Are there any user-facing changes?
Not API ones.

`limit(0)` now returns 0 rows
## Which issue does this PR close?

- Closes apache#21428.

## Rationale for this change

This PR adds `BuildHasher`-based variants for `hash_utils` so callers
can compute row hashes with a caller-provided hash builder instead of
always using DataFusion's default `RandomState`.

The main constraint is performance: `with_hashes` is a hot path,
especially for string, dictionary, and nested array hashing. A previous
version in apache#21429 caused measurable regressions in the default
`RandomState` path, for example `large_utf8: single, no nulls` regressed
from roughly `26.7us` to `36.3us`, and `large_utf8: multiple, no nulls`
from roughly `112us` to `127us`.

This version keeps the default path performance-oriented by avoiding a
fully generic `BuildHasher` rewrite of the existing hot loops.

## What changes are included in this PR?

This PR adds:

- `with_hashes_with_hasher`
- `create_hashes_with_hasher`
- custom-hasher implementations for primitive, string, binary,
byte-view, dictionary, and nested arrays
- tests covering custom hashers, multi-column hashing, and dictionary
equivalence

The implementation intentionally uses a hybrid design:

- Default `RandomState` leaf hot paths remain specialized.
- Custom `BuildHasher` leaf paths live separately in
`hash_utils/build_hasher.rs`.
- Nested/structural logic is shared through an internal child-hashing
adapter, so struct/list/map/union/run/dictionary behavior does not need
to be broadly duplicated.

The trade-off is that there is still some duplication for
primitive/string/binary leaf loops. That duplication is intentional:
those are the hottest loops, and keeping them separate prevents the
existing `RandomState` path from becoming generic over `BuildHasher` or
being perturbed by the custom-hasher implementation.

## Are these changes tested?

Yes.

---------

Co-authored-by: Dmitrii Blaginin <dmitrii@blaginin.me>
Co-authored-by: Andrew Lamb <andrew@nerdnetworks.org>
…t whether they keep the same ordering of the input (apache#23807)

## Which issue does this PR close?

- Closes apache#23798

Related to:
- apache#16217

## Rationale for this change

To be able to keep the same sorting order allowing for more
optimizations

Now comet own `cast` implementation can be recognized as not modifying
sort order in the same cases that datafusion cast does.

## What changes are included in this PR?
added `strictly_order_preserving` property to `ExprProperties` + varius
other places
and replaced the hard coded logic for cast (`substitute_cast_ordering`)
about keeping input order with more generic approach that now any
expression can implement and have the same advantage of sort elimination

also marked from_unixtime as keeping ordering to show a case of this
optimization

## Are these changes tested?
yes

## Are there any user-facing changes?
yes, breaking change, added `strictly_order_preserving` property to
`ExprProperties` and to `FFI_ExprProperties`


this property means that given expression `f` and 2 values from the
input column `a` and `b` the following variants are kept:
1. `a.cmp(b) == f(a).cmp(f(b))`
2. nulls maps to nulls

Example of satisfying expression: 
`cast(col_a as BIGINT)` where `col_a` is `INT` it is keeping the
properties

Example of not satisfying:

`floor` - floor can not

`array_repeat(my_col, 2)` which might look like at first glance as
keeping the property as well but in fact it does not.

the reason is that `array_repeat(null, 2)` will output list of 2 nulls
which breaks the 2nd property that nulls must be kept as nulls


how to migrate:

Option 1 - keeping the old behavior (safest but least performant)
set `strictly_order_preserving` to false

Option 2 - Using the new optimization that this opens up:
set `strictly_order_preserving` only if the expression keep both
variants
## Which issue does this PR close?

Part of apache#22330.

## Rationale for this change

`ForeignScalarUDF` inherits the default `preserves_lex_ordering`, so
producer overrides are lost across the FFI boundary.

## What changes are included in this PR?

- Forward `preserves_lex_ordering` through `FFI_ScalarUDF`.
- Reuse the existing placement UDF for unit and dynamic-library
coverage.

## Are these changes tested?

- `cargo test -p datafusion-ffi --features integration-tests`
- `cargo clippy --all-targets --all-features -- -D warnings`

## Are there any user-facing changes?

The FFI ABI changes. Foreign libraries must rebuild against the new
DataFusion version.

---------

Signed-off-by: Amogh Ramesh <ramogh2404@gmail.com>
…tonic deques (apache#23826) (apache#23827)

## Which issue does this PR close?

- Closes apache#23826

### What changes are included in this PR?

This PR optimizes sliding window `MIN`/`MAX` aggregate functions using a
**Sequence-Numbered Monotonic Deque** instead of a Two-Stack Queue.

**Revised Design:**
- We store `(sequence_number, value)` pairs in a single `VecDeque`, kept
in strictly monotonic order.
- **`push(val)`**: Evicts dominated elements from the back, then pushes
the new value with the current `push_seq` and increments `push_seq`.
- **`pop()`**: Increments `pop_seq`. If the front elements sequence
number equals the old `pop_seq`, it is expired and popped.
- **Benefits**: No secondary FIFO queue (lower memory) and no `clone()`
overhead.

### Are these changes tested?
Yes, existing tests pass. Added tests for empty-window `pop()` and
duplicate-heavy scenarios.

### Are there any user-facing changes?
**API Change**: `MovingMin` and `MovingMax` were changed to `pub(crate)`
visibility, and their `pop()` methods now return `()` instead of
returning a value.
Performance is significantly improved (2x-3.5x throughput).

---------

Co-authored-by: Pavan <pavan@Lakshmis-MacBook-Air.local>
## Which issue does this PR close?

NA

## Rationale for this change

Adds [datapress](https://docs.datap-rs.org) to the list of known users
in the documentation.

## What changes are included in this PR?

NA

## Are these changes tested?

NA

## Are there any user-facing changes?

NA
## Which issue does this PR close?

- Closes apache#11748

## Rationale for this change

`EliminateGroupByConstant` removes GROUP BY expressions that are
constants. If all of the GROUP BY expressions are constants, eliminating
all of them results in converting a grouped aggregate into an ungrouped
(global) aggregate. This changes the semantics of the query: a grouped
aggregate query on an empty input returns zero rows, whereas an
ungrouped aggregate query produces a single row.

## What changes are included in this PR?

`EliminateGroupByConstant` now declines to eliminate constant GROUP BY
expressions, if doing so would result in removing all of the grouping
expressions.

## Are these changes tested?

Yes. Sqllogictest reproducer for the original issue, updated unit test
and optimizer SLT expectations.

## Are there any user-facing changes?

Queries with all-constant GROUP BY may now return fewer (correct) rows.
…pache#23874)

## Which issue does this PR close?

- Closes apache#23872 

## Rationale for this change

`SlidingMinAccumulator::update_batch` skipped NULL values. This meant
that if all the non-NULL values in a window frame were retracted, the
window frame would not be empty but the `MovingMin` data structure by
the `SlidingMinAccumulator` would not contain any values. This resulted
in incorrectly returning a stale non-NULL value for a sliding `min()`
over a window frame consisting of only NULL values.

Along the way, optimize the min and max sliding window accumulators to
make them both more efficient and more symmetric with one another. In
the original coding, `SlidingMaxAccumulator` included NULL values but
`SlidingMinAccumulator` omitted them, in part because omitting NULLs
made the original `retract_batch` implementation more expensive. This PR
optimizes `retract_batch`, so we can now use the same scheme for both
the min and max sliding accumulators:

* Omit NULLs on `update_batch` (this improves on the prior behavior of
`max`)
* Efficiently account for NULLs in `retract_batch` (this improves on the
prior behavior of `min`)
* Ensure correct results for all-NULL window frames (this fixes the
prior bug in `min`).

## What changes are included in this PR?

* Fix bug in `min()` over all-NULL window frames
* Optimize `SlidingMinAccumulator::retract_batch` (avoid materializing
values just to count NULLs)
* Optimize `SlidingMinAccumulator::update_batch` (omit NULLs), also
improving symmetry with `min`
* Optimize both accumulators to stop caching the current `min` / `max`;
this saves a few clones, but perhaps more importantly it is simpler and
avoids the risk of inconsistency between the cached value and the
underlying `MovingMin` / `MovingMax` data structure

## Are these changes tested?

Yes, new tests added.

## Are there any user-facing changes?

No, aside from the bug fix.
## Which issue does this PR close?

- N/A

## Rationale for this change

```
$ cargo test -p datafusion-functions-aggregate
[...]
  warning: associated function `from_parts` is never used
     --> datafusion/physical-expr/src/expressions/dynamic_filters/mod.rs:464:8
[...]
```

Fix this by appropriately gating the compilation of `from_parts`.

## What changes are included in this PR?

* Squelch unused code warning

## Are these changes tested?

Yes.

## Are there any user-facing changes?

No.
…w retract (apache#23913)

## Which issue does this PR close?

- Closes apache#23912.

## Rationale for this change

`DistinctPercentileContAccumulator` reused the shared set-based
`GenericDistinctBuffer`, which doesn't fit it: (1) the buffer asserts a
single input column but `percentile_cont` passes two (value +
percentile), so every `percentile_cont(DISTINCT ...)` panicked; (2) the
buffer is a plain `HashSet` with no multiplicity, so sliding-window
`retract_batch` dropped a value while duplicates were still in the
frame.

## What changes are included in this PR?

- Replace the shared buffer in this accumulator with a per-accumulator
`HashMap<Hashable, usize>` count map: `update_batch` reads only the
value column and increments; `retract_batch` decrements and removes a
key only at zero; `state`/`merge_batch` keep the same List state shape.
Other `GenericDistinctBuffer` users are untouched.
- Regression tests in `aggregate.slt` for the plain distinct query and
the sliding-window duplicate-retract case.

## Are these changes tested?

Yes — new regression tests; the full `aggregate.slt` suite passes.

## Are there any user-facing changes?

`percentile_cont(DISTINCT ...)` now works instead of panicking, and
returns correct results in sliding windows.
## Which issue does this PR close?

- Closes apache#23717.

## Rationale for this change

Spark and Java format negative numeric values with parentheses when the
`(` flag is present. The decimal formatting path always emitted a minus
sign because it ignored `negative_in_parentheses`, while the
floating-point path already handled the flag. This made `format_string`
inconsistent across numeric input types.

## What changes are included in this PR?

- Add a closing suffix when a negative decimal uses parentheses
formatting.
- Include the suffix when calculating width, left alignment, and zero
padding.
- Add regression coverage for grouped negative decimals with and without
an explicit width.

## Are these changes tested?

Yes. The following checks pass locally:

- `cargo fmt --all -- --check`
- `cargo clippy --all-targets --all-features -- -D warnings`
- `cargo test -p datafusion-spark --lib`
- The full workspace test command from the contributor guide with
`avro,json,backtrace,extended_tests,recursive_protection,parquet_encryption`
enabled

## Are there any user-facing changes?

Yes. Spark-compatible formatting of negative decimal values now honors
the parentheses flag. There are no public API or breaking changes.
## Which issue does this PR close?

<!--
We generally require a GitHub issue to be filed for all bug fixes and
enhancements and this helps us generate change logs for our releases.
You can link an issue to this PR using the GitHub syntax. For example
`Closes #123` indicates that this PR will close issue #123.
-->

- Closes #.

## Rationale for this change

<!--
Why are you proposing this change? If this is already explained clearly
in the issue then this section is not needed.
Explaining clearly why changes are proposed helps reviewers understand
your changes and offer better suggestions for fixes.
-->

Precedes apache#23720.

While displaying a plan with metrics, allow filtering them by name, not
just by category or type.

This is very useful when writing snapshot tests where only a specific
metrics needs to be asserted, without other unrelated metrics polluting
the snapshot assertion.

Regardless of what happens with
apache#23720, I think this PR is
still worth it, as it allows creating some really nice `insta` tests
asserting runtime properties reliably by just cherry picking the runtime
metrics relevant for that specific test. This is relevant not only
within DataFusion codebase, but also for other people's codebases using
DataFusion and relying on `insta` for snapshot testing, See an example
of this here:
https://github.com/apache/datafusion/pull/23720/changes#diff-8281405e117428077c19c07afbbe59fed303a3460096e958f761d34f0899b618

## What changes are included in this PR?

<!--
There is no need to duplicate the description in the issue here but it
is sometimes worth providing a summary of the individual changes in this
PR.
-->

While displaying metrics, allows filtering them by name

## Are these changes tested?

<!--
We typically require tests for all PRs in order to:
1. Prevent the code from being accidentally broken by subsequent changes
2. Serve as another way to document the expected behavior of the code

If tests are not included in your PR, please explain why (for example,
are they covered by existing tests)?
-->

Yes, by a new small test.

## Are there any user-facing changes?

<!--
If there are user-facing changes then we may require documentation to be
updated before approving the PR.
-->

<!--
If there are any breaking changes to public APIs, please add the `api
change` label.
-->

People using the metrics display API will be able to filter metrics by
name.
## Which issue does this PR close?

N/A

## Rationale for this change

I saw that `empty` implementation is inefficient.
I wrote review comments for more why inefficient.

there are some optimizations that do not require benchmark since they
are obvious once you understand, this is one of them

## What changes are included in this PR?
rewrote the function to be fast

## Are these changes tested?
existing tests

## Are there any user-facing changes?
no
Bumps the codeql-actions group with 2 updates:
[github/codeql-action/init](https://github.com/github/codeql-action) and
[github/codeql-action/analyze](https://github.com/github/codeql-action).

Updates `github/codeql-action/init` from 4.37.1 to 4.37.3
<details>
<summary>Release notes</summary>
<p><em>Sourced from <a
href="https://github.com/github/codeql-action/releases">github/codeql-action/init's
releases</a>.</em></p>
<blockquote>
<h2>v4.37.3</h2>
<p>No user facing changes.</p>
<h2>v4.37.2</h2>
<ul>
<li>The new address format for the <code>config-file</code> input that
was introduced in CodeQL Action 4.37.0 is now enabled by default. In
addition to the format described there, the <code>remote=</code> prefix
can now be used to explicitly indicate that the input refers to a remote
file. All previous input formats continue to be accepted as well. <a
href="https://redirect.github.com/github/codeql-action/pull/4023">#4023</a></li>
<li>The CodeQL Action can now make use of <a
href="https://docs.github.com/en/code-security/how-tos/secure-at-scale/configure-organization-security/manage-usage-and-access/giving-org-access-private-registries">configured
private registries</a> in Default Setup to retrieve CodeQL configuration
files from remote repositories that require authentication. This will
allow customers to store their CodeQL configuration in a single
repository that can then be referenced by Default Setup workflows in
other repositories. We expect to roll this and other, related changes
out to everyone in July. <a
href="https://redirect.github.com/github/codeql-action/pull/4007">#4007</a></li>
</ul>
</blockquote>
</details>
<details>
<summary>Changelog</summary>
<p><em>Sourced from <a
href="https://github.com/github/codeql-action/blob/main/CHANGELOG.md">github/codeql-action/init's
changelog</a>.</em></p>
<blockquote>
<h1>CodeQL Action Changelog</h1>
<p>See the <a
href="https://github.com/github/codeql-action/releases">releases
page</a> for the relevant changes to the CodeQL CLI and language
packs.</p>
<h2>[UNRELEASED]</h2>
<ul>
<li>This version of the CodeQL Action adds support for the
<code>tools</code> input for the <code>codeql-action/init</code> step to
be specified using a <code>github-codeql-tools</code> <a
href="https://docs.github.com/en/organizations/managing-organization-settings/managing-custom-properties-for-repositories-in-your-organization">repository
property</a>. This feature will gradually be rolled out following the
release of this version. Once rolled out, this allows for the CodeQL CLI
version that is used in GitHub-managed workflows, such as Default Setup,
to be set to a custom value. For example, customers who run into issues
with rate limits when a new CodeQL CLI version is released can set the
value to <code>toolcache</code> to always use the CodeQL CLI version
that is available in the runner toolcache. For Advanced Setup workflows,
the value provided for <code>tools</code> in the workflow definition
always takes precedence unless the value of the repository property
starts with <code>!</code>. <a
href="https://redirect.github.com/github/codeql-action/pull/4037">#4037</a></li>
</ul>
<h2>4.37.3 - 22 Jul 2026</h2>
<p>No user facing changes.</p>
<h2>4.37.2 - 21 Jul 2026</h2>
<ul>
<li>The new address format for the <code>config-file</code> input that
was introduced in CodeQL Action 4.37.0 is now enabled by default. In
addition to the format described there, the <code>remote=</code> prefix
can now be used to explicitly indicate that the input refers to a remote
file. All previous input formats continue to be accepted as well. <a
href="https://redirect.github.com/github/codeql-action/pull/4023">#4023</a></li>
<li>The CodeQL Action can now make use of <a
href="https://docs.github.com/en/code-security/how-tos/secure-at-scale/configure-organization-security/manage-usage-and-access/giving-org-access-private-registries">configured
private registries</a> in Default Setup to retrieve CodeQL configuration
files from remote repositories that require authentication. This will
allow customers to store their CodeQL configuration in a single
repository that can then be referenced by Default Setup workflows in
other repositories. We expect to roll this and other, related changes
out to everyone in July. <a
href="https://redirect.github.com/github/codeql-action/pull/4007">#4007</a></li>
</ul>
<h2>4.37.1 - 16 Jul 2026</h2>
<ul>
<li><em>Upcoming breaking change</em>: Add a deprecation warning for
customers using CodeQL version 2.20.6 and earlier. These versions of
CodeQL were discontinued on 1 July 2026 alongside GitHub Enterprise
Server 3.16, and will be unsupported by the next minor release of the
CodeQL Action. <a
href="https://redirect.github.com/github/codeql-action/pull/3956">#3956</a></li>
<li>Update default CodeQL bundle version to <a
href="https://github.com/github/codeql-action/releases/tag/codeql-bundle-v2.26.1">2.26.1</a>.
<a
href="https://redirect.github.com/github/codeql-action/pull/4019">#4019</a></li>
</ul>
<h2>4.37.0 - 08 Jul 2026</h2>
<ul>
<li>Update default CodeQL bundle version to <a
href="https://github.com/github/codeql-action/releases/tag/codeql-bundle-v2.26.0">2.26.0</a>.
<a
href="https://redirect.github.com/github/codeql-action/pull/3995">#3995</a></li>
<li>In addition to the existing input format, the
<code>config-file</code> input for the <code>codeql-action/init</code>
step will soon support a new <code>[owner/]repo[@ref][:path]</code>
format. All components except the repository name are optional. If
omitted, <code>owner</code> defaults to the same owner as the repository
the analysis is running for, <code>ref</code> to <code>main</code>, and
<code>path</code> to <code>.github/codeql-action.yaml</code>. Support
for this format ships in this version of the CodeQL Action, but will
only be enabled over the coming weeks. <a
href="https://redirect.github.com/github/codeql-action/pull/3973">#3973</a></li>
</ul>
<h2>4.36.3 - 01 Jul 2026</h2>
<p>No user facing changes.</p>
<h2>4.36.2 - 04 Jun 2026</h2>
<ul>
<li>Cache CodeQL CLI version information across Actions steps. <a
href="https://redirect.github.com/github/codeql-action/pull/3943">#3943</a></li>
<li>Reduce requests while waiting for analysis processing by using
exponential backoff when polling SARIF processing status. <a
href="https://redirect.github.com/github/codeql-action/pull/3937">#3937</a></li>
<li>Update default CodeQL bundle version to <a
href="https://github.com/github/codeql-action/releases/tag/codeql-bundle-v2.25.6">2.25.6</a>.
<a
href="https://redirect.github.com/github/codeql-action/pull/3948">#3948</a></li>
</ul>
<h2>4.36.1 - 02 Jun 2026</h2>
<p>No user facing changes.</p>
<h2>4.36.0 - 22 May 2026</h2>
<ul>
<li><em>Breaking change</em>: Bump the minimum required CodeQL bundle
version to 2.19.4. <a
href="https://redirect.github.com/github/codeql-action/pull/3894">#3894</a></li>
<li>Add support for SHA-256 Git object IDs. <a
href="https://redirect.github.com/github/codeql-action/pull/3893">#3893</a></li>
<li>Update default CodeQL bundle version to <a
href="https://github.com/github/codeql-action/releases/tag/codeql-bundle-v2.25.5">2.25.5</a>.
<a
href="https://redirect.github.com/github/codeql-action/pull/3926">#3926</a></li>
</ul>
<h2>4.35.5 - 15 May 2026</h2>
<ul>
<li>We have improved how the JavaScript bundles for the CodeQL Action
are generated to avoid duplication across bundles and reduce the size of
the repository by around 70%. This should have no effect on the runtime
behaviour of the CodeQL Action. <a
href="https://redirect.github.com/github/codeql-action/pull/3899">#3899</a></li>
</ul>
<!-- raw HTML omitted -->
</blockquote>
<p>... (truncated)</p>
</details>
<details>
<summary>Commits</summary>
<ul>
<li><a
href="https://github.com/github/codeql-action/commit/e4fba868fa4b1b91e1fdab776edc8cfbe6e9fb81"><code>e4fba86</code></a>
Merge pull request <a
href="https://redirect.github.com/github/codeql-action/issues/4031">#4031</a>
from github/update-v4.37.3-72f6a9da0</li>
<li><a
href="https://github.com/github/codeql-action/commit/fb50ab5d62a274adf3ef3e22cfe750ae87a0ede7"><code>fb50ab5</code></a>
Update changelog for v4.37.3</li>
<li><a
href="https://github.com/github/codeql-action/commit/72f6a9da0def52d9193d6a758f0378b65091f8d1"><code>72f6a9d</code></a>
Merge pull request <a
href="https://redirect.github.com/github/codeql-action/issues/4030">#4030</a>
from github/mbg/fix/no-proxy</li>
<li><a
href="https://github.com/github/codeql-action/commit/3b5ee58597653d9cc6785f3f1277f796d81f3646"><code>3b5ee58</code></a>
Use default <code>request</code> options instead of
<code>undefined</code></li>
<li><a
href="https://github.com/github/codeql-action/commit/bfb6be4b5ecd3650f02f530571453e8c64ef0778"><code>bfb6be4</code></a>
Merge pull request <a
href="https://redirect.github.com/github/codeql-action/issues/4028">#4028</a>
from github/mergeback/v4.37.2-to-main-e0647621</li>
<li><a
href="https://github.com/github/codeql-action/commit/526ab84f9858816d9cf5f7b9df4dd5e2235f0eba"><code>526ab84</code></a>
Rebuild</li>
<li><a
href="https://github.com/github/codeql-action/commit/d6217b9b8c14166e4851db94c11155d03bd13c07"><code>d6217b9</code></a>
Update changelog and version after v4.37.2</li>
<li><a
href="https://github.com/github/codeql-action/commit/e0647621c2984b5ed2f768cb892365bf2a616ad1"><code>e064762</code></a>
Merge pull request <a
href="https://redirect.github.com/github/codeql-action/issues/4027">#4027</a>
from github/update-v4.37.2-385bcdc5a</li>
<li><a
href="https://github.com/github/codeql-action/commit/e0faed839190caa67a5cd42f1cc16246028ca3df"><code>e0faed8</code></a>
Add a couple of change notes</li>
<li><a
href="https://github.com/github/codeql-action/commit/73aad0eaa9df172668665a150d17b8bc5a650c20"><code>73aad0e</code></a>
Update changelog for v4.37.2</li>
<li>Additional commits viewable in <a
href="https://github.com/github/codeql-action/compare/7188fc363630916deb702c7fdcf4e481b751f97a...e4fba868fa4b1b91e1fdab776edc8cfbe6e9fb81">compare
view</a></li>
</ul>
</details>
<br />

Updates `github/codeql-action/analyze` from 4.37.1 to 4.37.3
<details>
<summary>Release notes</summary>
<p><em>Sourced from <a
href="https://github.com/github/codeql-action/releases">github/codeql-action/analyze's
releases</a>.</em></p>
<blockquote>
<h2>v4.37.3</h2>
<p>No user facing changes.</p>
<h2>v4.37.2</h2>
<ul>
<li>The new address format for the <code>config-file</code> input that
was introduced in CodeQL Action 4.37.0 is now enabled by default. In
addition to the format described there, the <code>remote=</code> prefix
can now be used to explicitly indicate that the input refers to a remote
file. All previous input formats continue to be accepted as well. <a
href="https://redirect.github.com/github/codeql-action/pull/4023">#4023</a></li>
<li>The CodeQL Action can now make use of <a
href="https://docs.github.com/en/code-security/how-tos/secure-at-scale/configure-organization-security/manage-usage-and-access/giving-org-access-private-registries">configured
private registries</a> in Default Setup to retrieve CodeQL configuration
files from remote repositories that require authentication. This will
allow customers to store their CodeQL configuration in a single
repository that can then be referenced by Default Setup workflows in
other repositories. We expect to roll this and other, related changes
out to everyone in July. <a
href="https://redirect.github.com/github/codeql-action/pull/4007">#4007</a></li>
</ul>
</blockquote>
</details>
<details>
<summary>Changelog</summary>
<p><em>Sourced from <a
href="https://github.com/github/codeql-action/blob/main/CHANGELOG.md">github/codeql-action/analyze's
changelog</a>.</em></p>
<blockquote>
<h1>CodeQL Action Changelog</h1>
<p>See the <a
href="https://github.com/github/codeql-action/releases">releases
page</a> for the relevant changes to the CodeQL CLI and language
packs.</p>
<h2>[UNRELEASED]</h2>
<ul>
<li>This version of the CodeQL Action adds support for the
<code>tools</code> input for the <code>codeql-action/init</code> step to
be specified using a <code>github-codeql-tools</code> <a
href="https://docs.github.com/en/organizations/managing-organization-settings/managing-custom-properties-for-repositories-in-your-organization">repository
property</a>. This feature will gradually be rolled out following the
release of this version. Once rolled out, this allows for the CodeQL CLI
version that is used in GitHub-managed workflows, such as Default Setup,
to be set to a custom value. For example, customers who run into issues
with rate limits when a new CodeQL CLI version is released can set the
value to <code>toolcache</code> to always use the CodeQL CLI version
that is available in the runner toolcache. For Advanced Setup workflows,
the value provided for <code>tools</code> in the workflow definition
always takes precedence unless the value of the repository property
starts with <code>!</code>. <a
href="https://redirect.github.com/github/codeql-action/pull/4037">#4037</a></li>
</ul>
<h2>4.37.3 - 22 Jul 2026</h2>
<p>No user facing changes.</p>
<h2>4.37.2 - 21 Jul 2026</h2>
<ul>
<li>The new address format for the <code>config-file</code> input that
was introduced in CodeQL Action 4.37.0 is now enabled by default. In
addition to the format described there, the <code>remote=</code> prefix
can now be used to explicitly indicate that the input refers to a remote
file. All previous input formats continue to be accepted as well. <a
href="https://redirect.github.com/github/codeql-action/pull/4023">#4023</a></li>
<li>The CodeQL Action can now make use of <a
href="https://docs.github.com/en/code-security/how-tos/secure-at-scale/configure-organization-security/manage-usage-and-access/giving-org-access-private-registries">configured
private registries</a> in Default Setup to retrieve CodeQL configuration
files from remote repositories that require authentication. This will
allow customers to store their CodeQL configuration in a single
repository that can then be referenced by Default Setup workflows in
other repositories. We expect to roll this and other, related changes
out to everyone in July. <a
href="https://redirect.github.com/github/codeql-action/pull/4007">#4007</a></li>
</ul>
<h2>4.37.1 - 16 Jul 2026</h2>
<ul>
<li><em>Upcoming breaking change</em>: Add a deprecation warning for
customers using CodeQL version 2.20.6 and earlier. These versions of
CodeQL were discontinued on 1 July 2026 alongside GitHub Enterprise
Server 3.16, and will be unsupported by the next minor release of the
CodeQL Action. <a
href="https://redirect.github.com/github/codeql-action/pull/3956">#3956</a></li>
<li>Update default CodeQL bundle version to <a
href="https://github.com/github/codeql-action/releases/tag/codeql-bundle-v2.26.1">2.26.1</a>.
<a
href="https://redirect.github.com/github/codeql-action/pull/4019">#4019</a></li>
</ul>
<h2>4.37.0 - 08 Jul 2026</h2>
<ul>
<li>Update default CodeQL bundle version to <a
href="https://github.com/github/codeql-action/releases/tag/codeql-bundle-v2.26.0">2.26.0</a>.
<a
href="https://redirect.github.com/github/codeql-action/pull/3995">#3995</a></li>
<li>In addition to the existing input format, the
<code>config-file</code> input for the <code>codeql-action/init</code>
step will soon support a new <code>[owner/]repo[@ref][:path]</code>
format. All components except the repository name are optional. If
omitted, <code>owner</code> defaults to the same owner as the repository
the analysis is running for, <code>ref</code> to <code>main</code>, and
<code>path</code> to <code>.github/codeql-action.yaml</code>. Support
for this format ships in this version of the CodeQL Action, but will
only be enabled over the coming weeks. <a
href="https://redirect.github.com/github/codeql-action/pull/3973">#3973</a></li>
</ul>
<h2>4.36.3 - 01 Jul 2026</h2>
<p>No user facing changes.</p>
<h2>4.36.2 - 04 Jun 2026</h2>
<ul>
<li>Cache CodeQL CLI version information across Actions steps. <a
href="https://redirect.github.com/github/codeql-action/pull/3943">#3943</a></li>
<li>Reduce requests while waiting for analysis processing by using
exponential backoff when polling SARIF processing status. <a
href="https://redirect.github.com/github/codeql-action/pull/3937">#3937</a></li>
<li>Update default CodeQL bundle version to <a
href="https://github.com/github/codeql-action/releases/tag/codeql-bundle-v2.25.6">2.25.6</a>.
<a
href="https://redirect.github.com/github/codeql-action/pull/3948">#3948</a></li>
</ul>
<h2>4.36.1 - 02 Jun 2026</h2>
<p>No user facing changes.</p>
<h2>4.36.0 - 22 May 2026</h2>
<ul>
<li><em>Breaking change</em>: Bump the minimum required CodeQL bundle
version to 2.19.4. <a
href="https://redirect.github.com/github/codeql-action/pull/3894">#3894</a></li>
<li>Add support for SHA-256 Git object IDs. <a
href="https://redirect.github.com/github/codeql-action/pull/3893">#3893</a></li>
<li>Update default CodeQL bundle version to <a
href="https://github.com/github/codeql-action/releases/tag/codeql-bundle-v2.25.5">2.25.5</a>.
<a
href="https://redirect.github.com/github/codeql-action/pull/3926">#3926</a></li>
</ul>
<h2>4.35.5 - 15 May 2026</h2>
<ul>
<li>We have improved how the JavaScript bundles for the CodeQL Action
are generated to avoid duplication across bundles and reduce the size of
the repository by around 70%. This should have no effect on the runtime
behaviour of the CodeQL Action. <a
href="https://redirect.github.com/github/codeql-action/pull/3899">#3899</a></li>
</ul>
<!-- raw HTML omitted -->
</blockquote>
<p>... (truncated)</p>
</details>
<details>
<summary>Commits</summary>
<ul>
<li><a
href="https://github.com/github/codeql-action/commit/e4fba868fa4b1b91e1fdab776edc8cfbe6e9fb81"><code>e4fba86</code></a>
Merge pull request <a
href="https://redirect.github.com/github/codeql-action/issues/4031">#4031</a>
from github/update-v4.37.3-72f6a9da0</li>
<li><a
href="https://github.com/github/codeql-action/commit/fb50ab5d62a274adf3ef3e22cfe750ae87a0ede7"><code>fb50ab5</code></a>
Update changelog for v4.37.3</li>
<li><a
href="https://github.com/github/codeql-action/commit/72f6a9da0def52d9193d6a758f0378b65091f8d1"><code>72f6a9d</code></a>
Merge pull request <a
href="https://redirect.github.com/github/codeql-action/issues/4030">#4030</a>
from github/mbg/fix/no-proxy</li>
<li><a
href="https://github.com/github/codeql-action/commit/3b5ee58597653d9cc6785f3f1277f796d81f3646"><code>3b5ee58</code></a>
Use default <code>request</code> options instead of
<code>undefined</code></li>
<li><a
href="https://github.com/github/codeql-action/commit/bfb6be4b5ecd3650f02f530571453e8c64ef0778"><code>bfb6be4</code></a>
Merge pull request <a
href="https://redirect.github.com/github/codeql-action/issues/4028">#4028</a>
from github/mergeback/v4.37.2-to-main-e0647621</li>
<li><a
href="https://github.com/github/codeql-action/commit/526ab84f9858816d9cf5f7b9df4dd5e2235f0eba"><code>526ab84</code></a>
Rebuild</li>
<li><a
href="https://github.com/github/codeql-action/commit/d6217b9b8c14166e4851db94c11155d03bd13c07"><code>d6217b9</code></a>
Update changelog and version after v4.37.2</li>
<li><a
href="https://github.com/github/codeql-action/commit/e0647621c2984b5ed2f768cb892365bf2a616ad1"><code>e064762</code></a>
Merge pull request <a
href="https://redirect.github.com/github/codeql-action/issues/4027">#4027</a>
from github/update-v4.37.2-385bcdc5a</li>
<li><a
href="https://github.com/github/codeql-action/commit/e0faed839190caa67a5cd42f1cc16246028ca3df"><code>e0faed8</code></a>
Add a couple of change notes</li>
<li><a
href="https://github.com/github/codeql-action/commit/73aad0eaa9df172668665a150d17b8bc5a650c20"><code>73aad0e</code></a>
Update changelog for v4.37.2</li>
<li>Additional commits viewable in <a
href="https://github.com/github/codeql-action/compare/7188fc363630916deb702c7fdcf4e481b751f97a...e4fba868fa4b1b91e1fdab776edc8cfbe6e9fb81">compare
view</a></li>
</ul>
</details>
<br />


Dependabot will resolve any conflicts with this PR as long as you don't
alter it yourself. You can also trigger a rebase manually by commenting
`@dependabot rebase`.

[//]: # (dependabot-automerge-start)
[//]: # (dependabot-automerge-end)

---

<details>
<summary>Dependabot commands and options</summary>
<br />

You can trigger Dependabot actions by commenting on this PR:
- `@dependabot rebase` will rebase this PR
- `@dependabot recreate` will recreate this PR, overwriting any edits
that have been made to it
- `@dependabot show <dependency name> ignore conditions` will show all
of the ignore conditions of the specified dependency
- `@dependabot ignore <dependency name> major version` will close this
group update PR and stop Dependabot creating any more for the specific
dependency's major version (unless you unignore this specific
dependency's major version or upgrade to it yourself)
- `@dependabot ignore <dependency name> minor version` will close this
group update PR and stop Dependabot creating any more for the specific
dependency's minor version (unless you unignore this specific
dependency's minor version or upgrade to it yourself)
- `@dependabot ignore <dependency name>` will close this group update PR
and stop Dependabot creating any more for the specific dependency
(unless you unignore this specific dependency or upgrade to it yourself)
- `@dependabot unignore <dependency name>` will remove all of the ignore
conditions of the specified dependency
- `@dependabot unignore <dependency name> <ignore condition>` will
remove the ignore condition of the specified dependency and ignore
conditions


</details>

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
…e#23941)

Bumps
[taiki-e/install-action](https://github.com/taiki-e/install-action) from
2.84.0 to 2.85.2.
<details>
<summary>Release notes</summary>
<p><em>Sourced from <a
href="https://github.com/taiki-e/install-action/releases">taiki-e/install-action's
releases</a>.</em></p>
<blockquote>
<h2>2.85.2</h2>
<ul>
<li>
<p>Update <code>prek@latest</code> to 0.4.11.</p>
</li>
<li>
<p>Update <code>mise@latest</code> to 2026.7.13.</p>
</li>
<li>
<p>Update <code>kingfisher@latest</code> to 1.109.0.</p>
</li>
</ul>
<h2>2.85.1</h2>
<ul>
<li>
<p>Update <code>vacuum@latest</code> to 0.30.0.</p>
</li>
<li>
<p>Update <code>uv@latest</code> to 0.11.32.</p>
</li>
<li>
<p>Update <code>mise@latest</code> to 2026.7.12.</p>
</li>
<li>
<p>Update <code>cyclonedx@latest</code> to 0.33.1.</p>
</li>
<li>
<p>Update <code>cargo-neat@latest</code> to 0.5.2.</p>
</li>
</ul>
<h2>2.85.0</h2>
<ul>
<li>
<p>Support <code>wild</code> (alias: <code>wild-linker</code>). (<a
href="https://redirect.github.com/taiki-e/install-action/pull/1949">#1949</a>)</p>
</li>
<li>
<p>Support <code>bpf-linker</code>. (<a
href="https://redirect.github.com/taiki-e/install-action/pull/1950">#1950</a>)</p>
</li>
<li>
<p>Support <code>rafn</code>. (<a
href="https://redirect.github.com/taiki-e/install-action/pull/1935">#1935</a>,
thanks <a
href="https://github.com/DarkWanderer"><code>@​DarkWanderer</code></a>)</p>
</li>
<li>
<p>Update <code>cargo-neat@latest</code> to 0.5.1.</p>
</li>
<li>
<p>Update <code>zizmor@latest</code> to 1.28.0.</p>
</li>
<li>
<p>Update <code>wasmtime@latest</code> to 47.0.2.</p>
</li>
<li>
<p>Update <code>uv@latest</code> to 0.11.31.</p>
</li>
<li>
<p>Update <code>syft@latest</code> to 1.49.0.</p>
</li>
</ul>
<h2>2.84.1</h2>
<ul>
<li>
<p>Update <code>wasmtime@latest</code> to 47.0.1.</p>
</li>
<li>
<p>Update <code>wasm-tools@latest</code> to 1.254.0.</p>
</li>
<li>
<p>Update <code>uv@latest</code> to 0.11.30.</p>
</li>
<li>
<p>Update <code>mise@latest</code> to 2026.7.11.</p>
</li>
<li>
<p>Update <code>cargo-neat@latest</code> to 0.5.0.</p>
</li>
<li>
<p>Update <code>cargo-crap@latest</code> to 0.3.1.</p>
</li>
<li>
<p>Update <code>biome@latest</code> to 2.5.5.</p>
</li>
</ul>
</blockquote>
</details>
<details>
<summary>Changelog</summary>
<p><em>Sourced from <a
href="https://github.com/taiki-e/install-action/blob/main/CHANGELOG.md">taiki-e/install-action's
changelog</a>.</em></p>
<blockquote>
<h1>Changelog</h1>
<p>All notable changes to this project will be documented in this
file.</p>
<p>This project adheres to <a href="https://semver.org">Semantic
Versioning</a>.</p>
<!-- raw HTML omitted -->
<h2>[Unreleased]</h2>
<h2>[2.85.2] - 2026-07-26</h2>
<ul>
<li>
<p>Update <code>prek@latest</code> to 0.4.11.</p>
</li>
<li>
<p>Update <code>mise@latest</code> to 2026.7.13.</p>
</li>
<li>
<p>Update <code>kingfisher@latest</code> to 1.109.0.</p>
</li>
</ul>
<h2>[2.85.1] - 2026-07-25</h2>
<ul>
<li>
<p>Update <code>vacuum@latest</code> to 0.30.0.</p>
</li>
<li>
<p>Update <code>uv@latest</code> to 0.11.32.</p>
</li>
<li>
<p>Update <code>mise@latest</code> to 2026.7.12.</p>
</li>
<li>
<p>Update <code>cyclonedx@latest</code> to 0.33.1.</p>
</li>
<li>
<p>Update <code>cargo-neat@latest</code> to 0.5.2.</p>
</li>
</ul>
<h2>[2.85.0] - 2026-07-23</h2>
<ul>
<li>
<p>Support <code>wild</code> (alias: <code>wild-linker</code>). (<a
href="https://redirect.github.com/taiki-e/install-action/pull/1949">#1949</a>)</p>
</li>
<li>
<p>Support <code>bpf-linker</code>. (<a
href="https://redirect.github.com/taiki-e/install-action/pull/1950">#1950</a>)</p>
</li>
<li>
<p>Support <code>rafn</code>. (<a
href="https://redirect.github.com/taiki-e/install-action/pull/1935">#1935</a>,
thanks <a
href="https://github.com/DarkWanderer"><code>@​DarkWanderer</code></a>)</p>
</li>
<li>
<p>Update <code>cargo-neat@latest</code> to 0.5.1.</p>
</li>
<li>
<p>Update <code>zizmor@latest</code> to 1.28.0.</p>
</li>
<li>
<p>Update <code>wasmtime@latest</code> to 47.0.2.</p>
</li>
<li>
<p>Update <code>uv@latest</code> to 0.11.31.</p>
</li>
<li>
<p>Update <code>syft@latest</code> to 1.49.0.</p>
</li>
</ul>
<!-- raw HTML omitted -->
</blockquote>
<p>... (truncated)</p>
</details>
<details>
<summary>Commits</summary>
<ul>
<li><a
href="https://github.com/taiki-e/install-action/commit/41049aa56687c35e0afa74eed4f09cec4f9afabf"><code>41049aa</code></a>
Release 2.85.2</li>
<li><a
href="https://github.com/taiki-e/install-action/commit/dfcf36552b1b743e910b9b9b6b114a08e56a1e5e"><code>dfcf365</code></a>
Update <code>prek@latest</code> to 0.4.11</li>
<li><a
href="https://github.com/taiki-e/install-action/commit/eea03ccfa855b965a301aa52424915fddc32f855"><code>eea03cc</code></a>
Update <code>mise@latest</code> to 2026.7.13</li>
<li><a
href="https://github.com/taiki-e/install-action/commit/81ca2feb84748ed3fa514707ee1004f4f3ad715a"><code>81ca2fe</code></a>
Update martin manifest</li>
<li><a
href="https://github.com/taiki-e/install-action/commit/cc90ed04bc1a258ea90a83789f24b759a77c9671"><code>cc90ed0</code></a>
Update <code>kingfisher@latest</code> to 1.109.0</li>
<li><a
href="https://github.com/taiki-e/install-action/commit/55639a3362f508fda5154d3451072205f5c32ca6"><code>55639a3</code></a>
Update cargo-shear manifest</li>
<li><a
href="https://github.com/taiki-e/install-action/commit/3d7d7cd5ac7f994c1892ae0c06165095b9139094"><code>3d7d7cd</code></a>
Release 2.85.1</li>
<li><a
href="https://github.com/taiki-e/install-action/commit/d09ccb4fe2105ac9ead6376f7a423ff88defeb57"><code>d09ccb4</code></a>
Update <code>vacuum@latest</code> to 0.30.0</li>
<li><a
href="https://github.com/taiki-e/install-action/commit/ac43dee1a92e482c0e244bdb0bb126908159e245"><code>ac43dee</code></a>
Update <code>uv@latest</code> to 0.11.32</li>
<li><a
href="https://github.com/taiki-e/install-action/commit/49b16979f38f30e44b1f68af9115be9a6ffc5215"><code>49b1697</code></a>
Update prek manifest</li>
<li>Additional commits viewable in <a
href="https://github.com/taiki-e/install-action/compare/a6b2e2dcd845ddd7f509ce4f3ed3d922b80cc5d9...41049aa56687c35e0afa74eed4f09cec4f9afabf">compare
view</a></li>
</ul>
</details>
<br />


[![Dependabot compatibility
score](https://dependabot-badges.githubapp.com/badges/compatibility_score?dependency-name=taiki-e/install-action&package-manager=github_actions&previous-version=2.84.0&new-version=2.85.2)](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores)

Dependabot will resolve any conflicts with this PR as long as you don't
alter it yourself. You can also trigger a rebase manually by commenting
`@dependabot rebase`.

[//]: # (dependabot-automerge-start)
[//]: # (dependabot-automerge-end)

---

<details>
<summary>Dependabot commands and options</summary>
<br />

You can trigger Dependabot actions by commenting on this PR:
- `@dependabot rebase` will rebase this PR
- `@dependabot recreate` will recreate this PR, overwriting any edits
that have been made to it
- `@dependabot show <dependency name> ignore conditions` will show all
of the ignore conditions of the specified dependency
- `@dependabot ignore this major version` will close this PR and stop
Dependabot creating any more for this major version (unless you reopen
the PR or upgrade to it yourself)
- `@dependabot ignore this minor version` will close this PR and stop
Dependabot creating any more for this minor version (unless you reopen
the PR or upgrade to it yourself)
- `@dependabot ignore this dependency` will close this PR and stop
Dependabot creating any more for this dependency (unless you reopen the
PR or upgrade to it yourself)


</details>

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

The 900-commit, 1,720-file merge spans core execution, optimizer, API, tooling, and benchmark behavior and requires final human validation.

Review effort: Balanced
Findings: None

What changed in this PR

Merges DataFusion 55.1.0 into spiceai-55 while retaining Spice-specific correctness and compatibility patches.

Changes:

  • Integrates upstream optimizer, execution, API, dependency, and benchmark updates.
  • Fixes join sum statistics and dialect-aware Date32 unparsing.
  • Preserves Spark feature isolation and Rust 1.94 compatibility.
File Description
datafusion/​** Upstream engine changes and Spice fixes
benchmarks/​** New benchmark framework and suites
docs/​** 55.1 documentation and upgrade guidance
.github/​** CI and dependency automation updates
Root configuration files Toolchain, licensing, and lint updates

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Copilot AI balanced review requested due to automatic review settings October 2, 2026 01:29

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

The 1,720-file, 901-commit merge requires human validation, and its documented Arrow revision does not match the committed pin.

Review effort: Balanced
Findings: 1 Medium severity

Open (1)

Comment thread Cargo.toml
@krinart
krinart marked this pull request as draft October 2, 2026 02:09
@krinart
krinart marked this pull request as ready for review October 2, 2026 02:33
@phillipleblanc
phillipleblanc merged commit 02550cf into spiceai-55 Oct 2, 2026
2 checks passed
krinart added a commit to spiceai/datafusion-ballista that referenced this pull request Oct 2, 2026
krinart added a commit to spiceai/datafusion-federation that referenced this pull request Oct 2, 2026
krinart added a commit to spiceai/datafusion-table-providers that referenced this pull request Oct 2, 2026
krinart added a commit to spiceai/spiceai that referenced this pull request Oct 2, 2026
…nowflake-rs at their arrow-rs bumps

- DataFusion: 02550cf9, the merge of spiceai/datafusion#243 into spiceai-55 (same tree as the previous pin); no longer a temporary pin.
- iceberg-rust: 0284be3d and snowflake-rs: b95a74e8, which move each fork's own arrow-rs pin to 2e2cc330 (Cargo manifests only; Spice builds the same code).
krinart added a commit to spiceai/datafusion-ballista that referenced this pull request Oct 2, 2026
* build: pin arrow-rs at 2e2cc330, the merge of spiceai/arrow-rs#29 into spiceai-59

* build: pin DataFusion at 02550cf9, the merge of spiceai/datafusion#243 into spiceai-55
krinart added a commit to spiceai/datafusion-table-providers that referenced this pull request Oct 2, 2026
* build: pin arrow-rs at 2e2cc330, the merge of spiceai/arrow-rs#29 into spiceai-59

* build: pin DataFusion at 02550cf9, the merge of spiceai/datafusion#243 into spiceai-55
krinart added a commit to spiceai/datafusion-federation that referenced this pull request Oct 2, 2026
* build: pin arrow-rs at 2e2cc330, the merge of spiceai/arrow-rs#29 into spiceai-59

* build: pin DataFusion at 02550cf9, the merge of spiceai/datafusion#243 into spiceai-55
krinart added a commit to spiceai/iceberg-rust that referenced this pull request Oct 2, 2026
lukekim pushed a commit to spiceai/datafusion-table-providers that referenced this pull request Oct 2, 2026
* build: pin arrow-rs at 2e2cc330, the merge of spiceai/arrow-rs#29 into spiceai-59

* build: pin DataFusion at 02550cf9, the merge of spiceai/datafusion#243 into spiceai-55
lukekim pushed a commit to spiceai/iceberg-rust that referenced this pull request Oct 3, 2026
anush008 pushed a commit to anush008/spiceai that referenced this pull request Oct 6, 2026
* build: pin the DataFusion 55 / Arrow 59 fork lines

Moves every fork to its DataFusion 55 line: arrow-rs 59.3.0, DataFusion
55.1.0, Ballista (upstream main on DataFusion 55), federation,
table-providers, iceberg-rust 0.10.1, duckdb-rs 1.4.4 on arrow 59, ADBC 24,
snowflake-rs, spark-connect-rs, spice-rs and spicebench on arrow 59.
datafusion-functions-json comes from crates.io (0.55.6); delta-kernel stays
on 0.27.1 with its default engine split into delta_kernel_default_engine.

docs/dev/fork_patches.md is re-audited against the new lines; the pins stay
marked TEMPORARY until the fork pull requests merge.

* feat: upgrade to DataFusion 55 and Arrow 59

- Statistics flow through DataFusion 55's StatisticsContext. The deprecated
  partition_statistics returns Absent for built-in plans on 55, so every
  wrapper implements child_stats_requests / statistics_from_inputs and the
  optimizer rules, join sizing and Flight batch sizing read real statistics
  again.
- ExecutionPlan / TableProvider wrappers implement the new required and
  defaulted methods explicitly (apply_expressions, replace_children,
  try_to_proto, merge_into, ...), forwarding where they wrap.
- Arrow 59 refuses a MAP whose entries field is nullable: Flight SQL reads
  replace the declared schema before decoding and normalize the map
  afterwards, and legacy IPC maps decode as a list (arrow_tools map_entries).
- Delta Lake refuses Void and interval columns with a structured error
  instead of mis-reading them.
- Ballista 55: Vcores, the protocol-version handshake, task-id based
  cancellation and the append-only task-info model.
- Vortex 0.86 and the remaining API moves (TableReference, TableScanBuilder,
  TableSchema builder, codec proto converter, WriteOp::MergeInto).

* ci: validate the datafusion-table-providers pin against spiceai-55-patches

The DataFusion 55 line of the fork is spiceai-55-patches.

* revert: drop the unfinished guard tests d65aa29 picked up by accident

d65aa29 was meant to change only the CI branch check; it also committed
work-in-progress test edits. Restore those files to 98800a8; the finished
tests land in their own commit.

* fix(duckdb): push regexp_count down NULL-preserving, as DataFusion 55 evaluates it

DataFusion 55 (apache/datafusion#24239) makes regexp_count propagate NULL,
as PostgreSQL does. The DuckDB rendering still wrapped the count in
coalesce(.., 0) to match the old kernel, so a DuckDB-accelerated dataset
answered 0 where local evaluation answers NULL, and a WHERE on the count kept
rows local evaluation drops. Render it as len(regexp_extract_all(..)), which
is NULL for a NULL input or pattern.

* test(cayenne): count -0.0 as equal to 0.0 in the boundary-values test

DataFusion 55 compares floats in SQL by IEEE 754 value, so float_val = 0.0
matches the -0.0 row too (4 rows, not 3). A MemTable in the same session
returns the same rows, so this is the expectation, not Cayenne.

* build: repin the forks and guard their new patches

- Ballista 8bf37ad0: the executor poll loop accepts a vcore semaphore with no
  permits yet. Upstream's assert panicked the loop at startup, so no Spice
  executor registered and every cluster query failed.
- DataFusion d98ae645 (MSRV-clean unparser), federation 9ca84a39,
  table-providers 661d16d1, iceberg 575d81a1: their own tests now run
  against the Spice DataFusion (and, for table-providers, ADBC) forks.
- Guard tests for the unparser rescope, eager aggregation through
  StatisticsContext, SchemaCastScanExec statistics forwarding and Ballista
  TaskInfo fields 11/12; the ledger names them and records the Ballista fix.
- The distributed EXPLAIN snapshot now reads Ballista 55's
  UnknownPartitioning(1) where it printed None.

* build: point the fork pins at the spiceai-*-patches-2 branches

The DataFusion 55 lines of arrow-rs, datafusion, datafusion-ballista,
datafusion-federation, datafusion-table-providers, duckdb-rs, snowflake-rs and
spark-connect-rs are reviewed on new spiceai-*-patches-2 branches, leaving the
earlier spiceai-*-patches branches untouched. Same revisions; only the branch
names in the pin comments, the fork-patch ledger and the table-providers pin
check change.

* fix: re-root a delegated projection on the inner plan before swapping

DataFusion 55's projection pushdown reads a projection's input as the node
being swapped with: `ProjectionExec` collapses the chain starting at
`projection.input()`, `CoalescePartitionsExec` rebuilds from its child. A
wrapper that handed its inner plan the projection unchanged (its input being
the wrapper) got it back still above the wrapper, and the pushdown then nested
another copy on every step: the Cayenne maintained-aggregate soundness test
grew to gigabytes and never finished.

CayenneAccelerationExec, the DuckDB aggregate pushdown marker and
PartitionedUnionExec now rebuild the projection over their inner plan (whose
schema they share) before delegating. Regression tests cover the Cayenne
wrapper and the DuckDB marker.

* fix(cayenne): read legacy inline map data under Arrow 59 and keep the type policy

- Inline data written under a nullable map-entries declaration is read through
  arrow_tools::map_entries, since Arrow 59 refuses that declaration at decode.
- The primary-key fast-path test's control query no longer uses a bare SUM,
  which DataFusion 55 answers from statistics without a scan.
- RunEndEncoded columns stay refused: the Vortex pin can store them, but
  accepting them is a product change, so the drift test records the exception.

* fix: bring the remaining tests and connectors in line with DataFusion 55 and Arrow 59

- Flight: repair a server's nullable map-entries declaration before decoding
  (Arrow 59 refuses it), as the Flight SQL path does.
- Federation: SQLite and MySQL refuse a filter on a volatile projection output
  (DataFusion fork unparser), Turso forwards it, MSSQL opts out.
- Declared types: parse the decimal form directly so an out-of-range
  precision/scale still names the problem under arrow-rs's stricter parser.
- Expectations that DataFusion 55 changed: predicate reordering (reorder_q8),
  binary literals rendered as X'..', sorts moved below projections (vector
  scans), the NSQL context's version and Spark function list, and the
  null-aware anti join corrected in Ballista's distributed planner.
- json_get tests follow datafusion-functions-json 0.55 (negative array
  indices, json_as_text for a string cast, integral floats in json_get_int).

* fix(vortex): port the merged point-lookup and key-block code to DataFusion 55 and Vortex 0.86

Trunk's point-lookup and key-block changes were written against DataFusion 54 and
the earlier Vortex pin. The scan now binds its projection and filter before
handing them to Vortex (the reader takes bound expressions), the point read keeps
trunk's optimize-then-bind of its conjuncts, dynamic filters are detected through
DynamicFilterTracking, and TableSchema is built with From. Two merged tests move
to the current Vortex Arrow conversion and DataFusion's non-zero batch_size.

* fix(arrow_tools): keep the producer's dictionary ids when repairing a map declaration

Repairing a nullable map-entries declaration re-encoded the IPC schema, and
re-encoding renumbers dictionaries from zero while the dictionary batches that
follow keep the producer's ids. A producer that numbers them differently then
failed to decode, or decoded each dictionary column against another's dictionary
and returned wrong values with no error (reproduced on the Flight SQL read, the
Flight read and the Flight subscribe paths).

The repair now patches the schema message in place: one byte per map field (the
entries nullability, or the Map-to-List type tag for the decodable form),
verified with the flatbuffers verifier and checked against the expected schema;
an unexpected layout is refused rather than passed through.

* fix(runtime): keep the built-in semantics of functions datafusion-spark 55 would shadow

datafusion-spark 55 registers Spark versions of power/pow, atan2 and concat_ws
over the built-ins (infinity instead of an error for zero to a negative power,
Float64 widening, array flattening) and adds hypot, monthname, quote and weekday.
None of that is a deliberate change, so those functions join trunc, date_trunc and
date_part in the withheld list, and a test pins exactly which names resolve to a
Spark implementation. The NSQL context now lists the functions the session
actually registers, so it no longer advertises the withheld ones.

* fix(duckdb): keep a binary literal out of SQL sent to DuckDB

DataFusion 55's unparser renders a binary literal as X'ff', which DuckDB 1.4.4
reads as the text 'xff': through DuckDB federation, b = X'ff' returned the row
holding the bytes 'xff' instead of 0xFF (and X'' missed the empty blob), with no
error. Expressions with a binary literal now stay local; a NULL still pushes
down. Covered end to end by duckdb_binary_literal_filters_agree_with_local_evaluation.

* build: pin DataFusion at 006f3d21, which drops input sums from join statistics

DataFusion 55 answers a bare SUM from exact statistics, and inner/outer join
statistics carried each input's sum through, so SUM over a join returned the whole
input table's sum (silent wrong results; caught by the Cayenne result-correctness
parity tests). The fork fix is spiceai/datafusion#239; the ledger records it with
those parity tests as its guard.

* test(runtime-datafusion): drop a redundant clone in the binary-literal test

* test: fix the Spark built-ins table width and report the pinned DataFusion revision

- the_built_session_keeps_the_built_in_math_and_string_functions: the expected
  table padded one column a character wider than the rendered output; the values
  were already right.
- spice-substrait-compliance reports DATAFUSION_FORK_REV, which still named the
  pre-upgrade revision; it now matches the workspace pin, as its test requires.

* test(oracle): re-record the pushdown plan for DataFusion 55's decimal literal display

DataFusion 55 prints a decimal literal as Decimal128(123.45,10,2) where 54 printed
Decimal128(Some(12345),10,2). The SQL sent to Oracle is unchanged.

* build: repin the forks to their merged fix PRs and carry spiceai-54 forward

- DataFusion f22d1a74: spiceai-54 forward-merged into the 55 line
  (spiceai/datafusion#241), bringing spiceai#237 (a Date32 literal's cast uses the
  dialect's date type; SQLite read CAST('…' AS DATE) as a number, so date ranges
  pushed to SQLite matched nothing) and metadata-column listing pruning.
- arrow-rs be4918e7, Ballista 7c4d54c2, table-providers 3045960a, iceberg 50e982dc,
  arrow-adbc 674d6184: the merged CI fix PRs.
- Guard for spiceai#237: a Date32 range unparsed for SQLite keeps the rows it selects
  (runs the SQL on SQLite with --features sqlite); it fails on the previous pin.
- Ledger rows for both; the substrait harness reports the new DataFusion pin.

* build: pin DataFusion at ae719330, the merge of spiceai/datafusion#241

Same tree as f22d1a74 (spiceai-54 forward-merged into the 55 line); now the head
of spiceai-55-patches-2 rather than a temporary branch.

* style: rustfmt the SQLite date-range federation test

* build: pin arrow-rs at its spiceai-59 merge and DataFusion at spiceai-55-patches-3

- arrow-rs: 2e2cc330, the merge of spiceai/arrow-rs#29 into spiceai-59 (same tree as the previous pin).
- DataFusion: fc84f2d1 on spiceai-55-patches-3, which carries the Date32 literal fix (spiceai/datafusion#237) but not the metadata-column listing pruning (spiceai#229); the ledger records that pruning as not carried on 55 until it has a guard here.
- Correct the spark-connect-rs branch annotation to spiceai-59-patches-2.

* build: pin DataFusion at the spiceai-55 merge, and iceberg-rust and snowflake-rs at their arrow-rs bumps

- DataFusion: 02550cf9, the merge of spiceai/datafusion#243 into spiceai-55 (same tree as the previous pin); no longer a temporary pin.
- iceberg-rust: 0284be3d and snowflake-rs: b95a74e8, which move each fork's own arrow-rs pin to 2e2cc330 (Cargo manifests only; Spice builds the same code).

* build: repin ballista, table-providers, iceberg-rust and spark-connect-rs to their branch heads

- datafusion-ballista 557ae4ea: arrow-rs/DataFusion landed pins, and the Rust 1.99 clippy fixes (spiceai#70, spiceai#71, spiceai#72).
- datafusion-table-providers 5f50cfd5 and iceberg-rust a0bd6b04: arrow-rs/DataFusion landed pins (spiceai#82, spiceai#53).
- spark-connect-rs 3bf6421a: regenerated lockfile.

datafusion-federation and duckdb-rs stay at 9ca84a39 and 8ee43073: table-providers still pins exactly those revisions, and a second revision of either would put two copies of the crate in the graph. Their newer heads only change their own Cargo manifests and lockfile.

* build: pin iceberg-rust, snowflake-rs, spark-connect-rs, spice-rs and spicebench at their merges

The fork PRs these pins pointed at have merged, mostly as squash merges, so
several pinned commits are no longer on their base branches. Each pin now
points at the merge on the fork's base branch:

- iceberg-rust spiceai-0.10.1-df-55 @ 3e2a14e (spiceai/iceberg-rust#50, spiceai#54)
- snowflake-rs spiceai-59 @ f555738 (spiceai/snowflake-rs#13, spiceai#15)
- spark-connect-rs spiceai-59-2 @ 18ae9bd (spiceai/spark-connect-rs#15, spiceai#16)
- spice-rs trunk @ 4429f39 (spiceai/spice-rs#103)
- spicebench trunk @ 8a90555 (spiceai/spicebench#284)

datafusion-federation stays at 9ca84a3, which is on spiceai-55 because
spiceai/datafusion-federation#88 was a merge commit; only its branch note
changes.

What changes for Spice: lint-only rewrites in snowflake-api, and tests in
iceberg-datafusion and spark-connect-core. The other differences are in
the forks' own [patch.crates-io] sections, which cargo ignores for a
dependency. spice-rs's tree is identical.

* build: pin the forks at their landed version-branch revisions

- arrow-adbc 6e4119ac (spiceai-24, spiceai#7), iceberg-rust 3e2a14e8 (spiceai-0.10.1-df-55, spiceai#50),
  datafusion-table-providers 465926a3 (spiceai-55, spiceai#80), snowflake-rs f5557381 (spiceai-59, spiceai#13),
  spark-connect-rs 18ae9bd3 (spiceai-59-2, spiceai#15), spice-rs 4429f395 and spicebench 8a90555d (trunk).
- datafusion-federation stays at 9ca84a39, now on spiceai-55 (spiceai#88): datafusion-table-providers pins
  exactly that revision.
- duckdb-rs stays at 8ee43073 and temporary: spiceai#49 was squash-merged, so that revision is not on
  spiceai-1.4.4, and datafusion-table-providers pins it, so both have to move together.
- datafusion-ballista is unchanged; its version PR is still open.

* build: pin datafusion-ballista at upstream 55.0.0-rc1

spiceai/datafusion-ballista#73 merged upstream's 55.0.0-rc1 tag into
spiceai-55-patches-2, the head of spiceai/datafusion-ballista#68. spiceai#68 is
still open, waiting on license-compliance decisions. Pin 718429f, the spiceai#73
merge, until spiceai#68 lands on spiceai-55.

ballista-core, -executor, -scheduler, -history and -api-types move from
54.0.0 to 55.0.0; no packages are added or removed in Cargo.lock.
`cargo check -p spiced --locked` passes on rustc 1.98.1.

* docs(fork_patches): record the datafusion-ballista pin at 718429f4 (upstream 55.0.0-rc1, spiceai#73)

* build: pin datafusion-ballista at a7c4c585, the merge of spiceai/datafusion-ballista#68 into spiceai-55

* build: pin datafusion-ballista at the spiceai-55 merge of spiceai#68

spiceai/datafusion-ballista#68 merged into spiceai-55 as a7c4c58, a merge
commit whose tree is identical to 718429f, the previous pin
(`git diff 718429f a7c4c58` is empty). Repoint the pin at the landed
revision and drop the TEMPORARY marker from the fork-patch ledger, as
scripts/check_fork_patches.py asks. The ledger's ballista section now
names spiceai-55 and upstream 55.0.0-rc1 as its base.

Cargo.lock changes only the source revision of the five ballista packages.

* perf(cayenne): keep small Vortex scans whole on the main query path

DataFusion 55 lowered `repartition_file_min_size` from 10 MiB to 1 MiB
(apache/datafusion#22439), so the main Cayenne scan byte-range-splits a 1-10 MiB
dimension table `target_partitions` ways and pays a Vortex footer open per range.
The main scan now opts small scans out below 10 MiB, as the internal and
protected-snapshot scans already do below 256 MiB; Parquet listing scans keep the
new default. `small_snapshot_groups_opt_out_of_repartitioning` now asserts the
main scan's opt-out and fails with the threshold back at 0.

* docs(fork_patches): record duckdb-rs on the protected spiceai-1.4.4-patches-2 head, which table-providers pins

* test(duckdb): build CreateExternalTable with DataFusion 55's locations field

* ci: validate the datafusion-table-providers pin against spiceai-55, where spiceai#80 landed

* test(cayenne): size the point-lookup fan-out control above the main scan's 10 MiB opt-out

* fix(data-accelerator-api): implement DataFusion 55's apply_expressions for KeepFirstExec

* test: port trunk's new tests to DataFusion 55 (CreateExternalTable locations, common::TableReference)

* fix(cayenne): keep -0.0 and 0.0 equal in pushed-down float comparisons and memory-mode indexes

DataFusion 55 compares floats with the two zeros equal (apache/datafusion#22835),
while Vortex and the Arrow row encoding order them apart. Spell out both zeros
when a float comparison or IN list with a zero literal is pushed into a Vortex
scan, keep float column-vs-column comparisons above the scan, and hash
memory-mode index keys with the file-mode KeyEncoder.

* Fix lint

---------

Co-authored-by: Luke Kim <80174+lukekim@users.noreply.github.com>
peasee pushed a commit to spiceai/spiceai that referenced this pull request Oct 8, 2026
…n fixes backported from 54 (#14800)

* build: pin the DataFusion 55 / Arrow 59 fork lines

Moves every fork to its DataFusion 55 line: arrow-rs 59.3.0, DataFusion
55.1.0, Ballista (upstream main on DataFusion 55), federation,
table-providers, iceberg-rust 0.10.1, duckdb-rs 1.4.4 on arrow 59, ADBC 24,
snowflake-rs, spark-connect-rs, spice-rs and spicebench on arrow 59.
datafusion-functions-json comes from crates.io (0.55.6); delta-kernel stays
on 0.27.1 with its default engine split into delta_kernel_default_engine.

docs/dev/fork_patches.md is re-audited against the new lines; the pins stay
marked TEMPORARY until the fork pull requests merge.

* feat: upgrade to DataFusion 55 and Arrow 59

- Statistics flow through DataFusion 55's StatisticsContext. The deprecated
  partition_statistics returns Absent for built-in plans on 55, so every
  wrapper implements child_stats_requests / statistics_from_inputs and the
  optimizer rules, join sizing and Flight batch sizing read real statistics
  again.
- ExecutionPlan / TableProvider wrappers implement the new required and
  defaulted methods explicitly (apply_expressions, replace_children,
  try_to_proto, merge_into, ...), forwarding where they wrap.
- Arrow 59 refuses a MAP whose entries field is nullable: Flight SQL reads
  replace the declared schema before decoding and normalize the map
  afterwards, and legacy IPC maps decode as a list (arrow_tools map_entries).
- Delta Lake refuses Void and interval columns with a structured error
  instead of mis-reading them.
- Ballista 55: Vcores, the protocol-version handshake, task-id based
  cancellation and the append-only task-info model.
- Vortex 0.86 and the remaining API moves (TableReference, TableScanBuilder,
  TableSchema builder, codec proto converter, WriteOp::MergeInto).

* ci: validate the datafusion-table-providers pin against spiceai-55-patches

The DataFusion 55 line of the fork is spiceai-55-patches.

* revert: drop the unfinished guard tests d65aa29 picked up by accident

d65aa29 was meant to change only the CI branch check; it also committed
work-in-progress test edits. Restore those files to 98800a8; the finished
tests land in their own commit.

* fix(duckdb): push regexp_count down NULL-preserving, as DataFusion 55 evaluates it

DataFusion 55 (apache/datafusion#24239) makes regexp_count propagate NULL,
as PostgreSQL does. The DuckDB rendering still wrapped the count in
coalesce(.., 0) to match the old kernel, so a DuckDB-accelerated dataset
answered 0 where local evaluation answers NULL, and a WHERE on the count kept
rows local evaluation drops. Render it as len(regexp_extract_all(..)), which
is NULL for a NULL input or pattern.

* test(cayenne): count -0.0 as equal to 0.0 in the boundary-values test

DataFusion 55 compares floats in SQL by IEEE 754 value, so float_val = 0.0
matches the -0.0 row too (4 rows, not 3). A MemTable in the same session
returns the same rows, so this is the expectation, not Cayenne.

* build: repin the forks and guard their new patches

- Ballista 8bf37ad0: the executor poll loop accepts a vcore semaphore with no
  permits yet. Upstream's assert panicked the loop at startup, so no Spice
  executor registered and every cluster query failed.
- DataFusion d98ae645 (MSRV-clean unparser), federation 9ca84a39,
  table-providers 661d16d1, iceberg 575d81a1: their own tests now run
  against the Spice DataFusion (and, for table-providers, ADBC) forks.
- Guard tests for the unparser rescope, eager aggregation through
  StatisticsContext, SchemaCastScanExec statistics forwarding and Ballista
  TaskInfo fields 11/12; the ledger names them and records the Ballista fix.
- The distributed EXPLAIN snapshot now reads Ballista 55's
  UnknownPartitioning(1) where it printed None.

* build: point the fork pins at the spiceai-*-patches-2 branches

The DataFusion 55 lines of arrow-rs, datafusion, datafusion-ballista,
datafusion-federation, datafusion-table-providers, duckdb-rs, snowflake-rs and
spark-connect-rs are reviewed on new spiceai-*-patches-2 branches, leaving the
earlier spiceai-*-patches branches untouched. Same revisions; only the branch
names in the pin comments, the fork-patch ledger and the table-providers pin
check change.

* fix: re-root a delegated projection on the inner plan before swapping

DataFusion 55's projection pushdown reads a projection's input as the node
being swapped with: `ProjectionExec` collapses the chain starting at
`projection.input()`, `CoalescePartitionsExec` rebuilds from its child. A
wrapper that handed its inner plan the projection unchanged (its input being
the wrapper) got it back still above the wrapper, and the pushdown then nested
another copy on every step: the Cayenne maintained-aggregate soundness test
grew to gigabytes and never finished.

CayenneAccelerationExec, the DuckDB aggregate pushdown marker and
PartitionedUnionExec now rebuild the projection over their inner plan (whose
schema they share) before delegating. Regression tests cover the Cayenne
wrapper and the DuckDB marker.

* fix(cayenne): read legacy inline map data under Arrow 59 and keep the type policy

- Inline data written under a nullable map-entries declaration is read through
  arrow_tools::map_entries, since Arrow 59 refuses that declaration at decode.
- The primary-key fast-path test's control query no longer uses a bare SUM,
  which DataFusion 55 answers from statistics without a scan.
- RunEndEncoded columns stay refused: the Vortex pin can store them, but
  accepting them is a product change, so the drift test records the exception.

* fix: bring the remaining tests and connectors in line with DataFusion 55 and Arrow 59

- Flight: repair a server's nullable map-entries declaration before decoding
  (Arrow 59 refuses it), as the Flight SQL path does.
- Federation: SQLite and MySQL refuse a filter on a volatile projection output
  (DataFusion fork unparser), Turso forwards it, MSSQL opts out.
- Declared types: parse the decimal form directly so an out-of-range
  precision/scale still names the problem under arrow-rs's stricter parser.
- Expectations that DataFusion 55 changed: predicate reordering (reorder_q8),
  binary literals rendered as X'..', sorts moved below projections (vector
  scans), the NSQL context's version and Spark function list, and the
  null-aware anti join corrected in Ballista's distributed planner.
- json_get tests follow datafusion-functions-json 0.55 (negative array
  indices, json_as_text for a string cast, integral floats in json_get_int).

* fix(vortex): port the merged point-lookup and key-block code to DataFusion 55 and Vortex 0.86

Trunk's point-lookup and key-block changes were written against DataFusion 54 and
the earlier Vortex pin. The scan now binds its projection and filter before
handing them to Vortex (the reader takes bound expressions), the point read keeps
trunk's optimize-then-bind of its conjuncts, dynamic filters are detected through
DynamicFilterTracking, and TableSchema is built with From. Two merged tests move
to the current Vortex Arrow conversion and DataFusion's non-zero batch_size.

* fix(arrow_tools): keep the producer's dictionary ids when repairing a map declaration

Repairing a nullable map-entries declaration re-encoded the IPC schema, and
re-encoding renumbers dictionaries from zero while the dictionary batches that
follow keep the producer's ids. A producer that numbers them differently then
failed to decode, or decoded each dictionary column against another's dictionary
and returned wrong values with no error (reproduced on the Flight SQL read, the
Flight read and the Flight subscribe paths).

The repair now patches the schema message in place: one byte per map field (the
entries nullability, or the Map-to-List type tag for the decodable form),
verified with the flatbuffers verifier and checked against the expected schema;
an unexpected layout is refused rather than passed through.

* fix(runtime): keep the built-in semantics of functions datafusion-spark 55 would shadow

datafusion-spark 55 registers Spark versions of power/pow, atan2 and concat_ws
over the built-ins (infinity instead of an error for zero to a negative power,
Float64 widening, array flattening) and adds hypot, monthname, quote and weekday.
None of that is a deliberate change, so those functions join trunc, date_trunc and
date_part in the withheld list, and a test pins exactly which names resolve to a
Spark implementation. The NSQL context now lists the functions the session
actually registers, so it no longer advertises the withheld ones.

* fix(duckdb): keep a binary literal out of SQL sent to DuckDB

DataFusion 55's unparser renders a binary literal as X'ff', which DuckDB 1.4.4
reads as the text 'xff': through DuckDB federation, b = X'ff' returned the row
holding the bytes 'xff' instead of 0xFF (and X'' missed the empty blob), with no
error. Expressions with a binary literal now stay local; a NULL still pushes
down. Covered end to end by duckdb_binary_literal_filters_agree_with_local_evaluation.

* build: pin DataFusion at 006f3d21, which drops input sums from join statistics

DataFusion 55 answers a bare SUM from exact statistics, and inner/outer join
statistics carried each input's sum through, so SUM over a join returned the whole
input table's sum (silent wrong results; caught by the Cayenne result-correctness
parity tests). The fork fix is spiceai/datafusion#239; the ledger records it with
those parity tests as its guard.

* test(runtime-datafusion): drop a redundant clone in the binary-literal test

* test: fix the Spark built-ins table width and report the pinned DataFusion revision

- the_built_session_keeps_the_built_in_math_and_string_functions: the expected
  table padded one column a character wider than the rendered output; the values
  were already right.
- spice-substrait-compliance reports DATAFUSION_FORK_REV, which still named the
  pre-upgrade revision; it now matches the workspace pin, as its test requires.

* test(oracle): re-record the pushdown plan for DataFusion 55's decimal literal display

DataFusion 55 prints a decimal literal as Decimal128(123.45,10,2) where 54 printed
Decimal128(Some(12345),10,2). The SQL sent to Oracle is unchanged.

* build: repin the forks to their merged fix PRs and carry spiceai-54 forward

- DataFusion f22d1a74: spiceai-54 forward-merged into the 55 line
  (spiceai/datafusion#241), bringing #237 (a Date32 literal's cast uses the
  dialect's date type; SQLite read CAST('…' AS DATE) as a number, so date ranges
  pushed to SQLite matched nothing) and metadata-column listing pruning.
- arrow-rs be4918e7, Ballista 7c4d54c2, table-providers 3045960a, iceberg 50e982dc,
  arrow-adbc 674d6184: the merged CI fix PRs.
- Guard for #237: a Date32 range unparsed for SQLite keeps the rows it selects
  (runs the SQL on SQLite with --features sqlite); it fails on the previous pin.
- Ledger rows for both; the substrait harness reports the new DataFusion pin.

* build: pin DataFusion at ae719330, the merge of spiceai/datafusion#241

Same tree as f22d1a74 (spiceai-54 forward-merged into the 55 line); now the head
of spiceai-55-patches-2 rather than a temporary branch.

* style: rustfmt the SQLite date-range federation test

* build: pin arrow-rs at its spiceai-59 merge and DataFusion at spiceai-55-patches-3

- arrow-rs: 2e2cc330, the merge of spiceai/arrow-rs#29 into spiceai-59 (same tree as the previous pin).
- DataFusion: fc84f2d1 on spiceai-55-patches-3, which carries the Date32 literal fix (spiceai/datafusion#237) but not the metadata-column listing pruning (#229); the ledger records that pruning as not carried on 55 until it has a guard here.
- Correct the spark-connect-rs branch annotation to spiceai-59-patches-2.

* build: pin DataFusion at the spiceai-55 merge, and iceberg-rust and snowflake-rs at their arrow-rs bumps

- DataFusion: 02550cf9, the merge of spiceai/datafusion#243 into spiceai-55 (same tree as the previous pin); no longer a temporary pin.
- iceberg-rust: 0284be3d and snowflake-rs: b95a74e8, which move each fork's own arrow-rs pin to 2e2cc330 (Cargo manifests only; Spice builds the same code).

* build: repin ballista, table-providers, iceberg-rust and spark-connect-rs to their branch heads

- datafusion-ballista 557ae4ea: arrow-rs/DataFusion landed pins, and the Rust 1.99 clippy fixes (#70, #71, #72).
- datafusion-table-providers 5f50cfd5 and iceberg-rust a0bd6b04: arrow-rs/DataFusion landed pins (#82, #53).
- spark-connect-rs 3bf6421a: regenerated lockfile.

datafusion-federation and duckdb-rs stay at 9ca84a39 and 8ee43073: table-providers still pins exactly those revisions, and a second revision of either would put two copies of the crate in the graph. Their newer heads only change their own Cargo manifests and lockfile.

* build: pin iceberg-rust, snowflake-rs, spark-connect-rs, spice-rs and spicebench at their merges

The fork PRs these pins pointed at have merged, mostly as squash merges, so
several pinned commits are no longer on their base branches. Each pin now
points at the merge on the fork's base branch:

- iceberg-rust spiceai-0.10.1-df-55 @ 3e2a14e (spiceai/iceberg-rust#50, #54)
- snowflake-rs spiceai-59 @ f555738 (spiceai/snowflake-rs#13, #15)
- spark-connect-rs spiceai-59-2 @ 18ae9bd (spiceai/spark-connect-rs#15, #16)
- spice-rs trunk @ 4429f39 (spiceai/spice-rs#103)
- spicebench trunk @ 8a90555 (spiceai/spicebench#284)

datafusion-federation stays at 9ca84a3, which is on spiceai-55 because
spiceai/datafusion-federation#88 was a merge commit; only its branch note
changes.

What changes for Spice: lint-only rewrites in snowflake-api, and tests in
iceberg-datafusion and spark-connect-core. The other differences are in
the forks' own [patch.crates-io] sections, which cargo ignores for a
dependency. spice-rs's tree is identical.

* build: pin the forks at their landed version-branch revisions

- arrow-adbc 6e4119ac (spiceai-24, #7), iceberg-rust 3e2a14e8 (spiceai-0.10.1-df-55, #50),
  datafusion-table-providers 465926a3 (spiceai-55, #80), snowflake-rs f5557381 (spiceai-59, #13),
  spark-connect-rs 18ae9bd3 (spiceai-59-2, #15), spice-rs 4429f395 and spicebench 8a90555d (trunk).
- datafusion-federation stays at 9ca84a39, now on spiceai-55 (#88): datafusion-table-providers pins
  exactly that revision.
- duckdb-rs stays at 8ee43073 and temporary: #49 was squash-merged, so that revision is not on
  spiceai-1.4.4, and datafusion-table-providers pins it, so both have to move together.
- datafusion-ballista is unchanged; its version PR is still open.

* build: pin datafusion-ballista at upstream 55.0.0-rc1

spiceai/datafusion-ballista#73 merged upstream's 55.0.0-rc1 tag into
spiceai-55-patches-2, the head of spiceai/datafusion-ballista#68. #68 is
still open, waiting on license-compliance decisions. Pin 718429f, the #73
merge, until #68 lands on spiceai-55.

ballista-core, -executor, -scheduler, -history and -api-types move from
54.0.0 to 55.0.0; no packages are added or removed in Cargo.lock.
`cargo check -p spiced --locked` passes on rustc 1.98.1.

* docs(fork_patches): record the datafusion-ballista pin at 718429f4 (upstream 55.0.0-rc1, #73)

* build: pin datafusion-ballista at a7c4c585, the merge of spiceai/datafusion-ballista#68 into spiceai-55

* build: pin datafusion-ballista at the spiceai-55 merge of #68

spiceai/datafusion-ballista#68 merged into spiceai-55 as a7c4c58, a merge
commit whose tree is identical to 718429f, the previous pin
(`git diff 718429f a7c4c58` is empty). Repoint the pin at the landed
revision and drop the TEMPORARY marker from the fork-patch ledger, as
scripts/check_fork_patches.py asks. The ledger's ballista section now
names spiceai-55 and upstream 55.0.0-rc1 as its base.

Cargo.lock changes only the source revision of the five ballista packages.

* perf(cayenne): keep small Vortex scans whole on the main query path

DataFusion 55 lowered `repartition_file_min_size` from 10 MiB to 1 MiB
(apache/datafusion#22439), so the main Cayenne scan byte-range-splits a 1-10 MiB
dimension table `target_partitions` ways and pays a Vortex footer open per range.
The main scan now opts small scans out below 10 MiB, as the internal and
protected-snapshot scans already do below 256 MiB; Parquet listing scans keep the
new default. `small_snapshot_groups_opt_out_of_repartitioning` now asserts the
main scan's opt-out and fails with the threshold back at 0.

* docs(fork_patches): record duckdb-rs on the protected spiceai-1.4.4-patches-2 head, which table-providers pins

* test(duckdb): build CreateExternalTable with DataFusion 55's locations field

* ci: validate the datafusion-table-providers pin against spiceai-55, where #80 landed

* test(cayenne): size the point-lookup fan-out control above the main scan's 10 MiB opt-out

* fix(data-accelerator-api): implement DataFusion 55's apply_expressions for KeepFirstExec

* test: port trunk's new tests to DataFusion 55 (CreateExternalTable locations, common::TableReference)

* fix(cayenne): keep -0.0 and 0.0 equal in pushed-down float comparisons and memory-mode indexes

DataFusion 55 compares floats with the two zeros equal (apache/datafusion#22835),
while Vortex and the Arrow row encoding order them apart. Spell out both zeros
when a float comparison or IN list with a zero literal is pushed into a Vortex
scan, keep float column-vs-column comparisons above the scan, and hash
memory-mode index keys with the file-mode KeyEncoder.

* build: move to iceberg-rust 0.11.0 (RC) and opendal 0.58

Points the iceberg crates at spiceai/iceberg-rust#55, which merges the
upstream 0.11.x release branch onto the DataFusion 55 fork line.

opendal moves to 0.58 to match iceberg-storage-opendal, so one opendal
stack is built instead of two. 0.58 removes the OperatorBuilder step:
Operator::new returns the operator with the same default layers that
finish() applied.

* Bump datafusion to 55.2-rc1

* build: pin iceberg-rust at spiceai-0.11.0-df-55 (6317d03e)

spiceai/iceberg-rust#55 merged into spiceai-0.11.0-df-55. Same Rust code as
the previous pin (e02547cb); the branch adds only a Python CI pin and a
doc-comment fix. fork_patches.md records the new branch and that SigV4 is
now an AuthManager in the fork's REST catalog.

* build: keep the lockfile's unrelated resolutions from the DataFusion bump

The previous commit's lockfile update also re-resolved five crates with
range requirements (hashbrown, heck, base64). Restore them so the iceberg
repin changes only the iceberg-rust source lines.

* build: pin iceberg-rust at spiceai-0.11.0-df-55 (7735caa2)

Picks up spiceai/iceberg-rust#49 (IcebergTableProvider::catalog and
::table_ident) and #56 (the fork's own DataFusion pin, which does not
affect this build). #49's fork_patches row and its compile guard land
with #14588, which first calls the getters.

* build: pin datafusion at spiceai/datafusion#250 (55.2.0-rc1 and the 54 backports)

Moves the DataFusion fork from 02550cf9 (55.1.0) to 1749ce4e, the head of
spiceai/datafusion#250, which merges into #249: `spiceai-55` at 55.2.0-rc1
(spiceai/datafusion#248) plus the `spiceai-54` patches #233, #234 and #245
carried onto 55, and #250's review fix. The guards for those follow in the
next commit.

The pin names a pull request branch, so the ledger row is marked TEMPORARY
and `scripts/check_fork_patches.py` refuses to let it land until #250 and
#249 merge and the pin moves to the `spiceai-55` merge commit.

The 55.2.0-rc1 merge changes no Spice patch. Outside `Cargo.lock`, the
fork's 02550cf9..27351ab0 diff has the same patch-id as upstream's
55.1.0..55.2.0-rc1 diff, except `datafusion-spark`'s `quote.rs`, which the
fork's backport had already made identical to 55.2.0-rc1's. That backport's
ledger row is dropped; upstream carries the fix as apache/datafusion#25277.
`Cargo.lock` changes only the 36 DataFusion entries.

* test: guard the DataFusion unparser and footer-cache fixes backported from 54

Each backported fix gets a guard here, so the next re-cut of the fork
cannot drop it silently, plus a ledger row naming the guard.

- spiceai/datafusion#233 (#14373): a join that is another join's right
  input stays a parenthesised joined table on that join's right.
  `a_join_that_is_another_joins_right_input_stays_on_its_right` checks
  that `b` is introduced before the outer `ON` names it and that a LEFT
  JOIN keeps a filter from inside it out of `WHERE`. With the `sqlite`
  feature it runs the SQL on SQLite and compares the rows with
  DataFusion's for the same plan.
- spiceai/datafusion#234 (#14375): a `Limit` that is a join input bounds
  that input, not the join.
  `a_limit_on_a_join_input_bounds_that_input_rather_than_the_join`, with
  the same row comparison.
- spiceai/datafusion#250: a projection over a join that passes up a mark
  join's mark is refused rather than unparsed with the mark unbound.
  `a_mark_a_projection_over_a_join_passes_up_is_refused`.
- spiceai/datafusion#245 (#12952): the file metadata cache frees an
  evicted footer's hit counter. 55's `DefaultCache` already does, so 55
  carries no patch; `crates/cayenne/tests/footer_cache_hit_counter_test.rs`
  pins the behaviour by counting the live heap across 20,000 cold footer
  reads into a full cache.

At 02550cf9 the three unparser guards fail. SQLite rejects the nested
join's SQL ("ON clause references tables to its right"), returns 1 row
where the plan returns 2 for `b FULL JOIN (c LIMIT 1)`, and the mark join
renders as `… LEFT OUTER JOIN (b) ON a.id = b.id WHERE NOT c.mark`.
Against a local DataFusion build with the hit-counter prune removed, the
footer-cache guard measures 2,646,176 bytes of heap growth. At 1749ce4e all
four pass, and the footer-cache growth is 0 bytes.

* build: pin datafusion at spiceai/datafusion#249's merge commit

Moves the fork pin from #250's branch head to bebc4d58, the merge of #249
into `spiceai-55`. Its tree is #249's head, cbda233a: 55.2.0-rc1 plus the
backported #233, #234 and #245. The pin names a landed revision, so the
TEMPORARY marker comes off and `scripts/check_fork_patches.py` passes.

#250 is not part of this revision, so its guard and ledger row are removed.

* Update datafusion version to 55.2 in Cargo.toml

* docs(fork_patches): record the unparser frame refactor #249 carries

`cbda233a6` is the fourth Spice commit at the pinned revision. It moves the
join-input bookkeeping that #233 and #234 add out of
`select_to_sql_recursively_inner`, so an unoptimised build's frame stays
where it was. It had no ledger row, so a re-cut could drop it without the
audit noticing.

Its row is a GAP, with an Open gaps entry. What it buys is stack headroom
in a debug build: without it, the fork's `roundtrip_statement` needs
2,112 KiB and overflows a 2 MiB test thread. This repo runs tests on 8 MiB
threads, so no test here sees the difference deterministically. The entry
names the fork-side check to run at a re-cut instead.

* Fix tests

* build: pin DataFusion at spiceai/datafusion#251, so MAX of a partition column skips empty partitions (#14801)

* build: pin datafusion at spiceai/datafusion#251's merge commit

Moves the DataFusion fork pin from #249's merge commit to the head of
`spiceai-55`, which adds two patches on top of it:

- #251: a file that holds no rows gets no partition-column min/max. Without
  it, `MIN`/`MAX` of a partition column is answered from the listing's
  statistics with the value of a partition whose files are all empty:
  `max(p)` answers '99' for a `p=99` directory holding one empty file, where
  the rows give '3'.
- #240: the listing prunes files by metadata-column predicates before
  opening them and reports those filters `Exact` (#229 on `spiceai-54`).

Neither changes a manifest, so `Cargo.lock` only moves its source revision.
Each patch gets a ledger row and a guard in the listing connector's tests,
which the gate runs as library tests.

* fix(listing): keep every predicate residual on the `_location` fast path

The pin now carries spiceai/datafusion#240, so the inner listing reports a
metadata-column predicate such as `_size < 50` `Exact`.
`MetadataPruningListingTable` handed pushdown classification to the inner
listing whenever a `_location` predicate was present, but its `_location`
fast path opens the named objects and applies no other predicate, so
`_location = x AND _size < 50` returned the rows of a file the `_size`
predicate excludes.

Report every predicate `Inexact` on that path as well, so a residual
`FilterExec` re-applies it. This is #14790's change, which makes the same fix
for partition predicates, taken byte for byte so that whichever of the two
lands second merges cleanly. The new test goes through
`create_listing_table` with `_location` and `_size` enabled.

* test(metadata): expect the listing to answer `_size = 2319` exactly

With spiceai/datafusion#240 the listing prunes files by a metadata-column
predicate and reports it `Exact`, so the plan for `met_size`, which enables
only `_size`, has no residual `Filter`. The scan lists only the four 2319-byte
files; `data_0.parquet` is 2317 bytes, as its S3 `Content-Length` shows.

Regenerated by running `metadata::s3_metadata_columns` and accepted with
`cargo insta accept`. The rest of the test passes unchanged at this pin,
including the `_location` plans and the returned rows.

* fix(data-connector-api): satisfy pedantic clippy in the metadata-pruning listing test

`make lint-rust`'s `--tests` pass rejected two lints in
`listing_table_prunes_files_by_a_metadata_column_predicate`:
`redundant_closure_for_method_calls` (`|group| group.len()`) and
`format_collect` (`map(format!).collect::<String>()`). Use
`FileGroup::len` and build the file contents with `writeln!` into one
`String`. The test is unchanged in what it checks.

---------

Co-authored-by: Luke Kim <80174+lukekim@users.noreply.github.com>
Co-authored-by: Sergei Grebnov <sergei.grebnov@gmail.com>
wangyusheng1985 pushed a commit to wangyusheng1985/spiceai that referenced this pull request Oct 8, 2026
…n fixes backported from 54 (spiceai#14800)

* build: pin the DataFusion 55 / Arrow 59 fork lines

Moves every fork to its DataFusion 55 line: arrow-rs 59.3.0, DataFusion
55.1.0, Ballista (upstream main on DataFusion 55), federation,
table-providers, iceberg-rust 0.10.1, duckdb-rs 1.4.4 on arrow 59, ADBC 24,
snowflake-rs, spark-connect-rs, spice-rs and spicebench on arrow 59.
datafusion-functions-json comes from crates.io (0.55.6); delta-kernel stays
on 0.27.1 with its default engine split into delta_kernel_default_engine.

docs/dev/fork_patches.md is re-audited against the new lines; the pins stay
marked TEMPORARY until the fork pull requests merge.

* feat: upgrade to DataFusion 55 and Arrow 59

- Statistics flow through DataFusion 55's StatisticsContext. The deprecated
  partition_statistics returns Absent for built-in plans on 55, so every
  wrapper implements child_stats_requests / statistics_from_inputs and the
  optimizer rules, join sizing and Flight batch sizing read real statistics
  again.
- ExecutionPlan / TableProvider wrappers implement the new required and
  defaulted methods explicitly (apply_expressions, replace_children,
  try_to_proto, merge_into, ...), forwarding where they wrap.
- Arrow 59 refuses a MAP whose entries field is nullable: Flight SQL reads
  replace the declared schema before decoding and normalize the map
  afterwards, and legacy IPC maps decode as a list (arrow_tools map_entries).
- Delta Lake refuses Void and interval columns with a structured error
  instead of mis-reading them.
- Ballista 55: Vcores, the protocol-version handshake, task-id based
  cancellation and the append-only task-info model.
- Vortex 0.86 and the remaining API moves (TableReference, TableScanBuilder,
  TableSchema builder, codec proto converter, WriteOp::MergeInto).

* ci: validate the datafusion-table-providers pin against spiceai-55-patches

The DataFusion 55 line of the fork is spiceai-55-patches.

* revert: drop the unfinished guard tests d65aa29 picked up by accident

d65aa29 was meant to change only the CI branch check; it also committed
work-in-progress test edits. Restore those files to 98800a8; the finished
tests land in their own commit.

* fix(duckdb): push regexp_count down NULL-preserving, as DataFusion 55 evaluates it

DataFusion 55 (apache/datafusion#24239) makes regexp_count propagate NULL,
as PostgreSQL does. The DuckDB rendering still wrapped the count in
coalesce(.., 0) to match the old kernel, so a DuckDB-accelerated dataset
answered 0 where local evaluation answers NULL, and a WHERE on the count kept
rows local evaluation drops. Render it as len(regexp_extract_all(..)), which
is NULL for a NULL input or pattern.

* test(cayenne): count -0.0 as equal to 0.0 in the boundary-values test

DataFusion 55 compares floats in SQL by IEEE 754 value, so float_val = 0.0
matches the -0.0 row too (4 rows, not 3). A MemTable in the same session
returns the same rows, so this is the expectation, not Cayenne.

* build: repin the forks and guard their new patches

- Ballista 8bf37ad0: the executor poll loop accepts a vcore semaphore with no
  permits yet. Upstream's assert panicked the loop at startup, so no Spice
  executor registered and every cluster query failed.
- DataFusion d98ae645 (MSRV-clean unparser), federation 9ca84a39,
  table-providers 661d16d1, iceberg 575d81a1: their own tests now run
  against the Spice DataFusion (and, for table-providers, ADBC) forks.
- Guard tests for the unparser rescope, eager aggregation through
  StatisticsContext, SchemaCastScanExec statistics forwarding and Ballista
  TaskInfo fields 11/12; the ledger names them and records the Ballista fix.
- The distributed EXPLAIN snapshot now reads Ballista 55's
  UnknownPartitioning(1) where it printed None.

* build: point the fork pins at the spiceai-*-patches-2 branches

The DataFusion 55 lines of arrow-rs, datafusion, datafusion-ballista,
datafusion-federation, datafusion-table-providers, duckdb-rs, snowflake-rs and
spark-connect-rs are reviewed on new spiceai-*-patches-2 branches, leaving the
earlier spiceai-*-patches branches untouched. Same revisions; only the branch
names in the pin comments, the fork-patch ledger and the table-providers pin
check change.

* fix: re-root a delegated projection on the inner plan before swapping

DataFusion 55's projection pushdown reads a projection's input as the node
being swapped with: `ProjectionExec` collapses the chain starting at
`projection.input()`, `CoalescePartitionsExec` rebuilds from its child. A
wrapper that handed its inner plan the projection unchanged (its input being
the wrapper) got it back still above the wrapper, and the pushdown then nested
another copy on every step: the Cayenne maintained-aggregate soundness test
grew to gigabytes and never finished.

CayenneAccelerationExec, the DuckDB aggregate pushdown marker and
PartitionedUnionExec now rebuild the projection over their inner plan (whose
schema they share) before delegating. Regression tests cover the Cayenne
wrapper and the DuckDB marker.

* fix(cayenne): read legacy inline map data under Arrow 59 and keep the type policy

- Inline data written under a nullable map-entries declaration is read through
  arrow_tools::map_entries, since Arrow 59 refuses that declaration at decode.
- The primary-key fast-path test's control query no longer uses a bare SUM,
  which DataFusion 55 answers from statistics without a scan.
- RunEndEncoded columns stay refused: the Vortex pin can store them, but
  accepting them is a product change, so the drift test records the exception.

* fix: bring the remaining tests and connectors in line with DataFusion 55 and Arrow 59

- Flight: repair a server's nullable map-entries declaration before decoding
  (Arrow 59 refuses it), as the Flight SQL path does.
- Federation: SQLite and MySQL refuse a filter on a volatile projection output
  (DataFusion fork unparser), Turso forwards it, MSSQL opts out.
- Declared types: parse the decimal form directly so an out-of-range
  precision/scale still names the problem under arrow-rs's stricter parser.
- Expectations that DataFusion 55 changed: predicate reordering (reorder_q8),
  binary literals rendered as X'..', sorts moved below projections (vector
  scans), the NSQL context's version and Spark function list, and the
  null-aware anti join corrected in Ballista's distributed planner.
- json_get tests follow datafusion-functions-json 0.55 (negative array
  indices, json_as_text for a string cast, integral floats in json_get_int).

* fix(vortex): port the merged point-lookup and key-block code to DataFusion 55 and Vortex 0.86

Trunk's point-lookup and key-block changes were written against DataFusion 54 and
the earlier Vortex pin. The scan now binds its projection and filter before
handing them to Vortex (the reader takes bound expressions), the point read keeps
trunk's optimize-then-bind of its conjuncts, dynamic filters are detected through
DynamicFilterTracking, and TableSchema is built with From. Two merged tests move
to the current Vortex Arrow conversion and DataFusion's non-zero batch_size.

* fix(arrow_tools): keep the producer's dictionary ids when repairing a map declaration

Repairing a nullable map-entries declaration re-encoded the IPC schema, and
re-encoding renumbers dictionaries from zero while the dictionary batches that
follow keep the producer's ids. A producer that numbers them differently then
failed to decode, or decoded each dictionary column against another's dictionary
and returned wrong values with no error (reproduced on the Flight SQL read, the
Flight read and the Flight subscribe paths).

The repair now patches the schema message in place: one byte per map field (the
entries nullability, or the Map-to-List type tag for the decodable form),
verified with the flatbuffers verifier and checked against the expected schema;
an unexpected layout is refused rather than passed through.

* fix(runtime): keep the built-in semantics of functions datafusion-spark 55 would shadow

datafusion-spark 55 registers Spark versions of power/pow, atan2 and concat_ws
over the built-ins (infinity instead of an error for zero to a negative power,
Float64 widening, array flattening) and adds hypot, monthname, quote and weekday.
None of that is a deliberate change, so those functions join trunc, date_trunc and
date_part in the withheld list, and a test pins exactly which names resolve to a
Spark implementation. The NSQL context now lists the functions the session
actually registers, so it no longer advertises the withheld ones.

* fix(duckdb): keep a binary literal out of SQL sent to DuckDB

DataFusion 55's unparser renders a binary literal as X'ff', which DuckDB 1.4.4
reads as the text 'xff': through DuckDB federation, b = X'ff' returned the row
holding the bytes 'xff' instead of 0xFF (and X'' missed the empty blob), with no
error. Expressions with a binary literal now stay local; a NULL still pushes
down. Covered end to end by duckdb_binary_literal_filters_agree_with_local_evaluation.

* build: pin DataFusion at 006f3d21, which drops input sums from join statistics

DataFusion 55 answers a bare SUM from exact statistics, and inner/outer join
statistics carried each input's sum through, so SUM over a join returned the whole
input table's sum (silent wrong results; caught by the Cayenne result-correctness
parity tests). The fork fix is spiceai/datafusion#239; the ledger records it with
those parity tests as its guard.

* test(runtime-datafusion): drop a redundant clone in the binary-literal test

* test: fix the Spark built-ins table width and report the pinned DataFusion revision

- the_built_session_keeps_the_built_in_math_and_string_functions: the expected
  table padded one column a character wider than the rendered output; the values
  were already right.
- spice-substrait-compliance reports DATAFUSION_FORK_REV, which still named the
  pre-upgrade revision; it now matches the workspace pin, as its test requires.

* test(oracle): re-record the pushdown plan for DataFusion 55's decimal literal display

DataFusion 55 prints a decimal literal as Decimal128(123.45,10,2) where 54 printed
Decimal128(Some(12345),10,2). The SQL sent to Oracle is unchanged.

* build: repin the forks to their merged fix PRs and carry spiceai-54 forward

- DataFusion f22d1a74: spiceai-54 forward-merged into the 55 line
  (spiceai/datafusion#241), bringing spiceai#237 (a Date32 literal's cast uses the
  dialect's date type; SQLite read CAST('…' AS DATE) as a number, so date ranges
  pushed to SQLite matched nothing) and metadata-column listing pruning.
- arrow-rs be4918e7, Ballista 7c4d54c2, table-providers 3045960a, iceberg 50e982dc,
  arrow-adbc 674d6184: the merged CI fix PRs.
- Guard for spiceai#237: a Date32 range unparsed for SQLite keeps the rows it selects
  (runs the SQL on SQLite with --features sqlite); it fails on the previous pin.
- Ledger rows for both; the substrait harness reports the new DataFusion pin.

* build: pin DataFusion at ae719330, the merge of spiceai/datafusion#241

Same tree as f22d1a74 (spiceai-54 forward-merged into the 55 line); now the head
of spiceai-55-patches-2 rather than a temporary branch.

* style: rustfmt the SQLite date-range federation test

* build: pin arrow-rs at its spiceai-59 merge and DataFusion at spiceai-55-patches-3

- arrow-rs: 2e2cc330, the merge of spiceai/arrow-rs#29 into spiceai-59 (same tree as the previous pin).
- DataFusion: fc84f2d1 on spiceai-55-patches-3, which carries the Date32 literal fix (spiceai/datafusion#237) but not the metadata-column listing pruning (spiceai#229); the ledger records that pruning as not carried on 55 until it has a guard here.
- Correct the spark-connect-rs branch annotation to spiceai-59-patches-2.

* build: pin DataFusion at the spiceai-55 merge, and iceberg-rust and snowflake-rs at their arrow-rs bumps

- DataFusion: 02550cf9, the merge of spiceai/datafusion#243 into spiceai-55 (same tree as the previous pin); no longer a temporary pin.
- iceberg-rust: 0284be3d and snowflake-rs: b95a74e8, which move each fork's own arrow-rs pin to 2e2cc330 (Cargo manifests only; Spice builds the same code).

* build: repin ballista, table-providers, iceberg-rust and spark-connect-rs to their branch heads

- datafusion-ballista 557ae4ea: arrow-rs/DataFusion landed pins, and the Rust 1.99 clippy fixes (spiceai#70, spiceai#71, spiceai#72).
- datafusion-table-providers 5f50cfd5 and iceberg-rust a0bd6b04: arrow-rs/DataFusion landed pins (spiceai#82, spiceai#53).
- spark-connect-rs 3bf6421a: regenerated lockfile.

datafusion-federation and duckdb-rs stay at 9ca84a39 and 8ee43073: table-providers still pins exactly those revisions, and a second revision of either would put two copies of the crate in the graph. Their newer heads only change their own Cargo manifests and lockfile.

* build: pin iceberg-rust, snowflake-rs, spark-connect-rs, spice-rs and spicebench at their merges

The fork PRs these pins pointed at have merged, mostly as squash merges, so
several pinned commits are no longer on their base branches. Each pin now
points at the merge on the fork's base branch:

- iceberg-rust spiceai-0.10.1-df-55 @ 3e2a14e (spiceai/iceberg-rust#50, spiceai#54)
- snowflake-rs spiceai-59 @ f555738 (spiceai/snowflake-rs#13, spiceai#15)
- spark-connect-rs spiceai-59-2 @ 18ae9bd (spiceai/spark-connect-rs#15, spiceai#16)
- spice-rs trunk @ 4429f39 (spiceai/spice-rs#103)
- spicebench trunk @ 8a90555 (spiceai/spicebench#284)

datafusion-federation stays at 9ca84a3, which is on spiceai-55 because
spiceai/datafusion-federation#88 was a merge commit; only its branch note
changes.

What changes for Spice: lint-only rewrites in snowflake-api, and tests in
iceberg-datafusion and spark-connect-core. The other differences are in
the forks' own [patch.crates-io] sections, which cargo ignores for a
dependency. spice-rs's tree is identical.

* build: pin the forks at their landed version-branch revisions

- arrow-adbc 6e4119ac (spiceai-24, spiceai#7), iceberg-rust 3e2a14e8 (spiceai-0.10.1-df-55, spiceai#50),
  datafusion-table-providers 465926a3 (spiceai-55, spiceai#80), snowflake-rs f5557381 (spiceai-59, spiceai#13),
  spark-connect-rs 18ae9bd3 (spiceai-59-2, spiceai#15), spice-rs 4429f395 and spicebench 8a90555d (trunk).
- datafusion-federation stays at 9ca84a39, now on spiceai-55 (spiceai#88): datafusion-table-providers pins
  exactly that revision.
- duckdb-rs stays at 8ee43073 and temporary: spiceai#49 was squash-merged, so that revision is not on
  spiceai-1.4.4, and datafusion-table-providers pins it, so both have to move together.
- datafusion-ballista is unchanged; its version PR is still open.

* build: pin datafusion-ballista at upstream 55.0.0-rc1

spiceai/datafusion-ballista#73 merged upstream's 55.0.0-rc1 tag into
spiceai-55-patches-2, the head of spiceai/datafusion-ballista#68. spiceai#68 is
still open, waiting on license-compliance decisions. Pin 718429f, the spiceai#73
merge, until spiceai#68 lands on spiceai-55.

ballista-core, -executor, -scheduler, -history and -api-types move from
54.0.0 to 55.0.0; no packages are added or removed in Cargo.lock.
`cargo check -p spiced --locked` passes on rustc 1.98.1.

* docs(fork_patches): record the datafusion-ballista pin at 718429f4 (upstream 55.0.0-rc1, spiceai#73)

* build: pin datafusion-ballista at a7c4c585, the merge of spiceai/datafusion-ballista#68 into spiceai-55

* build: pin datafusion-ballista at the spiceai-55 merge of spiceai#68

spiceai/datafusion-ballista#68 merged into spiceai-55 as a7c4c58, a merge
commit whose tree is identical to 718429f, the previous pin
(`git diff 718429f a7c4c58` is empty). Repoint the pin at the landed
revision and drop the TEMPORARY marker from the fork-patch ledger, as
scripts/check_fork_patches.py asks. The ledger's ballista section now
names spiceai-55 and upstream 55.0.0-rc1 as its base.

Cargo.lock changes only the source revision of the five ballista packages.

* perf(cayenne): keep small Vortex scans whole on the main query path

DataFusion 55 lowered `repartition_file_min_size` from 10 MiB to 1 MiB
(apache/datafusion#22439), so the main Cayenne scan byte-range-splits a 1-10 MiB
dimension table `target_partitions` ways and pays a Vortex footer open per range.
The main scan now opts small scans out below 10 MiB, as the internal and
protected-snapshot scans already do below 256 MiB; Parquet listing scans keep the
new default. `small_snapshot_groups_opt_out_of_repartitioning` now asserts the
main scan's opt-out and fails with the threshold back at 0.

* docs(fork_patches): record duckdb-rs on the protected spiceai-1.4.4-patches-2 head, which table-providers pins

* test(duckdb): build CreateExternalTable with DataFusion 55's locations field

* ci: validate the datafusion-table-providers pin against spiceai-55, where spiceai#80 landed

* test(cayenne): size the point-lookup fan-out control above the main scan's 10 MiB opt-out

* fix(data-accelerator-api): implement DataFusion 55's apply_expressions for KeepFirstExec

* test: port trunk's new tests to DataFusion 55 (CreateExternalTable locations, common::TableReference)

* fix(cayenne): keep -0.0 and 0.0 equal in pushed-down float comparisons and memory-mode indexes

DataFusion 55 compares floats with the two zeros equal (apache/datafusion#22835),
while Vortex and the Arrow row encoding order them apart. Spell out both zeros
when a float comparison or IN list with a zero literal is pushed into a Vortex
scan, keep float column-vs-column comparisons above the scan, and hash
memory-mode index keys with the file-mode KeyEncoder.

* build: move to iceberg-rust 0.11.0 (RC) and opendal 0.58

Points the iceberg crates at spiceai/iceberg-rust#55, which merges the
upstream 0.11.x release branch onto the DataFusion 55 fork line.

opendal moves to 0.58 to match iceberg-storage-opendal, so one opendal
stack is built instead of two. 0.58 removes the OperatorBuilder step:
Operator::new returns the operator with the same default layers that
finish() applied.

* Bump datafusion to 55.2-rc1

* build: pin iceberg-rust at spiceai-0.11.0-df-55 (6317d03e)

spiceai/iceberg-rust#55 merged into spiceai-0.11.0-df-55. Same Rust code as
the previous pin (e02547cb); the branch adds only a Python CI pin and a
doc-comment fix. fork_patches.md records the new branch and that SigV4 is
now an AuthManager in the fork's REST catalog.

* build: keep the lockfile's unrelated resolutions from the DataFusion bump

The previous commit's lockfile update also re-resolved five crates with
range requirements (hashbrown, heck, base64). Restore them so the iceberg
repin changes only the iceberg-rust source lines.

* build: pin iceberg-rust at spiceai-0.11.0-df-55 (7735caa2)

Picks up spiceai/iceberg-rust#49 (IcebergTableProvider::catalog and
::table_ident) and spiceai#56 (the fork's own DataFusion pin, which does not
affect this build). spiceai#49's fork_patches row and its compile guard land
with spiceai#14588, which first calls the getters.

* build: pin datafusion at spiceai/datafusion#250 (55.2.0-rc1 and the 54 backports)

Moves the DataFusion fork from 02550cf9 (55.1.0) to 1749ce4e, the head of
spiceai/datafusion#250, which merges into spiceai#249: `spiceai-55` at 55.2.0-rc1
(spiceai/datafusion#248) plus the `spiceai-54` patches spiceai#233, spiceai#234 and spiceai#245
carried onto 55, and spiceai#250's review fix. The guards for those follow in the
next commit.

The pin names a pull request branch, so the ledger row is marked TEMPORARY
and `scripts/check_fork_patches.py` refuses to let it land until spiceai#250 and
spiceai#249 merge and the pin moves to the `spiceai-55` merge commit.

The 55.2.0-rc1 merge changes no Spice patch. Outside `Cargo.lock`, the
fork's 02550cf9..27351ab0 diff has the same patch-id as upstream's
55.1.0..55.2.0-rc1 diff, except `datafusion-spark`'s `quote.rs`, which the
fork's backport had already made identical to 55.2.0-rc1's. That backport's
ledger row is dropped; upstream carries the fix as apache/datafusion#25277.
`Cargo.lock` changes only the 36 DataFusion entries.

* test: guard the DataFusion unparser and footer-cache fixes backported from 54

Each backported fix gets a guard here, so the next re-cut of the fork
cannot drop it silently, plus a ledger row naming the guard.

- spiceai/datafusion#233 (spiceai#14373): a join that is another join's right
  input stays a parenthesised joined table on that join's right.
  `a_join_that_is_another_joins_right_input_stays_on_its_right` checks
  that `b` is introduced before the outer `ON` names it and that a LEFT
  JOIN keeps a filter from inside it out of `WHERE`. With the `sqlite`
  feature it runs the SQL on SQLite and compares the rows with
  DataFusion's for the same plan.
- spiceai/datafusion#234 (spiceai#14375): a `Limit` that is a join input bounds
  that input, not the join.
  `a_limit_on_a_join_input_bounds_that_input_rather_than_the_join`, with
  the same row comparison.
- spiceai/datafusion#250: a projection over a join that passes up a mark
  join's mark is refused rather than unparsed with the mark unbound.
  `a_mark_a_projection_over_a_join_passes_up_is_refused`.
- spiceai/datafusion#245 (spiceai#12952): the file metadata cache frees an
  evicted footer's hit counter. 55's `DefaultCache` already does, so 55
  carries no patch; `crates/cayenne/tests/footer_cache_hit_counter_test.rs`
  pins the behaviour by counting the live heap across 20,000 cold footer
  reads into a full cache.

At 02550cf9 the three unparser guards fail. SQLite rejects the nested
join's SQL ("ON clause references tables to its right"), returns 1 row
where the plan returns 2 for `b FULL JOIN (c LIMIT 1)`, and the mark join
renders as `… LEFT OUTER JOIN (b) ON a.id = b.id WHERE NOT c.mark`.
Against a local DataFusion build with the hit-counter prune removed, the
footer-cache guard measures 2,646,176 bytes of heap growth. At 1749ce4e all
four pass, and the footer-cache growth is 0 bytes.

* build: pin datafusion at spiceai/datafusion#249's merge commit

Moves the fork pin from spiceai#250's branch head to bebc4d58, the merge of spiceai#249
into `spiceai-55`. Its tree is spiceai#249's head, cbda233a: 55.2.0-rc1 plus the
backported spiceai#233, spiceai#234 and spiceai#245. The pin names a landed revision, so the
TEMPORARY marker comes off and `scripts/check_fork_patches.py` passes.

spiceai#250 is not part of this revision, so its guard and ledger row are removed.

* Update datafusion version to 55.2 in Cargo.toml

* docs(fork_patches): record the unparser frame refactor spiceai#249 carries

`cbda233a6` is the fourth Spice commit at the pinned revision. It moves the
join-input bookkeeping that spiceai#233 and spiceai#234 add out of
`select_to_sql_recursively_inner`, so an unoptimised build's frame stays
where it was. It had no ledger row, so a re-cut could drop it without the
audit noticing.

Its row is a GAP, with an Open gaps entry. What it buys is stack headroom
in a debug build: without it, the fork's `roundtrip_statement` needs
2,112 KiB and overflows a 2 MiB test thread. This repo runs tests on 8 MiB
threads, so no test here sees the difference deterministically. The entry
names the fork-side check to run at a re-cut instead.

* Fix tests

* build: pin DataFusion at spiceai/datafusion#251, so MAX of a partition column skips empty partitions (spiceai#14801)

* build: pin datafusion at spiceai/datafusion#251's merge commit

Moves the DataFusion fork pin from spiceai#249's merge commit to the head of
`spiceai-55`, which adds two patches on top of it:

- spiceai#251: a file that holds no rows gets no partition-column min/max. Without
  it, `MIN`/`MAX` of a partition column is answered from the listing's
  statistics with the value of a partition whose files are all empty:
  `max(p)` answers '99' for a `p=99` directory holding one empty file, where
  the rows give '3'.
- spiceai#240: the listing prunes files by metadata-column predicates before
  opening them and reports those filters `Exact` (spiceai#229 on `spiceai-54`).

Neither changes a manifest, so `Cargo.lock` only moves its source revision.
Each patch gets a ledger row and a guard in the listing connector's tests,
which the gate runs as library tests.

* fix(listing): keep every predicate residual on the `_location` fast path

The pin now carries spiceai/datafusion#240, so the inner listing reports a
metadata-column predicate such as `_size < 50` `Exact`.
`MetadataPruningListingTable` handed pushdown classification to the inner
listing whenever a `_location` predicate was present, but its `_location`
fast path opens the named objects and applies no other predicate, so
`_location = x AND _size < 50` returned the rows of a file the `_size`
predicate excludes.

Report every predicate `Inexact` on that path as well, so a residual
`FilterExec` re-applies it. This is spiceai#14790's change, which makes the same fix
for partition predicates, taken byte for byte so that whichever of the two
lands second merges cleanly. The new test goes through
`create_listing_table` with `_location` and `_size` enabled.

* test(metadata): expect the listing to answer `_size = 2319` exactly

With spiceai/datafusion#240 the listing prunes files by a metadata-column
predicate and reports it `Exact`, so the plan for `met_size`, which enables
only `_size`, has no residual `Filter`. The scan lists only the four 2319-byte
files; `data_0.parquet` is 2317 bytes, as its S3 `Content-Length` shows.

Regenerated by running `metadata::s3_metadata_columns` and accepted with
`cargo insta accept`. The rest of the test passes unchanged at this pin,
including the `_location` plans and the returned rows.

* fix(data-connector-api): satisfy pedantic clippy in the metadata-pruning listing test

`make lint-rust`'s `--tests` pass rejected two lints in
`listing_table_prunes_files_by_a_metadata_column_predicate`:
`redundant_closure_for_method_calls` (`|group| group.len()`) and
`format_collect` (`map(format!).collect::<String>()`). Use
`FileGroup::len` and build the file contents with `writeln!` into one
`String`. The test is unchanged in what it checks.

---------

Co-authored-by: Luke Kim <80174+lukekim@users.noreply.github.com>
Co-authored-by: Sergei Grebnov <sergei.grebnov@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.