Repository navigation
Merge upstream DataFusion 55.1.0 into spiceai-55 - #243
Merged
Merged
Conversation
…che#22980) ## Which issue does this PR close? * Part of apache#20835 ## Rationale for this change `FixedSizeList` containing `Struct` values was not handled by the existing recursive nested adaptation logic used for schema evolution. As a result, planner-time compatibility checks, nested cast detection, and runtime casting did not support additive struct evolution within `FixedSizeList` containers. This change adds `FixedSizeList` support and verifies planner/runtime parity so that planning allows exactly the cases runtime can adapt while continuing to reject incompatible schema changes. ## What changes are included in this PR? * Extend `cast_column` to support recursive casting of `FixedSizeList` values when source and target list sizes match. * Add `FixedSizeList` handling to: * `requires_nested_struct_cast` * `validate_data_type_compatibility` * Implement recursive casting of nested `Struct` values contained in `FixedSizeList`. * Preserve planner/runtime parity by validating child type compatibility before runtime fallback logic is applied. * Add handling for null-parent `FixedSizeList` entries by masking hidden child values before retrying casts, avoiding failures caused by semantically inaccessible child data. * Refactor list and list-view casting helpers to use Arrow `AsArray` accessors. ## Are these changes tested? Yes. The following tests were added: * `test_cast_fixed_size_list_struct` * `test_validate_fixed_size_list_struct_compatibility` * `test_validate_fixed_size_list_struct_missing_non_nullable_field_rejected` * `test_validate_fixed_size_list_struct_size_mismatch_rejected` * `test_cast_fixed_size_list_struct_all_null` * `test_fixed_size_list_struct_planner_runtime_parity_on_incompatible_type` * `test_cast_fixed_size_list_struct_missing_non_nullable_field_runtime_rejected` * `test_cast_fixed_size_list_struct_ignores_hidden_child_values_for_null_parent` Existing coverage in `test_requires_nested_struct_cast` was also extended to include `FixedSizeList` cases. These tests cover: * Additive nullable nested-field evolution * All-null and partially null list cases * Incompatible nested type changes * Non-nullable field addition rejection * Planner/runtime parity validation ## Are there any user-facing changes? No user-facing changes. This is an internal enhancement to nested schema adaptation and casting behavior for `FixedSizeList<Struct>` types. ## LLM-generated code disclosure This PR includes LLM-generated code and comments. All LLM-generated content has been manually reviewed. --------- Co-authored-by: Andrew Lamb <andrew@nerdnetworks.org>
… spark (apache#23766) ## Which issue does this PR close? N/A ## Rationale for this change Hex encoding was implemented six times across the workspace, at two different levels of optimization. Three copies used a fast byte-pair lookup table: - `datafusion/spark/src/function/math/hex.rs` - `datafusion/functions/src/string/to_hex.rs` - `datafusion/functions/src/encoding/inner.rs` (via the `hex` crate) The other three used a slower nibble-at-a-time loop pushing one character at a time: - `datafusion/functions/src/crypto/md5.rs` - `datafusion/spark/src/function/hash/sha1.rs` - `datafusion/spark/src/function/hash/sha2.rs` Beyond the duplication, this split meant the digest functions were paying for a slower encoder than the one already sitting elsewhere in the tree. Consolidating on a single implementation removes the duplication and moves `md5`, `sha1`, and `sha2` onto the fast path. ## What changes are included in this PR? A new `datafusion_common::utils::hex` module holds the only hex encoder in the workspace: ```rust pub enum HexCase { Lower, Upper } pub fn encode_bytes_into(bytes: &[u8], case: HexCase, out: &mut Vec<u8>); pub fn encode_bytes(bytes: &[u8], case: HexCase) -> String; pub fn encode_bytes_to_slice(bytes: &[u8], case: HexCase, out: &mut [u8]); pub fn encode_u64(v: u64, case: HexCase, buf: &mut [u8; 16]) -> &[u8]; ``` Four entry points rather than one because the call sites have genuinely different output needs, and forcing them through a single shape would cost an allocation somewhere: `to_hex` writes straight into a `StringArray` values buffer, Spark's `hex` appends to a reused scratch `Vec`, `encode` writes into a pre-sized slice, and the digest functions want an owned `String`. Migrated call sites, all producing bit-identical output: | File | Change | | --- | --- | | `functions/src/string/to_hex.rs` | local table and two write helpers deleted; eight trait impls collapsed to two macros | | `spark/src/function/math/hex.rs` | two nibble tables, two lookup tables, `build_hex_lookup` and `hex_int64` deleted | | `functions/src/crypto/md5.rs` | local table and `hex_encode` deleted | | `spark/src/function/hash/sha1.rs` | local table and inline nibble loop deleted | | `spark/src/function/hash/sha2.rs` | local table and `hex_encode` deleted; eight call sites migrated | | `functions/src/encoding/inner.rs` | both encode sites moved off the `hex` crate | Deliberately left alone: - `hex::decode` / `hex::decode_to_slice` in `encoding/inner.rs`, and Spark's `unhex`. The decode direction has different semantics — Spark's `unhex` left-pads odd-length input, the `hex` crate does not — so unifying it is a separate question. The `hex` dependency stays for those. - `ScalarValue`'s binary `Display` impl in `common/src/scalar/mod.rs`, which writes to a `fmt::Formatter` rather than a byte buffer, and is a cold path. Three details worth a reviewer's attention: - The two ancestors of `encode_u64` disagreed on zero: `to_hex` wrote `'0'` into the caller's buffer, Spark's `hex_int64` returned a `'static` `b"0"` that never touched it. The shared version always writes into the buffer and returns a subslice of it, so the lifetime is uniform. Both callers still produce `"0"`. - Spark's `hex_encode_bytes` guards large binary input with `checked_mul(2)` + `try_reserve`, returning a `DataFusionError` rather than aborting on allocation failure. That guard stays at the call site; `encode_bytes_into` performs no reservation of its own, so behaviour is unchanged. - The encoders are `#[inline(always)]`, not `#[inline]`. They are called once per row from two other crates, and plain `#[inline]` left them out-of-line across the crate boundary. Measured cost of that: +3% on Spark's byte paths and up to +18% on `to_hex`'s i32 path, i.e. the refactor was a net regression on those benchmarks until the attribute changed. ## Are these changes tested? Every migrated function keeps its existing unit and sqllogictest coverage, which is what pins bit-identical output — in particular `spark/hash/sha1.slt` and `sha2.slt` assert concrete digest strings for all four SHA-2 bit lengths across both the scalar and array paths, and `expr.slt` covers `to_hex` and `md5`. New unit tests in `common/src/utils/hex.rs` cover zero, `u64::MAX`, single-nibble values, the odd/even digit-count boundary, two's complement of negative input, empty input, all 256 byte values in both cases, appending into a non-empty buffer, and that a reused scratch buffer never leaks stale digits between calls. Tests cross-check against `format!("{:x}")` rather than restating the implementation. Two Spark tests changed. `test_hex_int64` now drives `hex_encode_int64` instead of the deleted private `hex_int64`, keeping all ten cases including `i64::MIN`, `-1`, and the uppercase expectations. `test_hex_lookup_table_covers_all_bytes` was deleted — it only cross-checked the raw lookup tables, and `encode_bytes_covers_every_byte_value` now does that exhaustively through the public API. A test was added for Spark's lowercase byte path, which previously had no coverage. ### Benchmarks Criterion, `apache/main` @ `eef101769` as baseline. Median of the reported change interval. `datafusion/functions/benches/to_hex.rs`: | Benchmark | 1024 | 4096 | 8192 | | --- | --- | --- | --- | | `i32_random` | −13.8% | −12.8% | −10.2% | | `i64_random` | −8.8% | −9.5% | −8.1% | | `i64_large_values` | −9.9% | −9.8% | −7.9% | `scalar_i32` −2.3%, `scalar_i64` −1.8%. `datafusion/spark/benches/sha2.rs`: | Benchmark | 1024 | 4096 | 8192 | | --- | --- | --- | --- | | `array_binary_256` | −20.7% | −18.2% | −19.1% | | `array_scalar_binary_256` | −14.6% | −13.2% | −13.1% | `scalar/size=1` −3.3%. `datafusion/spark/benches/hex.rs`: | Benchmark | 1024 | 4096 | 8192 | | --- | --- | --- | --- | | `hex_int64` | −2.9% | −3.6% | −3.5% | | `hex_int64_dict` | −3.2% | −1.5% | −0.2% | | `hex_utf8` | −0.2% | −0.3% | +0.7% | | `hex_binary` | −0.4% | −0.6% | −0.3% | The `hex_utf8` and `hex_binary` paths already used the byte-pair table before this PR, so they are expected to be flat; they are. `datafusion/functions/benches/crypto.rs`: `md5_array` −4.0%, `md5_scalar` −3.7%. The `sha224` / `sha256` / `sha384` / `sha512` cases in that file range from −0.1% to +2.2%, but they exercise `crypto/basic.rs`, which this PR does not modify and which contains no hex encoding — those numbers are run-to-run variance, not an effect of this change. No number is quoted for Spark `sha1`: it has no benchmark, and its change is the same substitution applied to `md5` and `sha2`. ## Are there any user-facing changes? No behaviour change — all migrated functions produce byte-identical output. `datafusion_common::utils::hex` is new public API on `datafusion-common`. --------- Co-authored-by: Andrew Lamb <andrew@nerdnetworks.org> Co-authored-by: Jeffrey Vo <jeffrey.vo.australia@gmail.com>
…t_batch` memory (apache#23873) ## Which issue does this PR close? - related to apache#23716 ## Rationale for this change It adds test coverage for two gaps found while reviewing apache#23716. ## What changes are included in this PR? Tests only, no functional change. - Dictionary inputs - memory usage on retract (make sure memory is released) Note I moved `array_agg` cases out `aggregate.slt` as it is already more than 9k lines long ## Are these changes tested? They are only tests ## Are there any user-facing changes? No. Tests only, no public API changes. --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
…3881) ## Which issue does this PR close? N/A ## Rationale for this change Two Spark functions allocated a `String` per row purely to render a small, bounded amount of text: - `bin` called `format!("{value:b}")` for every row. - `char` called `ch.to_string()` for every row — a heap allocation for a single character. In both cases the output has a known upper bound (64 binary digits for an `i64`, 4 bytes for a UTF-8 character), so the rendering fits in a stack buffer and the result can be appended straight to a pre-sized builder. ## What changes are included in this PR? `math/bin.rs`: - `spark_bin` now writes digits right-aligned into a caller-supplied `[u8; 64]` and returns a `&str` borrowed from it, instead of returning an owned `String`. - The `collect::<StringArray>()` becomes an explicit loop over a `StringBuilder::with_capacity`, sized at 8 digits per row. - Negative values still render as their two's-complement bit pattern, matching `{:b}`. The digit loop is a `loop`, not a `while`, so zero renders as `"0"` rather than the empty string. `string/char.rs`: - `ch.to_string()` becomes `ch.encode_utf8(&mut encoded)` against a `[u8; 4]` hoisted out of the loop. Output is unchanged in both cases. ## Are these changes tested? Existing coverage pins the behaviour. `spark/math/bin.slt` asserts concrete output for the cases the rewrite had to get right: zero, negative values, `i64::MIN` (`-9223372036854775808`), `i64::MAX`, and `-2147483648` / `-32768` widened from narrower integer types. `spark/string/char.slt` covers the negative-input empty string, the null path, and characters on both sides of the ASCII boundary (`char(256)` and above wrap via `% 256`). All 119 `spark/math` and `spark/string` sqllogictest files pass, along with the 258 `datafusion-spark` unit tests. The `bin` benchmark used below, `datafusion/spark/benches/bin.rs`, is added separately in apache#23882 so the baseline can be measured on `main` before this change lands. It covers 1024 and 8192 rows with 20% nulls over two value distributions: small values that render to a handful of digits, and full-range values that render to the maximum 64. `char` already had `datafusion/spark/benches/char.rs` on `main`, so no benchmark change is needed for it. ### Benchmarks Criterion, `apache/main` @ `f1ab86dad` as baseline. Median of the reported change interval. | Benchmark | 1024 | 8192 | | --- | --- | --- | | `bin/small` | −75.7% | −73.7% | | `bin/wide` | −47.8% | −49.8% | `char` (1024 rows): −76.1%. ## Are there any user-facing changes? No. Both functions produce byte-identical output; this is purely an allocation change.
…ts enabled (apache#23848) ## Which issue does this PR close? <!-- We generally require a GitHub issue to be filed for all bug fixes and enhancements and this helps us generate change logs for our releases. You can link an issue to this PR using the GitHub syntax. For example `Closes #123` indicates that this PR will close issue #123. --> - Closes apache#23847 . ## Rationale for this change Fixes a correctness bug, not sure when that config used/desirable. This is another thing I ran into during apache#21585 ## What changes are included in this PR? 1. New SLT test 2. Fix for bug in `datafusion-optimizer` ## Are these changes tested? Existing tests, and a new SLT test verifying the config change doesn't affect the query result and basic unit test. ## Are there any user-facing changes? None --------- Signed-off-by: Adam Gutglick <adamgsal@gmail.com>
…lter (apache#23901) ## Which issue does this PR close? - Closes apache#23900. ## Rationale for this change `push_down_filter` infers equi-key predicates across a join's ON keys and pushes them to the opposite side. For a null-aware join (the `LeftAnti` join produced by `NOT IN` with a nullable subquery), an outer predicate on the left key like `outer.id > 5` is rewritten to `sub.id > 5` and pushed onto the subquery input. Since the inferred predicate must be null-rejecting to be pushed, this drops the subquery's NULL rows and breaks the three-valued `NOT IN` semantics — a NULL in the subquery key must reach the join so the result is empty. Same class of bug as apache#23848, in a different rule. ## What changes are included in this PR? - Skip predicate inference in `infer_join_predicates` when `join.null_aware` is set (mirrors the apache#23848 guard on `FilterNullJoinKeys`). - A `push_down_filter` unit test asserting no predicate is inferred onto the subquery side of a null-aware `LeftAnti` join. - SLT coverage for the failing query, plus a `prefer_hash_join = false` / multi-partition variant. ## Are these changes tested? Yes — new unit test (verified it fails without the guard) and SLT cases. The full optimizer lib suite passes. ## Are there any user-facing changes? No, aside from the correctness fix.
## Which issue does this PR close? - Part of apache#19241. - Stacked on [apache#23311](apache#23311). - Next in stack: apache#23015. - Extracted from apache#19390. ## Rationale for this change For very small `IN` lists, building or probing a hash table can be more work than just comparing the input value with each constant. For example, for `x IN (10, 20, 30)`, the fast path can behave like: ```text x == 10 OR x == 20 OR x == 30 ``` Because the list is tiny, those comparisons are cheap. The implementation stores the constants in a fixed-size array and checks them with a compact comparison chain. “Branchless” here means the comparisons are combined without stopping at the first match. That can be faster for these small fixed-width lists because the CPU gets a predictable sequence of simple operations instead of hash-table setup and probe logic. For primitive values that are not already plain unsigned integers, this PR keeps the logical Arrow type explicit and uses a matching same-width comparison representation only inside the branchless filter. For example, `Float16` uses `UInt16` storage, `Float32` uses `UInt32` storage, and `TimestampNanosecond` uses `UInt64` storage. `Decimal128` and `IntervalMonthDayNano` use their own 16-byte native representation. This preserves bit-pattern equality while relying on Arrow's native primitive compatibility rules: timestamp timezone metadata and Decimal128 precision/scale metadata may differ, while incompatible primitive representations remain rejected. ## What changes are included in this PR? - Adds a const-generic `BranchlessFilter` for small primitive `IN` lists. - Adds thresholds for when this path is used: - up to 16 values for 1-byte types - up to 8 values for 2-byte types - up to 32 values for 4-byte types - up to 16 values for 8-byte types - up to 4 values for 16-byte types - Keeps dispatch concrete and explicit in `strategy.rs`. - Maps each optimized logical type to the comparison representation used by the branchless filter: - `Int8` -> `UInt8` - `Int16`, `Float16` -> `UInt16` - `Int32`, `Float32`, `Date32`, `Time32` -> `UInt32` - `Int64`, `Float64`, `Date64`, `Time64`, `Timestamp`, `Duration` -> `UInt64` - `Decimal128`, `IntervalMonthDayNano` -> their native 16-byte representation - Leaves larger 1-byte and 2-byte lists on the existing bitmap filters. - Leaves larger 4-byte and 8-byte lists on the existing hash/generic paths. - Leaves wider primitive types such as `Decimal256` and unsupported complex types on the generic path. - Keeps the same `IN` / `NOT IN` null behavior as the rest of the stack. - Adds focused coverage for branchless null handling, signed boundary values, slices, Float16/Float32/Float64 bit patterns, compatible timestamp/Decimal128 metadata, incompatible timestamp units, IntervalMonthDayNano values, and same-width wrong-type probe rejection. ## Are these changes tested? Yes. - `cargo fmt --all -- --check` - `cargo test -p datafusion-physical-expr expressions::in_list --lib` - `cargo test -p datafusion-physical-expr --bench in_list_strategy --no-run` - `cargo clippy --all-targets --all-features -- -D warnings` ## Are there any user-facing changes? No. This is an internal performance optimization only. ## Local benchmark snapshot Built and run with `release-nonlto`, filtered to the relevant small primitive-list rows: ```bash cargo bench -p datafusion-physical-expr --profile release-nonlto --bench in_list_strategy -- <filter> --save-baseline <baseline> ``` Filters used: `narrow_integer`, `primitive/i32/small_list`, `primitive/i64/small_list`, `f32/small_list`, `timestamp_ns/small_list`, and `interval_month_day_nano/small_list`. Method: directly compared Criterion's raw sample minima (`min(time / iterations)`) from `sample.json`. Lower is better; changes within +/-5% are treated as noise. Compared baselines: [apache#23311](apache#23311) -> [apache#23014](apache#23014) Relevant scope: small primitive-list rows. Summary: 39 relevant rows, 28 faster, 0 slower, 11 within +/-5%. Largest relevant deltas: | Benchmark | Before | After | Change | |---|---:|---:|---:| | `timestamp_ns/small_list/list=4/match=50%` | 46.55 us | 3.17 us | -93.2% (14.69x faster) | | `f32/small_list/list=4/match=50%` | 33.93 us | 3.04 us | -91.0% (11.15x faster) | | `primitive/i32/small_list/list=4/match=50%` | 32.63 us | 3.08 us | -90.5% (10.58x faster) | | `primitive/i64/small_list/list=4/match=50%` | 33.55 us | 3.18 us | -90.5% (10.54x faster) | | `timestamp_ns/small_list/list=4/match=0%` | 19.57 us | 3.18 us | -83.8% (6.16x faster) | | `f32/small_list/list=4/match=0%` | 18.14 us | 3.05 us | -83.2% (5.95x faster) | | `primitive/i32/small_list/list=4/match=0%` | 17.00 us | 3.04 us | -82.1% (5.59x faster) | | `primitive/i64/small_list/list=4/match=0%` | 17.12 us | 3.22 us | -81.2% (5.31x faster) | | `primitive/i32/small_list/list=16/match=50%/NOT_IN` | 31.98 us | 7.26 us | -77.3% (4.41x faster) | | `nulls/primitive/i32/small_list/list=16/match=50%/nulls=20%` | 29.35 us | 7.32 us | -75.1% (4.01x faster) | | `timestamp_ns/small_list/list=16/match=50%` | 45.32 us | 11.79 us | -74.0% (3.84x faster) | | `nulls/primitive/i32/small_list/list=16/match=50%/nulls=50%` | 25.89 us | 7.31 us | -71.8% (3.54x faster) | | `nulls/primitive/i32/small_list/list=16/match=50%/nulls=20%/NOT_IN` | 26.05 us | 7.42 us | -71.5% (3.51x faster) | | `interval_month_day_nano/small_list/list=4/match=50%` | 52.94 us | 15.52 us | -70.7% (3.41x faster) | | `f32/small_list/list=32/match=50%` | 38.78 us | 13.27 us | -65.8% (2.92x faster) | | `primitive/i64/small_list/list=16/match=50%` | 29.46 us | 11.76 us | -60.1% (2.50x faster) | <details> <summary>Full relevant table (39 rows)</summary> | Benchmark | Before | After | Change | |---|---:|---:|---:| | `narrow_integer/u8/list=4/match=0%` | 3.86 us | 2.79 us | -27.8% (1.38x faster) | | `narrow_integer/u8/list=4/match=50%` | 3.84 us | 2.78 us | -27.7% (1.38x faster) | | `narrow_integer/u8/list=16/match=0%` | 3.88 us | 3.85 us | -0.8% (within noise) | | `narrow_integer/u8/list=16/match=50%` | 3.84 us | 3.86 us | +0.5% (within noise) | | `narrow_integer/i16/list=4/match=0%` | 3.93 us | 3.18 us | -19.1% (1.24x faster) | | `narrow_integer/i16/list=4/match=50%` | 3.92 us | 3.16 us | -19.5% (1.24x faster) | | `narrow_integer/i16/list=64/match=0%` | 3.96 us | 3.82 us | -3.5% (within noise) | | `narrow_integer/i16/list=64/match=50%` | 3.91 us | 3.80 us | -2.9% (within noise) | | `narrow_integer/i16/list=256/match=0%` | 3.90 us | 3.81 us | -2.5% (within noise) | | `narrow_integer/i16/list=256/match=50%` | 3.97 us | 3.81 us | -4.1% (within noise) | | `narrow_integer/f16/list=4/match=0%` | 3.87 us | 3.16 us | -18.5% (1.23x faster) | | `narrow_integer/f16/list=4/match=50%` | 3.94 us | 3.15 us | -20.2% (1.25x faster) | | `narrow_integer/f16/list=64/match=0%` | 3.87 us | 3.84 us | -0.6% (within noise) | | `narrow_integer/f16/list=64/match=50%` | 3.93 us | 3.85 us | -1.9% (within noise) | | `narrow_integer/f16/list=256/match=0%` | 3.90 us | 3.84 us | -1.5% (within noise) | | `narrow_integer/f16/list=256/match=50%` | 3.87 us | 3.91 us | +1.2% (within noise) | | `nulls/narrow_integer/u8/list=16/match=50%/nulls=20%` | 3.92 us | 4.02 us | +2.5% (within noise) | | `primitive/i32/small_list/list=4/match=0%` | 17.00 us | 3.04 us | -82.1% (5.59x faster) | | `primitive/i32/small_list/list=4/match=50%` | 32.63 us | 3.08 us | -90.5% (10.58x faster) | | `primitive/i32/small_list/list=32/match=0%` | 16.34 us | 13.33 us | -18.5% (1.23x faster) | | `primitive/i32/small_list/list=32/match=50%` | 31.17 us | 13.31 us | -57.3% (2.34x faster) | | `primitive/i32/small_list/list=16/match=50%/NOT_IN` | 31.98 us | 7.26 us | -77.3% (4.41x faster) | | `nulls/primitive/i32/small_list/list=16/match=50%/nulls=20%` | 29.35 us | 7.32 us | -75.1% (4.01x faster) | | `nulls/primitive/i32/small_list/list=16/match=50%/nulls=20%/NOT_IN` | 26.05 us | 7.42 us | -71.5% (3.51x faster) | | `nulls/primitive/i32/small_list/list=16/match=50%/nulls=50%` | 25.89 us | 7.31 us | -71.8% (3.54x faster) | | `primitive/i64/small_list/list=4/match=0%` | 17.12 us | 3.22 us | -81.2% (5.31x faster) | | `primitive/i64/small_list/list=4/match=50%` | 33.55 us | 3.18 us | -90.5% (10.54x faster) | | `primitive/i64/small_list/list=16/match=0%` | 16.34 us | 11.93 us | -27.0% (1.37x faster) | | `primitive/i64/small_list/list=16/match=50%` | 29.46 us | 11.76 us | -60.1% (2.50x faster) | | `f32/small_list/list=4/match=0%` | 18.14 us | 3.05 us | -83.2% (5.95x faster) | | `f32/small_list/list=4/match=50%` | 33.93 us | 3.04 us | -91.0% (11.15x faster) | | `f32/small_list/list=32/match=0%` | 22.05 us | 13.43 us | -39.1% (1.64x faster) | | `f32/small_list/list=32/match=50%` | 38.78 us | 13.27 us | -65.8% (2.92x faster) | | `timestamp_ns/small_list/list=4/match=0%` | 19.57 us | 3.18 us | -83.8% (6.16x faster) | | `timestamp_ns/small_list/list=4/match=50%` | 46.55 us | 3.17 us | -93.2% (14.69x faster) | | `timestamp_ns/small_list/list=16/match=0%` | 19.73 us | 12.07 us | -38.8% (1.63x faster) | | `timestamp_ns/small_list/list=16/match=50%` | 45.32 us | 11.79 us | -74.0% (3.84x faster) | | `interval_month_day_nano/small_list/list=4/match=0%` | 20.12 us | 13.20 us | -34.4% (1.52x faster) | | `interval_month_day_nano/small_list/list=4/match=50%` | 52.94 us | 15.52 us | -70.7% (3.41x faster) | </details> --------- Co-authored-by: Andrew Lamb <andrew@nerdnetworks.org>
## Which issue does this PR close? - Closes apache#23770 - Part of apache#15914 ## Rationale for this change Spark provides [`hypot(expr1, expr2)`](https://spark.apache.org/docs/latest/api/sql/#hypot), which returns `sqrt(expr1^2 + expr2^2)` computed without intermediate overflow or underflow. It was not yet implemented in `datafusion-spark` — only an auto-generated test stub existed at `spark/math/hypot.slt` with its query commented out. ## What changes are included in this PR? - Add `SparkHypot` (implementing `ScalarUDFImpl`) in `datafusion/spark/src/function/math/hypot.rs`, backed by Rust's `f64::hypot` — the same overflow-safe algorithm as Java/Spark's `Math.hypot`. - Register it in `datafusion/spark/src/function/math/mod.rs`. - Enable the `hypot.slt` sqllogictest. The signature is `exact(Float64, Float64) -> Float64`, following the `datafusion-spark` convention of only accepting types Spark supports. Computation uses the Arrow `binary` kernel so NULL in either argument propagates to a NULL result, matching Spark. ## Are these changes tested? Yes — `datafusion/sqllogictest/test_files/spark/math/hypot.slt` covers: - scalar Pythagorean triples (`hypot(3, 4)` → 5, `hypot(5, 12)` → 13), - double inputs, - NULL propagation when either argument is NULL, - the array path (including a NULL row), - overflow-safety: `hypot(3e200, 4e200)` stays finite, whereas a naive `sqrt(a^2 + b^2)` would overflow to `Infinity`. ## Are there any user-facing changes? Yes — adds the Spark-compatible `hypot` scalar function to `datafusion-spark`. No breaking changes to public APIs.
## Which issue does this PR close? <!-- We generally require a GitHub issue to be filed for all bug fixes and enhancements and this helps us generate change logs for our releases. You can link an issue to this PR using the GitHub syntax. For example `Closes #123` indicates that this PR will close issue #123. --> - Part of apache#23393 . ## Rationale for this change `SlidingMinAccumulator::size` and `SlidingMaxAccumulator::size` only reported the stack size of their `ScalarValue` field plus its heap payload, ignoring the memory held by the underlying `MovingMin` / `MovingMax` sliding-window buffers. For windowed `MIN`/`MAX` over string or list data, the two per-element stacks can hold megabytes of `ScalarValue` payload that the memory pool was never told about, understating accumulator memory usage. ## What changes are included in this PR? - Add a private `heap_size(elem_heap)` method to `MovingMin<T>` and `MovingMax<T>` that reports the two stack buffers' capacity in bytes plus each stored element's heap payload. - Factor the shared implementation into a `moving_stacks_heap_size` free helper so the two types stay in sync. - Include the buffer bytes in `SlidingMinAccumulator::size` and `SlidingMaxAccumulator::size` via the new method. ## Are these changes tested? Yes. Two new unit tests in `datafusion/functions-aggregate/src/min_max.rs`: - `moving_min_max_heap_size_i32` — fixed-width `T`, verifies buffer-only accounting with and without pushed elements. - `moving_min_max_heap_size_counts_elems` — `T = String`, verifies each of the two slots in a `(T, T)` pair contributes independently to the heap payload (mirroring the two independent `Clone`s made by `push`).
## Which issue does this PR close? - Related to apache#21882 (comment). ## Rationale for this change PR apache#21882 adds custom spill file support. Having an example showing how a downstream application can use this new API will help make sure the API is good enough for our needs ## What changes are included in this PR? This PR adds a small example showing how users can back spill files with an `ObjectStore`, using a local object store for a runnable example while keeping the implementation applicable to remote stores. ## Are these changes tested? Yes by CI ## Are there any user-facing changes? Yes. This adds a new example for configuring ObjectStore-backed spill files.
## Which issue does this PR close? <!-- We generally require a GitHub issue to be filed for all bug fixes and enhancements and this helps us generate change logs for our releases. You can link an issue to this PR using the GitHub syntax. For example `Closes #123` indicates that this PR will close issue #123. --> - Closes #N/A. ## Rationale for this change <!-- Why are you proposing this change? If this is already explained clearly in the issue then this section is not needed. Explaining clearly why changes are proposed helps reviewers understand your changes and offer better suggestions for fixes. --> Since DataFusion doesn't typically guartantee wire format compatibility, I remove the backward compatiblity shim that introduced in PR apache#23189 FYI: apache#23189 (comment) ## What changes are included in this PR? <!-- There is no need to duplicate the description in the issue here but it is sometimes worth providing a summary of the individual changes in this PR. --> DataFusion does not guarantee serialized plans across versions. Keeping `partitioned_by_file_group` therefore leaves a dead schema field and decoder path after `output_partitioning` became the source of truth. Reserve the old field number and name to prevent future reuse. ## Are these changes tested? <!-- We typically require tests for all PRs in order to: 1. Prevent the code from being accidentally broken by subsequent changes 2. Serve as another way to document the expected behavior of the code If tests are not included in your PR, please explain why (for example, are they covered by existing tests)? --> Yes ## Are there any user-facing changes? <!-- If there are user-facing changes then we may require documentation to be updated before approving the PR. --> <!-- If there are any breaking changes to public APIs, please add the `api change` label. --> Yes Signed-off-by: Jiawei Zhao <Phoenix500526@163.com> Co-authored-by: Andrew Lamb <andrew@nerdnetworks.org>
…llable build key (apache#23173) ## Which issue does this close? - Closes apache#23126. ## Rationale for this change `x NOT IN (subquery)` plans to a null-aware `LeftAnti` hash join (build = outer `x`, probe = subquery). Join dynamic filter pushdown pushes a bounds + membership filter, built from the build keys, onto the probe scan. That filter can prune every probe row. A null-aware `LeftAnti` reads an empty probe as a genuinely-empty subquery, so it emits build-side NULL rows that should drop: `NULL NOT IN (non-empty)` is UNKNOWN, not TRUE. The result is scan-dependent, so it's a silent correctness bug. A `VALUES` scan ignores the pushed filter and stays correct; a parquet scan applies it and is wrong. apache#23103 (the probe-side NULL drop) is orthogonal; this is the build-side NULL. ## What changes are included in this PR? Skip join dynamic filter pushdown for a null-aware anti join when the build key can be NULL. The build-side NULL emission depends on whether the probe is truly empty, which the pushed filter can change by emptying it. A NOT NULL build key has no such NULL, so it keeps the pushdown. The check is static: a schema-nullable build key disables the pushdown even when the data contains no NULLs. A runtime alternative (keep the pushdown and neutralize the filter only when the build actually holds a NULL key) would restore the optimization for those cases. I'd leave that as a follow-up. ## Are these changes tested? Yes. A `push_down_filter_parquet.slt` case reproduces it (build-side NULL, a non-matching parquet probe) and asserts the single correct row. Without the change it returns the extra NULL. In addition, unit tests pin both directions of the guard: a nullable build key rejects the pushdown and a NOT NULL build key keeps it. ## Are there any user-facing changes? `NOT IN` over a parquet (or otherwise prunable) scan with a nullable outer key now returns correct results. Such joins lose the dynamic filter pushdown.
…cimal) (apache#23631) ## Which issue does this PR close? <!-- We generally require a GitHub issue to be filed for all bug fixes and enhancements and this helps us generate change logs for our releases. You can link an issue to this PR using the GitHub syntax. For example `Closes #123` indicates that this PR will close issue #123. --> N/A ## Rationale for this change <!-- Why are you proposing this change? If this is already explained clearly in the issue then this section is not needed. Explaining clearly why changes are proposed helps reviewers understand your changes and offer better suggestions for fixes. --> I noticed there were some subtle errors with how decimals were handled in scalar values, and also opportunity to remove power calls in favour of precomputed constant tables. Also filling out some other missing support. ## What changes are included in this PR? <!-- There is no need to duplicate the description in the issue here but it is sometimes worth providing a summary of the individual changes in this PR. --> I recommend looking at the commits as they are self contained with detailed messages for each. ## Are these changes tested? <!-- We typically require tests for all PRs in order to: 1. Prevent the code from being accidentally broken by subsequent changes 2. Serve as another way to document the expected behavior of the code If tests are not included in your PR, please explain why (for example, are they covered by existing tests)? --> Yes ## Are there any user-facing changes? <!-- If there are user-facing changes then we may require documentation to be updated before approving the PR. --> No <!-- If there are any breaking changes to public APIs, please add the `api change` label. -->
…lter pushdown enabled (apache#23638) ## Which issue does this PR close? - Closes apache#23531. ## Rationale for this change `input_file_name()` is `FileSource` dependent like `file_row_index()`, and therefore shouldn't be pushed down into a filter. ## What changes are included in this PR? `PushdownChecker` now handles both UDFs consistently. If we keep adding this sort of metadata functions, we might want a better API to detect them, but for now I think this is a reasonable approach that isn't very invasive. ## Are these changes tested? Additional SLT test that verifies that both function behave correctly with pushdown either enabled or disabled. ## Are there any user-facing changes? None --------- Signed-off-by: Adam Gutglick <adamgsal@gmail.com>
…xpr::check_bigger_cast (apache#23808) (apache#23809) ## Which issue does this PR close? - Closes apache#23808. ## Rationale for this change ref. apache#23807 `CastExpr::check_bigger_cast` is used to determine whether a cast is a widening (order-preserving) conversion. Currently, it classifies `Int32 -> Float32`, `UInt32 -> Float32`, `Int64 -> Float64`, and `UInt64 -> Float64` as widening casts. However, integer-to-float conversions for 32-bit and 64-bit integers lose precision when values exceed the mantissa bit limit (24 bits for `Float32`, 53 bits for `Float64`). For example: `16_777_216_i32 as f32 == 16_777_217_i32 as f32` (both yield 16777216.0f32). Because distinct integer inputs can collapse to the same float output, these casts are not strictly 1-to-1 (injective) and can break suffix sort key ordering in multi-column ordering analysis (e.g. `[CAST(a AS Float32), b]`). ## What changes are included in this PR? - Updated `CastExpr::check_bigger_cast` to exclude precision-losing integer-to-float conversions (`Int32/UInt32 -> Float32` and `Int64/UInt64 -> Float64`). - Added unit test `test_check_bigger_cast_precision_loss` to verify precision-losing casts return `false` while exact conversions (`Int16 -> Float32`, `Int32 -> Float64`, etc.) continue to return `true`. ## Are these changes tested? Yes, new unit test `test_check_bigger_cast_precision_loss` in `cast.rs`. ## Are there any user-facing changes? No breaking API changes. Internal optimizer behavior fix.
…ossible (apache#23702) ## Which issue does this PR close? N/A ## Rationale for this change `SortPreservingMergeStream` is a little complex, so add some guiding comments and make it as textbook-like as possible ## What changes are included in this PR? Added comments, reorder code While this was done this also fixed couple of bugs due to how it work: 1. leftover drain was not counted in the `elapsed_compute` 2. `limit(0)` returns 0 rows and not 1 ## Are these changes tested? Existing tests ## Are there any user-facing changes? Not API ones. `limit(0)` now returns 0 rows
## Which issue does this PR close? - Closes apache#21428. ## Rationale for this change This PR adds `BuildHasher`-based variants for `hash_utils` so callers can compute row hashes with a caller-provided hash builder instead of always using DataFusion's default `RandomState`. The main constraint is performance: `with_hashes` is a hot path, especially for string, dictionary, and nested array hashing. A previous version in apache#21429 caused measurable regressions in the default `RandomState` path, for example `large_utf8: single, no nulls` regressed from roughly `26.7us` to `36.3us`, and `large_utf8: multiple, no nulls` from roughly `112us` to `127us`. This version keeps the default path performance-oriented by avoiding a fully generic `BuildHasher` rewrite of the existing hot loops. ## What changes are included in this PR? This PR adds: - `with_hashes_with_hasher` - `create_hashes_with_hasher` - custom-hasher implementations for primitive, string, binary, byte-view, dictionary, and nested arrays - tests covering custom hashers, multi-column hashing, and dictionary equivalence The implementation intentionally uses a hybrid design: - Default `RandomState` leaf hot paths remain specialized. - Custom `BuildHasher` leaf paths live separately in `hash_utils/build_hasher.rs`. - Nested/structural logic is shared through an internal child-hashing adapter, so struct/list/map/union/run/dictionary behavior does not need to be broadly duplicated. The trade-off is that there is still some duplication for primitive/string/binary leaf loops. That duplication is intentional: those are the hottest loops, and keeping them separate prevents the existing `RandomState` path from becoming generic over `BuildHasher` or being perturbed by the custom-hasher implementation. ## Are these changes tested? Yes. --------- Co-authored-by: Dmitrii Blaginin <dmitrii@blaginin.me> Co-authored-by: Andrew Lamb <andrew@nerdnetworks.org>
…t whether they keep the same ordering of the input (apache#23807) ## Which issue does this PR close? - Closes apache#23798 Related to: - apache#16217 ## Rationale for this change To be able to keep the same sorting order allowing for more optimizations Now comet own `cast` implementation can be recognized as not modifying sort order in the same cases that datafusion cast does. ## What changes are included in this PR? added `strictly_order_preserving` property to `ExprProperties` + varius other places and replaced the hard coded logic for cast (`substitute_cast_ordering`) about keeping input order with more generic approach that now any expression can implement and have the same advantage of sort elimination also marked from_unixtime as keeping ordering to show a case of this optimization ## Are these changes tested? yes ## Are there any user-facing changes? yes, breaking change, added `strictly_order_preserving` property to `ExprProperties` and to `FFI_ExprProperties` this property means that given expression `f` and 2 values from the input column `a` and `b` the following variants are kept: 1. `a.cmp(b) == f(a).cmp(f(b))` 2. nulls maps to nulls Example of satisfying expression: `cast(col_a as BIGINT)` where `col_a` is `INT` it is keeping the properties Example of not satisfying: `floor` - floor can not `array_repeat(my_col, 2)` which might look like at first glance as keeping the property as well but in fact it does not. the reason is that `array_repeat(null, 2)` will output list of 2 nulls which breaks the 2nd property that nulls must be kept as nulls how to migrate: Option 1 - keeping the old behavior (safest but least performant) set `strictly_order_preserving` to false Option 2 - Using the new optimization that this opens up: set `strictly_order_preserving` only if the expression keep both variants
## Which issue does this PR close? Part of apache#22330. ## Rationale for this change `ForeignScalarUDF` inherits the default `preserves_lex_ordering`, so producer overrides are lost across the FFI boundary. ## What changes are included in this PR? - Forward `preserves_lex_ordering` through `FFI_ScalarUDF`. - Reuse the existing placement UDF for unit and dynamic-library coverage. ## Are these changes tested? - `cargo test -p datafusion-ffi --features integration-tests` - `cargo clippy --all-targets --all-features -- -D warnings` ## Are there any user-facing changes? The FFI ABI changes. Foreign libraries must rebuild against the new DataFusion version. --------- Signed-off-by: Amogh Ramesh <ramogh2404@gmail.com>
…tonic deques (apache#23826) (apache#23827) ## Which issue does this PR close? - Closes apache#23826 ### What changes are included in this PR? This PR optimizes sliding window `MIN`/`MAX` aggregate functions using a **Sequence-Numbered Monotonic Deque** instead of a Two-Stack Queue. **Revised Design:** - We store `(sequence_number, value)` pairs in a single `VecDeque`, kept in strictly monotonic order. - **`push(val)`**: Evicts dominated elements from the back, then pushes the new value with the current `push_seq` and increments `push_seq`. - **`pop()`**: Increments `pop_seq`. If the front elements sequence number equals the old `pop_seq`, it is expired and popped. - **Benefits**: No secondary FIFO queue (lower memory) and no `clone()` overhead. ### Are these changes tested? Yes, existing tests pass. Added tests for empty-window `pop()` and duplicate-heavy scenarios. ### Are there any user-facing changes? **API Change**: `MovingMin` and `MovingMax` were changed to `pub(crate)` visibility, and their `pop()` methods now return `()` instead of returning a value. Performance is significantly improved (2x-3.5x throughput). --------- Co-authored-by: Pavan <pavan@Lakshmis-MacBook-Air.local>
## Which issue does this PR close? NA ## Rationale for this change Adds [datapress](https://docs.datap-rs.org) to the list of known users in the documentation. ## What changes are included in this PR? NA ## Are these changes tested? NA ## Are there any user-facing changes? NA
## Which issue does this PR close? - Closes apache#11748 ## Rationale for this change `EliminateGroupByConstant` removes GROUP BY expressions that are constants. If all of the GROUP BY expressions are constants, eliminating all of them results in converting a grouped aggregate into an ungrouped (global) aggregate. This changes the semantics of the query: a grouped aggregate query on an empty input returns zero rows, whereas an ungrouped aggregate query produces a single row. ## What changes are included in this PR? `EliminateGroupByConstant` now declines to eliminate constant GROUP BY expressions, if doing so would result in removing all of the grouping expressions. ## Are these changes tested? Yes. Sqllogictest reproducer for the original issue, updated unit test and optimizer SLT expectations. ## Are there any user-facing changes? Queries with all-constant GROUP BY may now return fewer (correct) rows.
…pache#23874) ## Which issue does this PR close? - Closes apache#23872 ## Rationale for this change `SlidingMinAccumulator::update_batch` skipped NULL values. This meant that if all the non-NULL values in a window frame were retracted, the window frame would not be empty but the `MovingMin` data structure by the `SlidingMinAccumulator` would not contain any values. This resulted in incorrectly returning a stale non-NULL value for a sliding `min()` over a window frame consisting of only NULL values. Along the way, optimize the min and max sliding window accumulators to make them both more efficient and more symmetric with one another. In the original coding, `SlidingMaxAccumulator` included NULL values but `SlidingMinAccumulator` omitted them, in part because omitting NULLs made the original `retract_batch` implementation more expensive. This PR optimizes `retract_batch`, so we can now use the same scheme for both the min and max sliding accumulators: * Omit NULLs on `update_batch` (this improves on the prior behavior of `max`) * Efficiently account for NULLs in `retract_batch` (this improves on the prior behavior of `min`) * Ensure correct results for all-NULL window frames (this fixes the prior bug in `min`). ## What changes are included in this PR? * Fix bug in `min()` over all-NULL window frames * Optimize `SlidingMinAccumulator::retract_batch` (avoid materializing values just to count NULLs) * Optimize `SlidingMinAccumulator::update_batch` (omit NULLs), also improving symmetry with `min` * Optimize both accumulators to stop caching the current `min` / `max`; this saves a few clones, but perhaps more importantly it is simpler and avoids the risk of inconsistency between the cached value and the underlying `MovingMin` / `MovingMax` data structure ## Are these changes tested? Yes, new tests added. ## Are there any user-facing changes? No, aside from the bug fix.
## Which issue does this PR close?
- N/A
## Rationale for this change
```
$ cargo test -p datafusion-functions-aggregate
[...]
warning: associated function `from_parts` is never used
--> datafusion/physical-expr/src/expressions/dynamic_filters/mod.rs:464:8
[...]
```
Fix this by appropriately gating the compilation of `from_parts`.
## What changes are included in this PR?
* Squelch unused code warning
## Are these changes tested?
Yes.
## Are there any user-facing changes?
No.
…w retract (apache#23913) ## Which issue does this PR close? - Closes apache#23912. ## Rationale for this change `DistinctPercentileContAccumulator` reused the shared set-based `GenericDistinctBuffer`, which doesn't fit it: (1) the buffer asserts a single input column but `percentile_cont` passes two (value + percentile), so every `percentile_cont(DISTINCT ...)` panicked; (2) the buffer is a plain `HashSet` with no multiplicity, so sliding-window `retract_batch` dropped a value while duplicates were still in the frame. ## What changes are included in this PR? - Replace the shared buffer in this accumulator with a per-accumulator `HashMap<Hashable, usize>` count map: `update_batch` reads only the value column and increments; `retract_batch` decrements and removes a key only at zero; `state`/`merge_batch` keep the same List state shape. Other `GenericDistinctBuffer` users are untouched. - Regression tests in `aggregate.slt` for the plain distinct query and the sliding-window duplicate-retract case. ## Are these changes tested? Yes — new regression tests; the full `aggregate.slt` suite passes. ## Are there any user-facing changes? `percentile_cont(DISTINCT ...)` now works instead of panicking, and returns correct results in sliding windows.
## Which issue does this PR close? - Closes apache#23717. ## Rationale for this change Spark and Java format negative numeric values with parentheses when the `(` flag is present. The decimal formatting path always emitted a minus sign because it ignored `negative_in_parentheses`, while the floating-point path already handled the flag. This made `format_string` inconsistent across numeric input types. ## What changes are included in this PR? - Add a closing suffix when a negative decimal uses parentheses formatting. - Include the suffix when calculating width, left alignment, and zero padding. - Add regression coverage for grouped negative decimals with and without an explicit width. ## Are these changes tested? Yes. The following checks pass locally: - `cargo fmt --all -- --check` - `cargo clippy --all-targets --all-features -- -D warnings` - `cargo test -p datafusion-spark --lib` - The full workspace test command from the contributor guide with `avro,json,backtrace,extended_tests,recursive_protection,parquet_encryption` enabled ## Are there any user-facing changes? Yes. Spark-compatible formatting of negative decimal values now honors the parentheses flag. There are no public API or breaking changes.
## Which issue does this PR close? <!-- We generally require a GitHub issue to be filed for all bug fixes and enhancements and this helps us generate change logs for our releases. You can link an issue to this PR using the GitHub syntax. For example `Closes #123` indicates that this PR will close issue #123. --> - Closes #. ## Rationale for this change <!-- Why are you proposing this change? If this is already explained clearly in the issue then this section is not needed. Explaining clearly why changes are proposed helps reviewers understand your changes and offer better suggestions for fixes. --> Precedes apache#23720. While displaying a plan with metrics, allow filtering them by name, not just by category or type. This is very useful when writing snapshot tests where only a specific metrics needs to be asserted, without other unrelated metrics polluting the snapshot assertion. Regardless of what happens with apache#23720, I think this PR is still worth it, as it allows creating some really nice `insta` tests asserting runtime properties reliably by just cherry picking the runtime metrics relevant for that specific test. This is relevant not only within DataFusion codebase, but also for other people's codebases using DataFusion and relying on `insta` for snapshot testing, See an example of this here: https://github.com/apache/datafusion/pull/23720/changes#diff-8281405e117428077c19c07afbbe59fed303a3460096e958f761d34f0899b618 ## What changes are included in this PR? <!-- There is no need to duplicate the description in the issue here but it is sometimes worth providing a summary of the individual changes in this PR. --> While displaying metrics, allows filtering them by name ## Are these changes tested? <!-- We typically require tests for all PRs in order to: 1. Prevent the code from being accidentally broken by subsequent changes 2. Serve as another way to document the expected behavior of the code If tests are not included in your PR, please explain why (for example, are they covered by existing tests)? --> Yes, by a new small test. ## Are there any user-facing changes? <!-- If there are user-facing changes then we may require documentation to be updated before approving the PR. --> <!-- If there are any breaking changes to public APIs, please add the `api change` label. --> People using the metrics display API will be able to filter metrics by name.
## Which issue does this PR close? N/A ## Rationale for this change I saw that `empty` implementation is inefficient. I wrote review comments for more why inefficient. there are some optimizations that do not require benchmark since they are obvious once you understand, this is one of them ## What changes are included in this PR? rewrote the function to be fast ## Are these changes tested? existing tests ## Are there any user-facing changes? no
Bumps the codeql-actions group with 2 updates: [github/codeql-action/init](https://github.com/github/codeql-action) and [github/codeql-action/analyze](https://github.com/github/codeql-action). Updates `github/codeql-action/init` from 4.37.1 to 4.37.3 <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/github/codeql-action/releases">github/codeql-action/init's releases</a>.</em></p> <blockquote> <h2>v4.37.3</h2> <p>No user facing changes.</p> <h2>v4.37.2</h2> <ul> <li>The new address format for the <code>config-file</code> input that was introduced in CodeQL Action 4.37.0 is now enabled by default. In addition to the format described there, the <code>remote=</code> prefix can now be used to explicitly indicate that the input refers to a remote file. All previous input formats continue to be accepted as well. <a href="https://redirect.github.com/github/codeql-action/pull/4023">#4023</a></li> <li>The CodeQL Action can now make use of <a href="https://docs.github.com/en/code-security/how-tos/secure-at-scale/configure-organization-security/manage-usage-and-access/giving-org-access-private-registries">configured private registries</a> in Default Setup to retrieve CodeQL configuration files from remote repositories that require authentication. This will allow customers to store their CodeQL configuration in a single repository that can then be referenced by Default Setup workflows in other repositories. We expect to roll this and other, related changes out to everyone in July. <a href="https://redirect.github.com/github/codeql-action/pull/4007">#4007</a></li> </ul> </blockquote> </details> <details> <summary>Changelog</summary> <p><em>Sourced from <a href="https://github.com/github/codeql-action/blob/main/CHANGELOG.md">github/codeql-action/init's changelog</a>.</em></p> <blockquote> <h1>CodeQL Action Changelog</h1> <p>See the <a href="https://github.com/github/codeql-action/releases">releases page</a> for the relevant changes to the CodeQL CLI and language packs.</p> <h2>[UNRELEASED]</h2> <ul> <li>This version of the CodeQL Action adds support for the <code>tools</code> input for the <code>codeql-action/init</code> step to be specified using a <code>github-codeql-tools</code> <a href="https://docs.github.com/en/organizations/managing-organization-settings/managing-custom-properties-for-repositories-in-your-organization">repository property</a>. This feature will gradually be rolled out following the release of this version. Once rolled out, this allows for the CodeQL CLI version that is used in GitHub-managed workflows, such as Default Setup, to be set to a custom value. For example, customers who run into issues with rate limits when a new CodeQL CLI version is released can set the value to <code>toolcache</code> to always use the CodeQL CLI version that is available in the runner toolcache. For Advanced Setup workflows, the value provided for <code>tools</code> in the workflow definition always takes precedence unless the value of the repository property starts with <code>!</code>. <a href="https://redirect.github.com/github/codeql-action/pull/4037">#4037</a></li> </ul> <h2>4.37.3 - 22 Jul 2026</h2> <p>No user facing changes.</p> <h2>4.37.2 - 21 Jul 2026</h2> <ul> <li>The new address format for the <code>config-file</code> input that was introduced in CodeQL Action 4.37.0 is now enabled by default. In addition to the format described there, the <code>remote=</code> prefix can now be used to explicitly indicate that the input refers to a remote file. All previous input formats continue to be accepted as well. <a href="https://redirect.github.com/github/codeql-action/pull/4023">#4023</a></li> <li>The CodeQL Action can now make use of <a href="https://docs.github.com/en/code-security/how-tos/secure-at-scale/configure-organization-security/manage-usage-and-access/giving-org-access-private-registries">configured private registries</a> in Default Setup to retrieve CodeQL configuration files from remote repositories that require authentication. This will allow customers to store their CodeQL configuration in a single repository that can then be referenced by Default Setup workflows in other repositories. We expect to roll this and other, related changes out to everyone in July. <a href="https://redirect.github.com/github/codeql-action/pull/4007">#4007</a></li> </ul> <h2>4.37.1 - 16 Jul 2026</h2> <ul> <li><em>Upcoming breaking change</em>: Add a deprecation warning for customers using CodeQL version 2.20.6 and earlier. These versions of CodeQL were discontinued on 1 July 2026 alongside GitHub Enterprise Server 3.16, and will be unsupported by the next minor release of the CodeQL Action. <a href="https://redirect.github.com/github/codeql-action/pull/3956">#3956</a></li> <li>Update default CodeQL bundle version to <a href="https://github.com/github/codeql-action/releases/tag/codeql-bundle-v2.26.1">2.26.1</a>. <a href="https://redirect.github.com/github/codeql-action/pull/4019">#4019</a></li> </ul> <h2>4.37.0 - 08 Jul 2026</h2> <ul> <li>Update default CodeQL bundle version to <a href="https://github.com/github/codeql-action/releases/tag/codeql-bundle-v2.26.0">2.26.0</a>. <a href="https://redirect.github.com/github/codeql-action/pull/3995">#3995</a></li> <li>In addition to the existing input format, the <code>config-file</code> input for the <code>codeql-action/init</code> step will soon support a new <code>[owner/]repo[@ref][:path]</code> format. All components except the repository name are optional. If omitted, <code>owner</code> defaults to the same owner as the repository the analysis is running for, <code>ref</code> to <code>main</code>, and <code>path</code> to <code>.github/codeql-action.yaml</code>. Support for this format ships in this version of the CodeQL Action, but will only be enabled over the coming weeks. <a href="https://redirect.github.com/github/codeql-action/pull/3973">#3973</a></li> </ul> <h2>4.36.3 - 01 Jul 2026</h2> <p>No user facing changes.</p> <h2>4.36.2 - 04 Jun 2026</h2> <ul> <li>Cache CodeQL CLI version information across Actions steps. <a href="https://redirect.github.com/github/codeql-action/pull/3943">#3943</a></li> <li>Reduce requests while waiting for analysis processing by using exponential backoff when polling SARIF processing status. <a href="https://redirect.github.com/github/codeql-action/pull/3937">#3937</a></li> <li>Update default CodeQL bundle version to <a href="https://github.com/github/codeql-action/releases/tag/codeql-bundle-v2.25.6">2.25.6</a>. <a href="https://redirect.github.com/github/codeql-action/pull/3948">#3948</a></li> </ul> <h2>4.36.1 - 02 Jun 2026</h2> <p>No user facing changes.</p> <h2>4.36.0 - 22 May 2026</h2> <ul> <li><em>Breaking change</em>: Bump the minimum required CodeQL bundle version to 2.19.4. <a href="https://redirect.github.com/github/codeql-action/pull/3894">#3894</a></li> <li>Add support for SHA-256 Git object IDs. <a href="https://redirect.github.com/github/codeql-action/pull/3893">#3893</a></li> <li>Update default CodeQL bundle version to <a href="https://github.com/github/codeql-action/releases/tag/codeql-bundle-v2.25.5">2.25.5</a>. <a href="https://redirect.github.com/github/codeql-action/pull/3926">#3926</a></li> </ul> <h2>4.35.5 - 15 May 2026</h2> <ul> <li>We have improved how the JavaScript bundles for the CodeQL Action are generated to avoid duplication across bundles and reduce the size of the repository by around 70%. This should have no effect on the runtime behaviour of the CodeQL Action. <a href="https://redirect.github.com/github/codeql-action/pull/3899">#3899</a></li> </ul> <!-- raw HTML omitted --> </blockquote> <p>... (truncated)</p> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/github/codeql-action/commit/e4fba868fa4b1b91e1fdab776edc8cfbe6e9fb81"><code>e4fba86</code></a> Merge pull request <a href="https://redirect.github.com/github/codeql-action/issues/4031">#4031</a> from github/update-v4.37.3-72f6a9da0</li> <li><a href="https://github.com/github/codeql-action/commit/fb50ab5d62a274adf3ef3e22cfe750ae87a0ede7"><code>fb50ab5</code></a> Update changelog for v4.37.3</li> <li><a href="https://github.com/github/codeql-action/commit/72f6a9da0def52d9193d6a758f0378b65091f8d1"><code>72f6a9d</code></a> Merge pull request <a href="https://redirect.github.com/github/codeql-action/issues/4030">#4030</a> from github/mbg/fix/no-proxy</li> <li><a href="https://github.com/github/codeql-action/commit/3b5ee58597653d9cc6785f3f1277f796d81f3646"><code>3b5ee58</code></a> Use default <code>request</code> options instead of <code>undefined</code></li> <li><a href="https://github.com/github/codeql-action/commit/bfb6be4b5ecd3650f02f530571453e8c64ef0778"><code>bfb6be4</code></a> Merge pull request <a href="https://redirect.github.com/github/codeql-action/issues/4028">#4028</a> from github/mergeback/v4.37.2-to-main-e0647621</li> <li><a href="https://github.com/github/codeql-action/commit/526ab84f9858816d9cf5f7b9df4dd5e2235f0eba"><code>526ab84</code></a> Rebuild</li> <li><a href="https://github.com/github/codeql-action/commit/d6217b9b8c14166e4851db94c11155d03bd13c07"><code>d6217b9</code></a> Update changelog and version after v4.37.2</li> <li><a href="https://github.com/github/codeql-action/commit/e0647621c2984b5ed2f768cb892365bf2a616ad1"><code>e064762</code></a> Merge pull request <a href="https://redirect.github.com/github/codeql-action/issues/4027">#4027</a> from github/update-v4.37.2-385bcdc5a</li> <li><a href="https://github.com/github/codeql-action/commit/e0faed839190caa67a5cd42f1cc16246028ca3df"><code>e0faed8</code></a> Add a couple of change notes</li> <li><a href="https://github.com/github/codeql-action/commit/73aad0eaa9df172668665a150d17b8bc5a650c20"><code>73aad0e</code></a> Update changelog for v4.37.2</li> <li>Additional commits viewable in <a href="https://github.com/github/codeql-action/compare/7188fc363630916deb702c7fdcf4e481b751f97a...e4fba868fa4b1b91e1fdab776edc8cfbe6e9fb81">compare view</a></li> </ul> </details> <br /> Updates `github/codeql-action/analyze` from 4.37.1 to 4.37.3 <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/github/codeql-action/releases">github/codeql-action/analyze's releases</a>.</em></p> <blockquote> <h2>v4.37.3</h2> <p>No user facing changes.</p> <h2>v4.37.2</h2> <ul> <li>The new address format for the <code>config-file</code> input that was introduced in CodeQL Action 4.37.0 is now enabled by default. In addition to the format described there, the <code>remote=</code> prefix can now be used to explicitly indicate that the input refers to a remote file. All previous input formats continue to be accepted as well. <a href="https://redirect.github.com/github/codeql-action/pull/4023">#4023</a></li> <li>The CodeQL Action can now make use of <a href="https://docs.github.com/en/code-security/how-tos/secure-at-scale/configure-organization-security/manage-usage-and-access/giving-org-access-private-registries">configured private registries</a> in Default Setup to retrieve CodeQL configuration files from remote repositories that require authentication. This will allow customers to store their CodeQL configuration in a single repository that can then be referenced by Default Setup workflows in other repositories. We expect to roll this and other, related changes out to everyone in July. <a href="https://redirect.github.com/github/codeql-action/pull/4007">#4007</a></li> </ul> </blockquote> </details> <details> <summary>Changelog</summary> <p><em>Sourced from <a href="https://github.com/github/codeql-action/blob/main/CHANGELOG.md">github/codeql-action/analyze's changelog</a>.</em></p> <blockquote> <h1>CodeQL Action Changelog</h1> <p>See the <a href="https://github.com/github/codeql-action/releases">releases page</a> for the relevant changes to the CodeQL CLI and language packs.</p> <h2>[UNRELEASED]</h2> <ul> <li>This version of the CodeQL Action adds support for the <code>tools</code> input for the <code>codeql-action/init</code> step to be specified using a <code>github-codeql-tools</code> <a href="https://docs.github.com/en/organizations/managing-organization-settings/managing-custom-properties-for-repositories-in-your-organization">repository property</a>. This feature will gradually be rolled out following the release of this version. Once rolled out, this allows for the CodeQL CLI version that is used in GitHub-managed workflows, such as Default Setup, to be set to a custom value. For example, customers who run into issues with rate limits when a new CodeQL CLI version is released can set the value to <code>toolcache</code> to always use the CodeQL CLI version that is available in the runner toolcache. For Advanced Setup workflows, the value provided for <code>tools</code> in the workflow definition always takes precedence unless the value of the repository property starts with <code>!</code>. <a href="https://redirect.github.com/github/codeql-action/pull/4037">#4037</a></li> </ul> <h2>4.37.3 - 22 Jul 2026</h2> <p>No user facing changes.</p> <h2>4.37.2 - 21 Jul 2026</h2> <ul> <li>The new address format for the <code>config-file</code> input that was introduced in CodeQL Action 4.37.0 is now enabled by default. In addition to the format described there, the <code>remote=</code> prefix can now be used to explicitly indicate that the input refers to a remote file. All previous input formats continue to be accepted as well. <a href="https://redirect.github.com/github/codeql-action/pull/4023">#4023</a></li> <li>The CodeQL Action can now make use of <a href="https://docs.github.com/en/code-security/how-tos/secure-at-scale/configure-organization-security/manage-usage-and-access/giving-org-access-private-registries">configured private registries</a> in Default Setup to retrieve CodeQL configuration files from remote repositories that require authentication. This will allow customers to store their CodeQL configuration in a single repository that can then be referenced by Default Setup workflows in other repositories. We expect to roll this and other, related changes out to everyone in July. <a href="https://redirect.github.com/github/codeql-action/pull/4007">#4007</a></li> </ul> <h2>4.37.1 - 16 Jul 2026</h2> <ul> <li><em>Upcoming breaking change</em>: Add a deprecation warning for customers using CodeQL version 2.20.6 and earlier. These versions of CodeQL were discontinued on 1 July 2026 alongside GitHub Enterprise Server 3.16, and will be unsupported by the next minor release of the CodeQL Action. <a href="https://redirect.github.com/github/codeql-action/pull/3956">#3956</a></li> <li>Update default CodeQL bundle version to <a href="https://github.com/github/codeql-action/releases/tag/codeql-bundle-v2.26.1">2.26.1</a>. <a href="https://redirect.github.com/github/codeql-action/pull/4019">#4019</a></li> </ul> <h2>4.37.0 - 08 Jul 2026</h2> <ul> <li>Update default CodeQL bundle version to <a href="https://github.com/github/codeql-action/releases/tag/codeql-bundle-v2.26.0">2.26.0</a>. <a href="https://redirect.github.com/github/codeql-action/pull/3995">#3995</a></li> <li>In addition to the existing input format, the <code>config-file</code> input for the <code>codeql-action/init</code> step will soon support a new <code>[owner/]repo[@ref][:path]</code> format. All components except the repository name are optional. If omitted, <code>owner</code> defaults to the same owner as the repository the analysis is running for, <code>ref</code> to <code>main</code>, and <code>path</code> to <code>.github/codeql-action.yaml</code>. Support for this format ships in this version of the CodeQL Action, but will only be enabled over the coming weeks. <a href="https://redirect.github.com/github/codeql-action/pull/3973">#3973</a></li> </ul> <h2>4.36.3 - 01 Jul 2026</h2> <p>No user facing changes.</p> <h2>4.36.2 - 04 Jun 2026</h2> <ul> <li>Cache CodeQL CLI version information across Actions steps. <a href="https://redirect.github.com/github/codeql-action/pull/3943">#3943</a></li> <li>Reduce requests while waiting for analysis processing by using exponential backoff when polling SARIF processing status. <a href="https://redirect.github.com/github/codeql-action/pull/3937">#3937</a></li> <li>Update default CodeQL bundle version to <a href="https://github.com/github/codeql-action/releases/tag/codeql-bundle-v2.25.6">2.25.6</a>. <a href="https://redirect.github.com/github/codeql-action/pull/3948">#3948</a></li> </ul> <h2>4.36.1 - 02 Jun 2026</h2> <p>No user facing changes.</p> <h2>4.36.0 - 22 May 2026</h2> <ul> <li><em>Breaking change</em>: Bump the minimum required CodeQL bundle version to 2.19.4. <a href="https://redirect.github.com/github/codeql-action/pull/3894">#3894</a></li> <li>Add support for SHA-256 Git object IDs. <a href="https://redirect.github.com/github/codeql-action/pull/3893">#3893</a></li> <li>Update default CodeQL bundle version to <a href="https://github.com/github/codeql-action/releases/tag/codeql-bundle-v2.25.5">2.25.5</a>. <a href="https://redirect.github.com/github/codeql-action/pull/3926">#3926</a></li> </ul> <h2>4.35.5 - 15 May 2026</h2> <ul> <li>We have improved how the JavaScript bundles for the CodeQL Action are generated to avoid duplication across bundles and reduce the size of the repository by around 70%. This should have no effect on the runtime behaviour of the CodeQL Action. <a href="https://redirect.github.com/github/codeql-action/pull/3899">#3899</a></li> </ul> <!-- raw HTML omitted --> </blockquote> <p>... (truncated)</p> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/github/codeql-action/commit/e4fba868fa4b1b91e1fdab776edc8cfbe6e9fb81"><code>e4fba86</code></a> Merge pull request <a href="https://redirect.github.com/github/codeql-action/issues/4031">#4031</a> from github/update-v4.37.3-72f6a9da0</li> <li><a href="https://github.com/github/codeql-action/commit/fb50ab5d62a274adf3ef3e22cfe750ae87a0ede7"><code>fb50ab5</code></a> Update changelog for v4.37.3</li> <li><a href="https://github.com/github/codeql-action/commit/72f6a9da0def52d9193d6a758f0378b65091f8d1"><code>72f6a9d</code></a> Merge pull request <a href="https://redirect.github.com/github/codeql-action/issues/4030">#4030</a> from github/mbg/fix/no-proxy</li> <li><a href="https://github.com/github/codeql-action/commit/3b5ee58597653d9cc6785f3f1277f796d81f3646"><code>3b5ee58</code></a> Use default <code>request</code> options instead of <code>undefined</code></li> <li><a href="https://github.com/github/codeql-action/commit/bfb6be4b5ecd3650f02f530571453e8c64ef0778"><code>bfb6be4</code></a> Merge pull request <a href="https://redirect.github.com/github/codeql-action/issues/4028">#4028</a> from github/mergeback/v4.37.2-to-main-e0647621</li> <li><a href="https://github.com/github/codeql-action/commit/526ab84f9858816d9cf5f7b9df4dd5e2235f0eba"><code>526ab84</code></a> Rebuild</li> <li><a href="https://github.com/github/codeql-action/commit/d6217b9b8c14166e4851db94c11155d03bd13c07"><code>d6217b9</code></a> Update changelog and version after v4.37.2</li> <li><a href="https://github.com/github/codeql-action/commit/e0647621c2984b5ed2f768cb892365bf2a616ad1"><code>e064762</code></a> Merge pull request <a href="https://redirect.github.com/github/codeql-action/issues/4027">#4027</a> from github/update-v4.37.2-385bcdc5a</li> <li><a href="https://github.com/github/codeql-action/commit/e0faed839190caa67a5cd42f1cc16246028ca3df"><code>e0faed8</code></a> Add a couple of change notes</li> <li><a href="https://github.com/github/codeql-action/commit/73aad0eaa9df172668665a150d17b8bc5a650c20"><code>73aad0e</code></a> Update changelog for v4.37.2</li> <li>Additional commits viewable in <a href="https://github.com/github/codeql-action/compare/7188fc363630916deb702c7fdcf4e481b751f97a...e4fba868fa4b1b91e1fdab776edc8cfbe6e9fb81">compare view</a></li> </ul> </details> <br /> Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting `@dependabot rebase`. [//]: # (dependabot-automerge-start) [//]: # (dependabot-automerge-end) --- <details> <summary>Dependabot commands and options</summary> <br /> You can trigger Dependabot actions by commenting on this PR: - `@dependabot rebase` will rebase this PR - `@dependabot recreate` will recreate this PR, overwriting any edits that have been made to it - `@dependabot show <dependency name> ignore conditions` will show all of the ignore conditions of the specified dependency - `@dependabot ignore <dependency name> major version` will close this group update PR and stop Dependabot creating any more for the specific dependency's major version (unless you unignore this specific dependency's major version or upgrade to it yourself) - `@dependabot ignore <dependency name> minor version` will close this group update PR and stop Dependabot creating any more for the specific dependency's minor version (unless you unignore this specific dependency's minor version or upgrade to it yourself) - `@dependabot ignore <dependency name>` will close this group update PR and stop Dependabot creating any more for the specific dependency (unless you unignore this specific dependency or upgrade to it yourself) - `@dependabot unignore <dependency name>` will remove all of the ignore conditions of the specified dependency - `@dependabot unignore <dependency name> <ignore condition>` will remove the ignore condition of the specified dependency and ignore conditions </details> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
…e#23941) Bumps [taiki-e/install-action](https://github.com/taiki-e/install-action) from 2.84.0 to 2.85.2. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/taiki-e/install-action/releases">taiki-e/install-action's releases</a>.</em></p> <blockquote> <h2>2.85.2</h2> <ul> <li> <p>Update <code>prek@latest</code> to 0.4.11.</p> </li> <li> <p>Update <code>mise@latest</code> to 2026.7.13.</p> </li> <li> <p>Update <code>kingfisher@latest</code> to 1.109.0.</p> </li> </ul> <h2>2.85.1</h2> <ul> <li> <p>Update <code>vacuum@latest</code> to 0.30.0.</p> </li> <li> <p>Update <code>uv@latest</code> to 0.11.32.</p> </li> <li> <p>Update <code>mise@latest</code> to 2026.7.12.</p> </li> <li> <p>Update <code>cyclonedx@latest</code> to 0.33.1.</p> </li> <li> <p>Update <code>cargo-neat@latest</code> to 0.5.2.</p> </li> </ul> <h2>2.85.0</h2> <ul> <li> <p>Support <code>wild</code> (alias: <code>wild-linker</code>). (<a href="https://redirect.github.com/taiki-e/install-action/pull/1949">#1949</a>)</p> </li> <li> <p>Support <code>bpf-linker</code>. (<a href="https://redirect.github.com/taiki-e/install-action/pull/1950">#1950</a>)</p> </li> <li> <p>Support <code>rafn</code>. (<a href="https://redirect.github.com/taiki-e/install-action/pull/1935">#1935</a>, thanks <a href="https://github.com/DarkWanderer"><code>@DarkWanderer</code></a>)</p> </li> <li> <p>Update <code>cargo-neat@latest</code> to 0.5.1.</p> </li> <li> <p>Update <code>zizmor@latest</code> to 1.28.0.</p> </li> <li> <p>Update <code>wasmtime@latest</code> to 47.0.2.</p> </li> <li> <p>Update <code>uv@latest</code> to 0.11.31.</p> </li> <li> <p>Update <code>syft@latest</code> to 1.49.0.</p> </li> </ul> <h2>2.84.1</h2> <ul> <li> <p>Update <code>wasmtime@latest</code> to 47.0.1.</p> </li> <li> <p>Update <code>wasm-tools@latest</code> to 1.254.0.</p> </li> <li> <p>Update <code>uv@latest</code> to 0.11.30.</p> </li> <li> <p>Update <code>mise@latest</code> to 2026.7.11.</p> </li> <li> <p>Update <code>cargo-neat@latest</code> to 0.5.0.</p> </li> <li> <p>Update <code>cargo-crap@latest</code> to 0.3.1.</p> </li> <li> <p>Update <code>biome@latest</code> to 2.5.5.</p> </li> </ul> </blockquote> </details> <details> <summary>Changelog</summary> <p><em>Sourced from <a href="https://github.com/taiki-e/install-action/blob/main/CHANGELOG.md">taiki-e/install-action's changelog</a>.</em></p> <blockquote> <h1>Changelog</h1> <p>All notable changes to this project will be documented in this file.</p> <p>This project adheres to <a href="https://semver.org">Semantic Versioning</a>.</p> <!-- raw HTML omitted --> <h2>[Unreleased]</h2> <h2>[2.85.2] - 2026-07-26</h2> <ul> <li> <p>Update <code>prek@latest</code> to 0.4.11.</p> </li> <li> <p>Update <code>mise@latest</code> to 2026.7.13.</p> </li> <li> <p>Update <code>kingfisher@latest</code> to 1.109.0.</p> </li> </ul> <h2>[2.85.1] - 2026-07-25</h2> <ul> <li> <p>Update <code>vacuum@latest</code> to 0.30.0.</p> </li> <li> <p>Update <code>uv@latest</code> to 0.11.32.</p> </li> <li> <p>Update <code>mise@latest</code> to 2026.7.12.</p> </li> <li> <p>Update <code>cyclonedx@latest</code> to 0.33.1.</p> </li> <li> <p>Update <code>cargo-neat@latest</code> to 0.5.2.</p> </li> </ul> <h2>[2.85.0] - 2026-07-23</h2> <ul> <li> <p>Support <code>wild</code> (alias: <code>wild-linker</code>). (<a href="https://redirect.github.com/taiki-e/install-action/pull/1949">#1949</a>)</p> </li> <li> <p>Support <code>bpf-linker</code>. (<a href="https://redirect.github.com/taiki-e/install-action/pull/1950">#1950</a>)</p> </li> <li> <p>Support <code>rafn</code>. (<a href="https://redirect.github.com/taiki-e/install-action/pull/1935">#1935</a>, thanks <a href="https://github.com/DarkWanderer"><code>@DarkWanderer</code></a>)</p> </li> <li> <p>Update <code>cargo-neat@latest</code> to 0.5.1.</p> </li> <li> <p>Update <code>zizmor@latest</code> to 1.28.0.</p> </li> <li> <p>Update <code>wasmtime@latest</code> to 47.0.2.</p> </li> <li> <p>Update <code>uv@latest</code> to 0.11.31.</p> </li> <li> <p>Update <code>syft@latest</code> to 1.49.0.</p> </li> </ul> <!-- raw HTML omitted --> </blockquote> <p>... (truncated)</p> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/taiki-e/install-action/commit/41049aa56687c35e0afa74eed4f09cec4f9afabf"><code>41049aa</code></a> Release 2.85.2</li> <li><a href="https://github.com/taiki-e/install-action/commit/dfcf36552b1b743e910b9b9b6b114a08e56a1e5e"><code>dfcf365</code></a> Update <code>prek@latest</code> to 0.4.11</li> <li><a href="https://github.com/taiki-e/install-action/commit/eea03ccfa855b965a301aa52424915fddc32f855"><code>eea03cc</code></a> Update <code>mise@latest</code> to 2026.7.13</li> <li><a href="https://github.com/taiki-e/install-action/commit/81ca2feb84748ed3fa514707ee1004f4f3ad715a"><code>81ca2fe</code></a> Update martin manifest</li> <li><a href="https://github.com/taiki-e/install-action/commit/cc90ed04bc1a258ea90a83789f24b759a77c9671"><code>cc90ed0</code></a> Update <code>kingfisher@latest</code> to 1.109.0</li> <li><a href="https://github.com/taiki-e/install-action/commit/55639a3362f508fda5154d3451072205f5c32ca6"><code>55639a3</code></a> Update cargo-shear manifest</li> <li><a href="https://github.com/taiki-e/install-action/commit/3d7d7cd5ac7f994c1892ae0c06165095b9139094"><code>3d7d7cd</code></a> Release 2.85.1</li> <li><a href="https://github.com/taiki-e/install-action/commit/d09ccb4fe2105ac9ead6376f7a423ff88defeb57"><code>d09ccb4</code></a> Update <code>vacuum@latest</code> to 0.30.0</li> <li><a href="https://github.com/taiki-e/install-action/commit/ac43dee1a92e482c0e244bdb0bb126908159e245"><code>ac43dee</code></a> Update <code>uv@latest</code> to 0.11.32</li> <li><a href="https://github.com/taiki-e/install-action/commit/49b16979f38f30e44b1f68af9115be9a6ffc5215"><code>49b1697</code></a> Update prek manifest</li> <li>Additional commits viewable in <a href="https://github.com/taiki-e/install-action/compare/a6b2e2dcd845ddd7f509ce4f3ed3d922b80cc5d9...41049aa56687c35e0afa74eed4f09cec4f9afabf">compare view</a></li> </ul> </details> <br /> [](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores) Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting `@dependabot rebase`. [//]: # (dependabot-automerge-start) [//]: # (dependabot-automerge-end) --- <details> <summary>Dependabot commands and options</summary> <br /> You can trigger Dependabot actions by commenting on this PR: - `@dependabot rebase` will rebase this PR - `@dependabot recreate` will recreate this PR, overwriting any edits that have been made to it - `@dependabot show <dependency name> ignore conditions` will show all of the ignore conditions of the specified dependency - `@dependabot ignore this major version` will close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this minor version` will close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this dependency` will close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself) </details> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
9 of 58 tasks
There was a problem hiding this comment.
Copilot review overview
🔵 Needs a closer look
The 900-commit, 1,720-file merge spans core execution, optimizer, API, tooling, and benchmark behavior and requires final human validation.
Review effort: Balanced
Findings: None
What changed in this PR
Merges DataFusion 55.1.0 into spiceai-55 while retaining Spice-specific correctness and compatibility patches.
Changes:
- Integrates upstream optimizer, execution, API, dependency, and benchmark updates.
- Fixes join sum statistics and dialect-aware Date32 unparsing.
- Preserves Spark feature isolation and Rust 1.94 compatibility.
| File | Description |
|---|---|
datafusion/** |
Upstream engine changes and Spice fixes |
benchmarks/** |
New benchmark framework and suites |
docs/** |
55.1 documentation and upgrade guidance |
.github/** |
CI and dependency automation updates |
| Root configuration files | Toolchain, licensing, and lint updates |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
…o spiceai-59 (#244)
Jeadie
approved these changes
Oct 2, 2026
krinart
marked this pull request as draft
October 2, 2026 02:09
phillipleblanc
approved these changes
Oct 2, 2026
krinart
marked this pull request as ready for review
October 2, 2026 02:33
krinart
added a commit
to spiceai/datafusion-ballista
that referenced
this pull request
Oct 2, 2026
…into spiceai-55
krinart
added a commit
to spiceai/datafusion-federation
that referenced
this pull request
Oct 2, 2026
…into spiceai-55
krinart
added a commit
to spiceai/datafusion-table-providers
that referenced
this pull request
Oct 2, 2026
…into spiceai-55
This was referenced Oct 2, 2026
Merged
krinart
added a commit
to spiceai/spiceai
that referenced
this pull request
Oct 2, 2026
…nowflake-rs at their arrow-rs bumps - DataFusion: 02550cf9, the merge of spiceai/datafusion#243 into spiceai-55 (same tree as the previous pin); no longer a temporary pin. - iceberg-rust: 0284be3d and snowflake-rs: b95a74e8, which move each fork's own arrow-rs pin to 2e2cc330 (Cargo manifests only; Spice builds the same code).
krinart
added a commit
to spiceai/datafusion-ballista
that referenced
this pull request
Oct 2, 2026
* build: pin arrow-rs at 2e2cc330, the merge of spiceai/arrow-rs#29 into spiceai-59 * build: pin DataFusion at 02550cf9, the merge of spiceai/datafusion#243 into spiceai-55
krinart
added a commit
to spiceai/datafusion-table-providers
that referenced
this pull request
Oct 2, 2026
* build: pin arrow-rs at 2e2cc330, the merge of spiceai/arrow-rs#29 into spiceai-59 * build: pin DataFusion at 02550cf9, the merge of spiceai/datafusion#243 into spiceai-55
krinart
added a commit
to spiceai/datafusion-federation
that referenced
this pull request
Oct 2, 2026
* build: pin arrow-rs at 2e2cc330, the merge of spiceai/arrow-rs#29 into spiceai-59 * build: pin DataFusion at 02550cf9, the merge of spiceai/datafusion#243 into spiceai-55
krinart
added a commit
to spiceai/iceberg-rust
that referenced
this pull request
Oct 2, 2026
…into spiceai-55 (#53)
lukekim
pushed a commit
to spiceai/datafusion-table-providers
that referenced
this pull request
Oct 2, 2026
* build: pin arrow-rs at 2e2cc330, the merge of spiceai/arrow-rs#29 into spiceai-59 * build: pin DataFusion at 02550cf9, the merge of spiceai/datafusion#243 into spiceai-55
lukekim
pushed a commit
to spiceai/iceberg-rust
that referenced
this pull request
Oct 3, 2026
…into spiceai-55 (#53)
anush008
pushed a commit
to anush008/spiceai
that referenced
this pull request
Oct 6, 2026
* build: pin the DataFusion 55 / Arrow 59 fork lines Moves every fork to its DataFusion 55 line: arrow-rs 59.3.0, DataFusion 55.1.0, Ballista (upstream main on DataFusion 55), federation, table-providers, iceberg-rust 0.10.1, duckdb-rs 1.4.4 on arrow 59, ADBC 24, snowflake-rs, spark-connect-rs, spice-rs and spicebench on arrow 59. datafusion-functions-json comes from crates.io (0.55.6); delta-kernel stays on 0.27.1 with its default engine split into delta_kernel_default_engine. docs/dev/fork_patches.md is re-audited against the new lines; the pins stay marked TEMPORARY until the fork pull requests merge. * feat: upgrade to DataFusion 55 and Arrow 59 - Statistics flow through DataFusion 55's StatisticsContext. The deprecated partition_statistics returns Absent for built-in plans on 55, so every wrapper implements child_stats_requests / statistics_from_inputs and the optimizer rules, join sizing and Flight batch sizing read real statistics again. - ExecutionPlan / TableProvider wrappers implement the new required and defaulted methods explicitly (apply_expressions, replace_children, try_to_proto, merge_into, ...), forwarding where they wrap. - Arrow 59 refuses a MAP whose entries field is nullable: Flight SQL reads replace the declared schema before decoding and normalize the map afterwards, and legacy IPC maps decode as a list (arrow_tools map_entries). - Delta Lake refuses Void and interval columns with a structured error instead of mis-reading them. - Ballista 55: Vcores, the protocol-version handshake, task-id based cancellation and the append-only task-info model. - Vortex 0.86 and the remaining API moves (TableReference, TableScanBuilder, TableSchema builder, codec proto converter, WriteOp::MergeInto). * ci: validate the datafusion-table-providers pin against spiceai-55-patches The DataFusion 55 line of the fork is spiceai-55-patches. * revert: drop the unfinished guard tests d65aa29 picked up by accident d65aa29 was meant to change only the CI branch check; it also committed work-in-progress test edits. Restore those files to 98800a8; the finished tests land in their own commit. * fix(duckdb): push regexp_count down NULL-preserving, as DataFusion 55 evaluates it DataFusion 55 (apache/datafusion#24239) makes regexp_count propagate NULL, as PostgreSQL does. The DuckDB rendering still wrapped the count in coalesce(.., 0) to match the old kernel, so a DuckDB-accelerated dataset answered 0 where local evaluation answers NULL, and a WHERE on the count kept rows local evaluation drops. Render it as len(regexp_extract_all(..)), which is NULL for a NULL input or pattern. * test(cayenne): count -0.0 as equal to 0.0 in the boundary-values test DataFusion 55 compares floats in SQL by IEEE 754 value, so float_val = 0.0 matches the -0.0 row too (4 rows, not 3). A MemTable in the same session returns the same rows, so this is the expectation, not Cayenne. * build: repin the forks and guard their new patches - Ballista 8bf37ad0: the executor poll loop accepts a vcore semaphore with no permits yet. Upstream's assert panicked the loop at startup, so no Spice executor registered and every cluster query failed. - DataFusion d98ae645 (MSRV-clean unparser), federation 9ca84a39, table-providers 661d16d1, iceberg 575d81a1: their own tests now run against the Spice DataFusion (and, for table-providers, ADBC) forks. - Guard tests for the unparser rescope, eager aggregation through StatisticsContext, SchemaCastScanExec statistics forwarding and Ballista TaskInfo fields 11/12; the ledger names them and records the Ballista fix. - The distributed EXPLAIN snapshot now reads Ballista 55's UnknownPartitioning(1) where it printed None. * build: point the fork pins at the spiceai-*-patches-2 branches The DataFusion 55 lines of arrow-rs, datafusion, datafusion-ballista, datafusion-federation, datafusion-table-providers, duckdb-rs, snowflake-rs and spark-connect-rs are reviewed on new spiceai-*-patches-2 branches, leaving the earlier spiceai-*-patches branches untouched. Same revisions; only the branch names in the pin comments, the fork-patch ledger and the table-providers pin check change. * fix: re-root a delegated projection on the inner plan before swapping DataFusion 55's projection pushdown reads a projection's input as the node being swapped with: `ProjectionExec` collapses the chain starting at `projection.input()`, `CoalescePartitionsExec` rebuilds from its child. A wrapper that handed its inner plan the projection unchanged (its input being the wrapper) got it back still above the wrapper, and the pushdown then nested another copy on every step: the Cayenne maintained-aggregate soundness test grew to gigabytes and never finished. CayenneAccelerationExec, the DuckDB aggregate pushdown marker and PartitionedUnionExec now rebuild the projection over their inner plan (whose schema they share) before delegating. Regression tests cover the Cayenne wrapper and the DuckDB marker. * fix(cayenne): read legacy inline map data under Arrow 59 and keep the type policy - Inline data written under a nullable map-entries declaration is read through arrow_tools::map_entries, since Arrow 59 refuses that declaration at decode. - The primary-key fast-path test's control query no longer uses a bare SUM, which DataFusion 55 answers from statistics without a scan. - RunEndEncoded columns stay refused: the Vortex pin can store them, but accepting them is a product change, so the drift test records the exception. * fix: bring the remaining tests and connectors in line with DataFusion 55 and Arrow 59 - Flight: repair a server's nullable map-entries declaration before decoding (Arrow 59 refuses it), as the Flight SQL path does. - Federation: SQLite and MySQL refuse a filter on a volatile projection output (DataFusion fork unparser), Turso forwards it, MSSQL opts out. - Declared types: parse the decimal form directly so an out-of-range precision/scale still names the problem under arrow-rs's stricter parser. - Expectations that DataFusion 55 changed: predicate reordering (reorder_q8), binary literals rendered as X'..', sorts moved below projections (vector scans), the NSQL context's version and Spark function list, and the null-aware anti join corrected in Ballista's distributed planner. - json_get tests follow datafusion-functions-json 0.55 (negative array indices, json_as_text for a string cast, integral floats in json_get_int). * fix(vortex): port the merged point-lookup and key-block code to DataFusion 55 and Vortex 0.86 Trunk's point-lookup and key-block changes were written against DataFusion 54 and the earlier Vortex pin. The scan now binds its projection and filter before handing them to Vortex (the reader takes bound expressions), the point read keeps trunk's optimize-then-bind of its conjuncts, dynamic filters are detected through DynamicFilterTracking, and TableSchema is built with From. Two merged tests move to the current Vortex Arrow conversion and DataFusion's non-zero batch_size. * fix(arrow_tools): keep the producer's dictionary ids when repairing a map declaration Repairing a nullable map-entries declaration re-encoded the IPC schema, and re-encoding renumbers dictionaries from zero while the dictionary batches that follow keep the producer's ids. A producer that numbers them differently then failed to decode, or decoded each dictionary column against another's dictionary and returned wrong values with no error (reproduced on the Flight SQL read, the Flight read and the Flight subscribe paths). The repair now patches the schema message in place: one byte per map field (the entries nullability, or the Map-to-List type tag for the decodable form), verified with the flatbuffers verifier and checked against the expected schema; an unexpected layout is refused rather than passed through. * fix(runtime): keep the built-in semantics of functions datafusion-spark 55 would shadow datafusion-spark 55 registers Spark versions of power/pow, atan2 and concat_ws over the built-ins (infinity instead of an error for zero to a negative power, Float64 widening, array flattening) and adds hypot, monthname, quote and weekday. None of that is a deliberate change, so those functions join trunc, date_trunc and date_part in the withheld list, and a test pins exactly which names resolve to a Spark implementation. The NSQL context now lists the functions the session actually registers, so it no longer advertises the withheld ones. * fix(duckdb): keep a binary literal out of SQL sent to DuckDB DataFusion 55's unparser renders a binary literal as X'ff', which DuckDB 1.4.4 reads as the text 'xff': through DuckDB federation, b = X'ff' returned the row holding the bytes 'xff' instead of 0xFF (and X'' missed the empty blob), with no error. Expressions with a binary literal now stay local; a NULL still pushes down. Covered end to end by duckdb_binary_literal_filters_agree_with_local_evaluation. * build: pin DataFusion at 006f3d21, which drops input sums from join statistics DataFusion 55 answers a bare SUM from exact statistics, and inner/outer join statistics carried each input's sum through, so SUM over a join returned the whole input table's sum (silent wrong results; caught by the Cayenne result-correctness parity tests). The fork fix is spiceai/datafusion#239; the ledger records it with those parity tests as its guard. * test(runtime-datafusion): drop a redundant clone in the binary-literal test * test: fix the Spark built-ins table width and report the pinned DataFusion revision - the_built_session_keeps_the_built_in_math_and_string_functions: the expected table padded one column a character wider than the rendered output; the values were already right. - spice-substrait-compliance reports DATAFUSION_FORK_REV, which still named the pre-upgrade revision; it now matches the workspace pin, as its test requires. * test(oracle): re-record the pushdown plan for DataFusion 55's decimal literal display DataFusion 55 prints a decimal literal as Decimal128(123.45,10,2) where 54 printed Decimal128(Some(12345),10,2). The SQL sent to Oracle is unchanged. * build: repin the forks to their merged fix PRs and carry spiceai-54 forward - DataFusion f22d1a74: spiceai-54 forward-merged into the 55 line (spiceai/datafusion#241), bringing spiceai#237 (a Date32 literal's cast uses the dialect's date type; SQLite read CAST('…' AS DATE) as a number, so date ranges pushed to SQLite matched nothing) and metadata-column listing pruning. - arrow-rs be4918e7, Ballista 7c4d54c2, table-providers 3045960a, iceberg 50e982dc, arrow-adbc 674d6184: the merged CI fix PRs. - Guard for spiceai#237: a Date32 range unparsed for SQLite keeps the rows it selects (runs the SQL on SQLite with --features sqlite); it fails on the previous pin. - Ledger rows for both; the substrait harness reports the new DataFusion pin. * build: pin DataFusion at ae719330, the merge of spiceai/datafusion#241 Same tree as f22d1a74 (spiceai-54 forward-merged into the 55 line); now the head of spiceai-55-patches-2 rather than a temporary branch. * style: rustfmt the SQLite date-range federation test * build: pin arrow-rs at its spiceai-59 merge and DataFusion at spiceai-55-patches-3 - arrow-rs: 2e2cc330, the merge of spiceai/arrow-rs#29 into spiceai-59 (same tree as the previous pin). - DataFusion: fc84f2d1 on spiceai-55-patches-3, which carries the Date32 literal fix (spiceai/datafusion#237) but not the metadata-column listing pruning (spiceai#229); the ledger records that pruning as not carried on 55 until it has a guard here. - Correct the spark-connect-rs branch annotation to spiceai-59-patches-2. * build: pin DataFusion at the spiceai-55 merge, and iceberg-rust and snowflake-rs at their arrow-rs bumps - DataFusion: 02550cf9, the merge of spiceai/datafusion#243 into spiceai-55 (same tree as the previous pin); no longer a temporary pin. - iceberg-rust: 0284be3d and snowflake-rs: b95a74e8, which move each fork's own arrow-rs pin to 2e2cc330 (Cargo manifests only; Spice builds the same code). * build: repin ballista, table-providers, iceberg-rust and spark-connect-rs to their branch heads - datafusion-ballista 557ae4ea: arrow-rs/DataFusion landed pins, and the Rust 1.99 clippy fixes (spiceai#70, spiceai#71, spiceai#72). - datafusion-table-providers 5f50cfd5 and iceberg-rust a0bd6b04: arrow-rs/DataFusion landed pins (spiceai#82, spiceai#53). - spark-connect-rs 3bf6421a: regenerated lockfile. datafusion-federation and duckdb-rs stay at 9ca84a39 and 8ee43073: table-providers still pins exactly those revisions, and a second revision of either would put two copies of the crate in the graph. Their newer heads only change their own Cargo manifests and lockfile. * build: pin iceberg-rust, snowflake-rs, spark-connect-rs, spice-rs and spicebench at their merges The fork PRs these pins pointed at have merged, mostly as squash merges, so several pinned commits are no longer on their base branches. Each pin now points at the merge on the fork's base branch: - iceberg-rust spiceai-0.10.1-df-55 @ 3e2a14e (spiceai/iceberg-rust#50, spiceai#54) - snowflake-rs spiceai-59 @ f555738 (spiceai/snowflake-rs#13, spiceai#15) - spark-connect-rs spiceai-59-2 @ 18ae9bd (spiceai/spark-connect-rs#15, spiceai#16) - spice-rs trunk @ 4429f39 (spiceai/spice-rs#103) - spicebench trunk @ 8a90555 (spiceai/spicebench#284) datafusion-federation stays at 9ca84a3, which is on spiceai-55 because spiceai/datafusion-federation#88 was a merge commit; only its branch note changes. What changes for Spice: lint-only rewrites in snowflake-api, and tests in iceberg-datafusion and spark-connect-core. The other differences are in the forks' own [patch.crates-io] sections, which cargo ignores for a dependency. spice-rs's tree is identical. * build: pin the forks at their landed version-branch revisions - arrow-adbc 6e4119ac (spiceai-24, spiceai#7), iceberg-rust 3e2a14e8 (spiceai-0.10.1-df-55, spiceai#50), datafusion-table-providers 465926a3 (spiceai-55, spiceai#80), snowflake-rs f5557381 (spiceai-59, spiceai#13), spark-connect-rs 18ae9bd3 (spiceai-59-2, spiceai#15), spice-rs 4429f395 and spicebench 8a90555d (trunk). - datafusion-federation stays at 9ca84a39, now on spiceai-55 (spiceai#88): datafusion-table-providers pins exactly that revision. - duckdb-rs stays at 8ee43073 and temporary: spiceai#49 was squash-merged, so that revision is not on spiceai-1.4.4, and datafusion-table-providers pins it, so both have to move together. - datafusion-ballista is unchanged; its version PR is still open. * build: pin datafusion-ballista at upstream 55.0.0-rc1 spiceai/datafusion-ballista#73 merged upstream's 55.0.0-rc1 tag into spiceai-55-patches-2, the head of spiceai/datafusion-ballista#68. spiceai#68 is still open, waiting on license-compliance decisions. Pin 718429f, the spiceai#73 merge, until spiceai#68 lands on spiceai-55. ballista-core, -executor, -scheduler, -history and -api-types move from 54.0.0 to 55.0.0; no packages are added or removed in Cargo.lock. `cargo check -p spiced --locked` passes on rustc 1.98.1. * docs(fork_patches): record the datafusion-ballista pin at 718429f4 (upstream 55.0.0-rc1, spiceai#73) * build: pin datafusion-ballista at a7c4c585, the merge of spiceai/datafusion-ballista#68 into spiceai-55 * build: pin datafusion-ballista at the spiceai-55 merge of spiceai#68 spiceai/datafusion-ballista#68 merged into spiceai-55 as a7c4c58, a merge commit whose tree is identical to 718429f, the previous pin (`git diff 718429f a7c4c58` is empty). Repoint the pin at the landed revision and drop the TEMPORARY marker from the fork-patch ledger, as scripts/check_fork_patches.py asks. The ledger's ballista section now names spiceai-55 and upstream 55.0.0-rc1 as its base. Cargo.lock changes only the source revision of the five ballista packages. * perf(cayenne): keep small Vortex scans whole on the main query path DataFusion 55 lowered `repartition_file_min_size` from 10 MiB to 1 MiB (apache/datafusion#22439), so the main Cayenne scan byte-range-splits a 1-10 MiB dimension table `target_partitions` ways and pays a Vortex footer open per range. The main scan now opts small scans out below 10 MiB, as the internal and protected-snapshot scans already do below 256 MiB; Parquet listing scans keep the new default. `small_snapshot_groups_opt_out_of_repartitioning` now asserts the main scan's opt-out and fails with the threshold back at 0. * docs(fork_patches): record duckdb-rs on the protected spiceai-1.4.4-patches-2 head, which table-providers pins * test(duckdb): build CreateExternalTable with DataFusion 55's locations field * ci: validate the datafusion-table-providers pin against spiceai-55, where spiceai#80 landed * test(cayenne): size the point-lookup fan-out control above the main scan's 10 MiB opt-out * fix(data-accelerator-api): implement DataFusion 55's apply_expressions for KeepFirstExec * test: port trunk's new tests to DataFusion 55 (CreateExternalTable locations, common::TableReference) * fix(cayenne): keep -0.0 and 0.0 equal in pushed-down float comparisons and memory-mode indexes DataFusion 55 compares floats with the two zeros equal (apache/datafusion#22835), while Vortex and the Arrow row encoding order them apart. Spell out both zeros when a float comparison or IN list with a zero literal is pushed into a Vortex scan, keep float column-vs-column comparisons above the scan, and hash memory-mode index keys with the file-mode KeyEncoder. * Fix lint --------- Co-authored-by: Luke Kim <80174+lukekim@users.noreply.github.com>
peasee
pushed a commit
to spiceai/spiceai
that referenced
this pull request
Oct 8, 2026
…n fixes backported from 54 (#14800) * build: pin the DataFusion 55 / Arrow 59 fork lines Moves every fork to its DataFusion 55 line: arrow-rs 59.3.0, DataFusion 55.1.0, Ballista (upstream main on DataFusion 55), federation, table-providers, iceberg-rust 0.10.1, duckdb-rs 1.4.4 on arrow 59, ADBC 24, snowflake-rs, spark-connect-rs, spice-rs and spicebench on arrow 59. datafusion-functions-json comes from crates.io (0.55.6); delta-kernel stays on 0.27.1 with its default engine split into delta_kernel_default_engine. docs/dev/fork_patches.md is re-audited against the new lines; the pins stay marked TEMPORARY until the fork pull requests merge. * feat: upgrade to DataFusion 55 and Arrow 59 - Statistics flow through DataFusion 55's StatisticsContext. The deprecated partition_statistics returns Absent for built-in plans on 55, so every wrapper implements child_stats_requests / statistics_from_inputs and the optimizer rules, join sizing and Flight batch sizing read real statistics again. - ExecutionPlan / TableProvider wrappers implement the new required and defaulted methods explicitly (apply_expressions, replace_children, try_to_proto, merge_into, ...), forwarding where they wrap. - Arrow 59 refuses a MAP whose entries field is nullable: Flight SQL reads replace the declared schema before decoding and normalize the map afterwards, and legacy IPC maps decode as a list (arrow_tools map_entries). - Delta Lake refuses Void and interval columns with a structured error instead of mis-reading them. - Ballista 55: Vcores, the protocol-version handshake, task-id based cancellation and the append-only task-info model. - Vortex 0.86 and the remaining API moves (TableReference, TableScanBuilder, TableSchema builder, codec proto converter, WriteOp::MergeInto). * ci: validate the datafusion-table-providers pin against spiceai-55-patches The DataFusion 55 line of the fork is spiceai-55-patches. * revert: drop the unfinished guard tests d65aa29 picked up by accident d65aa29 was meant to change only the CI branch check; it also committed work-in-progress test edits. Restore those files to 98800a8; the finished tests land in their own commit. * fix(duckdb): push regexp_count down NULL-preserving, as DataFusion 55 evaluates it DataFusion 55 (apache/datafusion#24239) makes regexp_count propagate NULL, as PostgreSQL does. The DuckDB rendering still wrapped the count in coalesce(.., 0) to match the old kernel, so a DuckDB-accelerated dataset answered 0 where local evaluation answers NULL, and a WHERE on the count kept rows local evaluation drops. Render it as len(regexp_extract_all(..)), which is NULL for a NULL input or pattern. * test(cayenne): count -0.0 as equal to 0.0 in the boundary-values test DataFusion 55 compares floats in SQL by IEEE 754 value, so float_val = 0.0 matches the -0.0 row too (4 rows, not 3). A MemTable in the same session returns the same rows, so this is the expectation, not Cayenne. * build: repin the forks and guard their new patches - Ballista 8bf37ad0: the executor poll loop accepts a vcore semaphore with no permits yet. Upstream's assert panicked the loop at startup, so no Spice executor registered and every cluster query failed. - DataFusion d98ae645 (MSRV-clean unparser), federation 9ca84a39, table-providers 661d16d1, iceberg 575d81a1: their own tests now run against the Spice DataFusion (and, for table-providers, ADBC) forks. - Guard tests for the unparser rescope, eager aggregation through StatisticsContext, SchemaCastScanExec statistics forwarding and Ballista TaskInfo fields 11/12; the ledger names them and records the Ballista fix. - The distributed EXPLAIN snapshot now reads Ballista 55's UnknownPartitioning(1) where it printed None. * build: point the fork pins at the spiceai-*-patches-2 branches The DataFusion 55 lines of arrow-rs, datafusion, datafusion-ballista, datafusion-federation, datafusion-table-providers, duckdb-rs, snowflake-rs and spark-connect-rs are reviewed on new spiceai-*-patches-2 branches, leaving the earlier spiceai-*-patches branches untouched. Same revisions; only the branch names in the pin comments, the fork-patch ledger and the table-providers pin check change. * fix: re-root a delegated projection on the inner plan before swapping DataFusion 55's projection pushdown reads a projection's input as the node being swapped with: `ProjectionExec` collapses the chain starting at `projection.input()`, `CoalescePartitionsExec` rebuilds from its child. A wrapper that handed its inner plan the projection unchanged (its input being the wrapper) got it back still above the wrapper, and the pushdown then nested another copy on every step: the Cayenne maintained-aggregate soundness test grew to gigabytes and never finished. CayenneAccelerationExec, the DuckDB aggregate pushdown marker and PartitionedUnionExec now rebuild the projection over their inner plan (whose schema they share) before delegating. Regression tests cover the Cayenne wrapper and the DuckDB marker. * fix(cayenne): read legacy inline map data under Arrow 59 and keep the type policy - Inline data written under a nullable map-entries declaration is read through arrow_tools::map_entries, since Arrow 59 refuses that declaration at decode. - The primary-key fast-path test's control query no longer uses a bare SUM, which DataFusion 55 answers from statistics without a scan. - RunEndEncoded columns stay refused: the Vortex pin can store them, but accepting them is a product change, so the drift test records the exception. * fix: bring the remaining tests and connectors in line with DataFusion 55 and Arrow 59 - Flight: repair a server's nullable map-entries declaration before decoding (Arrow 59 refuses it), as the Flight SQL path does. - Federation: SQLite and MySQL refuse a filter on a volatile projection output (DataFusion fork unparser), Turso forwards it, MSSQL opts out. - Declared types: parse the decimal form directly so an out-of-range precision/scale still names the problem under arrow-rs's stricter parser. - Expectations that DataFusion 55 changed: predicate reordering (reorder_q8), binary literals rendered as X'..', sorts moved below projections (vector scans), the NSQL context's version and Spark function list, and the null-aware anti join corrected in Ballista's distributed planner. - json_get tests follow datafusion-functions-json 0.55 (negative array indices, json_as_text for a string cast, integral floats in json_get_int). * fix(vortex): port the merged point-lookup and key-block code to DataFusion 55 and Vortex 0.86 Trunk's point-lookup and key-block changes were written against DataFusion 54 and the earlier Vortex pin. The scan now binds its projection and filter before handing them to Vortex (the reader takes bound expressions), the point read keeps trunk's optimize-then-bind of its conjuncts, dynamic filters are detected through DynamicFilterTracking, and TableSchema is built with From. Two merged tests move to the current Vortex Arrow conversion and DataFusion's non-zero batch_size. * fix(arrow_tools): keep the producer's dictionary ids when repairing a map declaration Repairing a nullable map-entries declaration re-encoded the IPC schema, and re-encoding renumbers dictionaries from zero while the dictionary batches that follow keep the producer's ids. A producer that numbers them differently then failed to decode, or decoded each dictionary column against another's dictionary and returned wrong values with no error (reproduced on the Flight SQL read, the Flight read and the Flight subscribe paths). The repair now patches the schema message in place: one byte per map field (the entries nullability, or the Map-to-List type tag for the decodable form), verified with the flatbuffers verifier and checked against the expected schema; an unexpected layout is refused rather than passed through. * fix(runtime): keep the built-in semantics of functions datafusion-spark 55 would shadow datafusion-spark 55 registers Spark versions of power/pow, atan2 and concat_ws over the built-ins (infinity instead of an error for zero to a negative power, Float64 widening, array flattening) and adds hypot, monthname, quote and weekday. None of that is a deliberate change, so those functions join trunc, date_trunc and date_part in the withheld list, and a test pins exactly which names resolve to a Spark implementation. The NSQL context now lists the functions the session actually registers, so it no longer advertises the withheld ones. * fix(duckdb): keep a binary literal out of SQL sent to DuckDB DataFusion 55's unparser renders a binary literal as X'ff', which DuckDB 1.4.4 reads as the text 'xff': through DuckDB federation, b = X'ff' returned the row holding the bytes 'xff' instead of 0xFF (and X'' missed the empty blob), with no error. Expressions with a binary literal now stay local; a NULL still pushes down. Covered end to end by duckdb_binary_literal_filters_agree_with_local_evaluation. * build: pin DataFusion at 006f3d21, which drops input sums from join statistics DataFusion 55 answers a bare SUM from exact statistics, and inner/outer join statistics carried each input's sum through, so SUM over a join returned the whole input table's sum (silent wrong results; caught by the Cayenne result-correctness parity tests). The fork fix is spiceai/datafusion#239; the ledger records it with those parity tests as its guard. * test(runtime-datafusion): drop a redundant clone in the binary-literal test * test: fix the Spark built-ins table width and report the pinned DataFusion revision - the_built_session_keeps_the_built_in_math_and_string_functions: the expected table padded one column a character wider than the rendered output; the values were already right. - spice-substrait-compliance reports DATAFUSION_FORK_REV, which still named the pre-upgrade revision; it now matches the workspace pin, as its test requires. * test(oracle): re-record the pushdown plan for DataFusion 55's decimal literal display DataFusion 55 prints a decimal literal as Decimal128(123.45,10,2) where 54 printed Decimal128(Some(12345),10,2). The SQL sent to Oracle is unchanged. * build: repin the forks to their merged fix PRs and carry spiceai-54 forward - DataFusion f22d1a74: spiceai-54 forward-merged into the 55 line (spiceai/datafusion#241), bringing #237 (a Date32 literal's cast uses the dialect's date type; SQLite read CAST('…' AS DATE) as a number, so date ranges pushed to SQLite matched nothing) and metadata-column listing pruning. - arrow-rs be4918e7, Ballista 7c4d54c2, table-providers 3045960a, iceberg 50e982dc, arrow-adbc 674d6184: the merged CI fix PRs. - Guard for #237: a Date32 range unparsed for SQLite keeps the rows it selects (runs the SQL on SQLite with --features sqlite); it fails on the previous pin. - Ledger rows for both; the substrait harness reports the new DataFusion pin. * build: pin DataFusion at ae719330, the merge of spiceai/datafusion#241 Same tree as f22d1a74 (spiceai-54 forward-merged into the 55 line); now the head of spiceai-55-patches-2 rather than a temporary branch. * style: rustfmt the SQLite date-range federation test * build: pin arrow-rs at its spiceai-59 merge and DataFusion at spiceai-55-patches-3 - arrow-rs: 2e2cc330, the merge of spiceai/arrow-rs#29 into spiceai-59 (same tree as the previous pin). - DataFusion: fc84f2d1 on spiceai-55-patches-3, which carries the Date32 literal fix (spiceai/datafusion#237) but not the metadata-column listing pruning (#229); the ledger records that pruning as not carried on 55 until it has a guard here. - Correct the spark-connect-rs branch annotation to spiceai-59-patches-2. * build: pin DataFusion at the spiceai-55 merge, and iceberg-rust and snowflake-rs at their arrow-rs bumps - DataFusion: 02550cf9, the merge of spiceai/datafusion#243 into spiceai-55 (same tree as the previous pin); no longer a temporary pin. - iceberg-rust: 0284be3d and snowflake-rs: b95a74e8, which move each fork's own arrow-rs pin to 2e2cc330 (Cargo manifests only; Spice builds the same code). * build: repin ballista, table-providers, iceberg-rust and spark-connect-rs to their branch heads - datafusion-ballista 557ae4ea: arrow-rs/DataFusion landed pins, and the Rust 1.99 clippy fixes (#70, #71, #72). - datafusion-table-providers 5f50cfd5 and iceberg-rust a0bd6b04: arrow-rs/DataFusion landed pins (#82, #53). - spark-connect-rs 3bf6421a: regenerated lockfile. datafusion-federation and duckdb-rs stay at 9ca84a39 and 8ee43073: table-providers still pins exactly those revisions, and a second revision of either would put two copies of the crate in the graph. Their newer heads only change their own Cargo manifests and lockfile. * build: pin iceberg-rust, snowflake-rs, spark-connect-rs, spice-rs and spicebench at their merges The fork PRs these pins pointed at have merged, mostly as squash merges, so several pinned commits are no longer on their base branches. Each pin now points at the merge on the fork's base branch: - iceberg-rust spiceai-0.10.1-df-55 @ 3e2a14e (spiceai/iceberg-rust#50, #54) - snowflake-rs spiceai-59 @ f555738 (spiceai/snowflake-rs#13, #15) - spark-connect-rs spiceai-59-2 @ 18ae9bd (spiceai/spark-connect-rs#15, #16) - spice-rs trunk @ 4429f39 (spiceai/spice-rs#103) - spicebench trunk @ 8a90555 (spiceai/spicebench#284) datafusion-federation stays at 9ca84a3, which is on spiceai-55 because spiceai/datafusion-federation#88 was a merge commit; only its branch note changes. What changes for Spice: lint-only rewrites in snowflake-api, and tests in iceberg-datafusion and spark-connect-core. The other differences are in the forks' own [patch.crates-io] sections, which cargo ignores for a dependency. spice-rs's tree is identical. * build: pin the forks at their landed version-branch revisions - arrow-adbc 6e4119ac (spiceai-24, #7), iceberg-rust 3e2a14e8 (spiceai-0.10.1-df-55, #50), datafusion-table-providers 465926a3 (spiceai-55, #80), snowflake-rs f5557381 (spiceai-59, #13), spark-connect-rs 18ae9bd3 (spiceai-59-2, #15), spice-rs 4429f395 and spicebench 8a90555d (trunk). - datafusion-federation stays at 9ca84a39, now on spiceai-55 (#88): datafusion-table-providers pins exactly that revision. - duckdb-rs stays at 8ee43073 and temporary: #49 was squash-merged, so that revision is not on spiceai-1.4.4, and datafusion-table-providers pins it, so both have to move together. - datafusion-ballista is unchanged; its version PR is still open. * build: pin datafusion-ballista at upstream 55.0.0-rc1 spiceai/datafusion-ballista#73 merged upstream's 55.0.0-rc1 tag into spiceai-55-patches-2, the head of spiceai/datafusion-ballista#68. #68 is still open, waiting on license-compliance decisions. Pin 718429f, the #73 merge, until #68 lands on spiceai-55. ballista-core, -executor, -scheduler, -history and -api-types move from 54.0.0 to 55.0.0; no packages are added or removed in Cargo.lock. `cargo check -p spiced --locked` passes on rustc 1.98.1. * docs(fork_patches): record the datafusion-ballista pin at 718429f4 (upstream 55.0.0-rc1, #73) * build: pin datafusion-ballista at a7c4c585, the merge of spiceai/datafusion-ballista#68 into spiceai-55 * build: pin datafusion-ballista at the spiceai-55 merge of #68 spiceai/datafusion-ballista#68 merged into spiceai-55 as a7c4c58, a merge commit whose tree is identical to 718429f, the previous pin (`git diff 718429f a7c4c58` is empty). Repoint the pin at the landed revision and drop the TEMPORARY marker from the fork-patch ledger, as scripts/check_fork_patches.py asks. The ledger's ballista section now names spiceai-55 and upstream 55.0.0-rc1 as its base. Cargo.lock changes only the source revision of the five ballista packages. * perf(cayenne): keep small Vortex scans whole on the main query path DataFusion 55 lowered `repartition_file_min_size` from 10 MiB to 1 MiB (apache/datafusion#22439), so the main Cayenne scan byte-range-splits a 1-10 MiB dimension table `target_partitions` ways and pays a Vortex footer open per range. The main scan now opts small scans out below 10 MiB, as the internal and protected-snapshot scans already do below 256 MiB; Parquet listing scans keep the new default. `small_snapshot_groups_opt_out_of_repartitioning` now asserts the main scan's opt-out and fails with the threshold back at 0. * docs(fork_patches): record duckdb-rs on the protected spiceai-1.4.4-patches-2 head, which table-providers pins * test(duckdb): build CreateExternalTable with DataFusion 55's locations field * ci: validate the datafusion-table-providers pin against spiceai-55, where #80 landed * test(cayenne): size the point-lookup fan-out control above the main scan's 10 MiB opt-out * fix(data-accelerator-api): implement DataFusion 55's apply_expressions for KeepFirstExec * test: port trunk's new tests to DataFusion 55 (CreateExternalTable locations, common::TableReference) * fix(cayenne): keep -0.0 and 0.0 equal in pushed-down float comparisons and memory-mode indexes DataFusion 55 compares floats with the two zeros equal (apache/datafusion#22835), while Vortex and the Arrow row encoding order them apart. Spell out both zeros when a float comparison or IN list with a zero literal is pushed into a Vortex scan, keep float column-vs-column comparisons above the scan, and hash memory-mode index keys with the file-mode KeyEncoder. * build: move to iceberg-rust 0.11.0 (RC) and opendal 0.58 Points the iceberg crates at spiceai/iceberg-rust#55, which merges the upstream 0.11.x release branch onto the DataFusion 55 fork line. opendal moves to 0.58 to match iceberg-storage-opendal, so one opendal stack is built instead of two. 0.58 removes the OperatorBuilder step: Operator::new returns the operator with the same default layers that finish() applied. * Bump datafusion to 55.2-rc1 * build: pin iceberg-rust at spiceai-0.11.0-df-55 (6317d03e) spiceai/iceberg-rust#55 merged into spiceai-0.11.0-df-55. Same Rust code as the previous pin (e02547cb); the branch adds only a Python CI pin and a doc-comment fix. fork_patches.md records the new branch and that SigV4 is now an AuthManager in the fork's REST catalog. * build: keep the lockfile's unrelated resolutions from the DataFusion bump The previous commit's lockfile update also re-resolved five crates with range requirements (hashbrown, heck, base64). Restore them so the iceberg repin changes only the iceberg-rust source lines. * build: pin iceberg-rust at spiceai-0.11.0-df-55 (7735caa2) Picks up spiceai/iceberg-rust#49 (IcebergTableProvider::catalog and ::table_ident) and #56 (the fork's own DataFusion pin, which does not affect this build). #49's fork_patches row and its compile guard land with #14588, which first calls the getters. * build: pin datafusion at spiceai/datafusion#250 (55.2.0-rc1 and the 54 backports) Moves the DataFusion fork from 02550cf9 (55.1.0) to 1749ce4e, the head of spiceai/datafusion#250, which merges into #249: `spiceai-55` at 55.2.0-rc1 (spiceai/datafusion#248) plus the `spiceai-54` patches #233, #234 and #245 carried onto 55, and #250's review fix. The guards for those follow in the next commit. The pin names a pull request branch, so the ledger row is marked TEMPORARY and `scripts/check_fork_patches.py` refuses to let it land until #250 and #249 merge and the pin moves to the `spiceai-55` merge commit. The 55.2.0-rc1 merge changes no Spice patch. Outside `Cargo.lock`, the fork's 02550cf9..27351ab0 diff has the same patch-id as upstream's 55.1.0..55.2.0-rc1 diff, except `datafusion-spark`'s `quote.rs`, which the fork's backport had already made identical to 55.2.0-rc1's. That backport's ledger row is dropped; upstream carries the fix as apache/datafusion#25277. `Cargo.lock` changes only the 36 DataFusion entries. * test: guard the DataFusion unparser and footer-cache fixes backported from 54 Each backported fix gets a guard here, so the next re-cut of the fork cannot drop it silently, plus a ledger row naming the guard. - spiceai/datafusion#233 (#14373): a join that is another join's right input stays a parenthesised joined table on that join's right. `a_join_that_is_another_joins_right_input_stays_on_its_right` checks that `b` is introduced before the outer `ON` names it and that a LEFT JOIN keeps a filter from inside it out of `WHERE`. With the `sqlite` feature it runs the SQL on SQLite and compares the rows with DataFusion's for the same plan. - spiceai/datafusion#234 (#14375): a `Limit` that is a join input bounds that input, not the join. `a_limit_on_a_join_input_bounds_that_input_rather_than_the_join`, with the same row comparison. - spiceai/datafusion#250: a projection over a join that passes up a mark join's mark is refused rather than unparsed with the mark unbound. `a_mark_a_projection_over_a_join_passes_up_is_refused`. - spiceai/datafusion#245 (#12952): the file metadata cache frees an evicted footer's hit counter. 55's `DefaultCache` already does, so 55 carries no patch; `crates/cayenne/tests/footer_cache_hit_counter_test.rs` pins the behaviour by counting the live heap across 20,000 cold footer reads into a full cache. At 02550cf9 the three unparser guards fail. SQLite rejects the nested join's SQL ("ON clause references tables to its right"), returns 1 row where the plan returns 2 for `b FULL JOIN (c LIMIT 1)`, and the mark join renders as `… LEFT OUTER JOIN (b) ON a.id = b.id WHERE NOT c.mark`. Against a local DataFusion build with the hit-counter prune removed, the footer-cache guard measures 2,646,176 bytes of heap growth. At 1749ce4e all four pass, and the footer-cache growth is 0 bytes. * build: pin datafusion at spiceai/datafusion#249's merge commit Moves the fork pin from #250's branch head to bebc4d58, the merge of #249 into `spiceai-55`. Its tree is #249's head, cbda233a: 55.2.0-rc1 plus the backported #233, #234 and #245. The pin names a landed revision, so the TEMPORARY marker comes off and `scripts/check_fork_patches.py` passes. #250 is not part of this revision, so its guard and ledger row are removed. * Update datafusion version to 55.2 in Cargo.toml * docs(fork_patches): record the unparser frame refactor #249 carries `cbda233a6` is the fourth Spice commit at the pinned revision. It moves the join-input bookkeeping that #233 and #234 add out of `select_to_sql_recursively_inner`, so an unoptimised build's frame stays where it was. It had no ledger row, so a re-cut could drop it without the audit noticing. Its row is a GAP, with an Open gaps entry. What it buys is stack headroom in a debug build: without it, the fork's `roundtrip_statement` needs 2,112 KiB and overflows a 2 MiB test thread. This repo runs tests on 8 MiB threads, so no test here sees the difference deterministically. The entry names the fork-side check to run at a re-cut instead. * Fix tests * build: pin DataFusion at spiceai/datafusion#251, so MAX of a partition column skips empty partitions (#14801) * build: pin datafusion at spiceai/datafusion#251's merge commit Moves the DataFusion fork pin from #249's merge commit to the head of `spiceai-55`, which adds two patches on top of it: - #251: a file that holds no rows gets no partition-column min/max. Without it, `MIN`/`MAX` of a partition column is answered from the listing's statistics with the value of a partition whose files are all empty: `max(p)` answers '99' for a `p=99` directory holding one empty file, where the rows give '3'. - #240: the listing prunes files by metadata-column predicates before opening them and reports those filters `Exact` (#229 on `spiceai-54`). Neither changes a manifest, so `Cargo.lock` only moves its source revision. Each patch gets a ledger row and a guard in the listing connector's tests, which the gate runs as library tests. * fix(listing): keep every predicate residual on the `_location` fast path The pin now carries spiceai/datafusion#240, so the inner listing reports a metadata-column predicate such as `_size < 50` `Exact`. `MetadataPruningListingTable` handed pushdown classification to the inner listing whenever a `_location` predicate was present, but its `_location` fast path opens the named objects and applies no other predicate, so `_location = x AND _size < 50` returned the rows of a file the `_size` predicate excludes. Report every predicate `Inexact` on that path as well, so a residual `FilterExec` re-applies it. This is #14790's change, which makes the same fix for partition predicates, taken byte for byte so that whichever of the two lands second merges cleanly. The new test goes through `create_listing_table` with `_location` and `_size` enabled. * test(metadata): expect the listing to answer `_size = 2319` exactly With spiceai/datafusion#240 the listing prunes files by a metadata-column predicate and reports it `Exact`, so the plan for `met_size`, which enables only `_size`, has no residual `Filter`. The scan lists only the four 2319-byte files; `data_0.parquet` is 2317 bytes, as its S3 `Content-Length` shows. Regenerated by running `metadata::s3_metadata_columns` and accepted with `cargo insta accept`. The rest of the test passes unchanged at this pin, including the `_location` plans and the returned rows. * fix(data-connector-api): satisfy pedantic clippy in the metadata-pruning listing test `make lint-rust`'s `--tests` pass rejected two lints in `listing_table_prunes_files_by_a_metadata_column_predicate`: `redundant_closure_for_method_calls` (`|group| group.len()`) and `format_collect` (`map(format!).collect::<String>()`). Use `FileGroup::len` and build the file contents with `writeln!` into one `String`. The test is unchanged in what it checks. --------- Co-authored-by: Luke Kim <80174+lukekim@users.noreply.github.com> Co-authored-by: Sergei Grebnov <sergei.grebnov@gmail.com>
wangyusheng1985
pushed a commit
to wangyusheng1985/spiceai
that referenced
this pull request
Oct 8, 2026
…n fixes backported from 54 (spiceai#14800) * build: pin the DataFusion 55 / Arrow 59 fork lines Moves every fork to its DataFusion 55 line: arrow-rs 59.3.0, DataFusion 55.1.0, Ballista (upstream main on DataFusion 55), federation, table-providers, iceberg-rust 0.10.1, duckdb-rs 1.4.4 on arrow 59, ADBC 24, snowflake-rs, spark-connect-rs, spice-rs and spicebench on arrow 59. datafusion-functions-json comes from crates.io (0.55.6); delta-kernel stays on 0.27.1 with its default engine split into delta_kernel_default_engine. docs/dev/fork_patches.md is re-audited against the new lines; the pins stay marked TEMPORARY until the fork pull requests merge. * feat: upgrade to DataFusion 55 and Arrow 59 - Statistics flow through DataFusion 55's StatisticsContext. The deprecated partition_statistics returns Absent for built-in plans on 55, so every wrapper implements child_stats_requests / statistics_from_inputs and the optimizer rules, join sizing and Flight batch sizing read real statistics again. - ExecutionPlan / TableProvider wrappers implement the new required and defaulted methods explicitly (apply_expressions, replace_children, try_to_proto, merge_into, ...), forwarding where they wrap. - Arrow 59 refuses a MAP whose entries field is nullable: Flight SQL reads replace the declared schema before decoding and normalize the map afterwards, and legacy IPC maps decode as a list (arrow_tools map_entries). - Delta Lake refuses Void and interval columns with a structured error instead of mis-reading them. - Ballista 55: Vcores, the protocol-version handshake, task-id based cancellation and the append-only task-info model. - Vortex 0.86 and the remaining API moves (TableReference, TableScanBuilder, TableSchema builder, codec proto converter, WriteOp::MergeInto). * ci: validate the datafusion-table-providers pin against spiceai-55-patches The DataFusion 55 line of the fork is spiceai-55-patches. * revert: drop the unfinished guard tests d65aa29 picked up by accident d65aa29 was meant to change only the CI branch check; it also committed work-in-progress test edits. Restore those files to 98800a8; the finished tests land in their own commit. * fix(duckdb): push regexp_count down NULL-preserving, as DataFusion 55 evaluates it DataFusion 55 (apache/datafusion#24239) makes regexp_count propagate NULL, as PostgreSQL does. The DuckDB rendering still wrapped the count in coalesce(.., 0) to match the old kernel, so a DuckDB-accelerated dataset answered 0 where local evaluation answers NULL, and a WHERE on the count kept rows local evaluation drops. Render it as len(regexp_extract_all(..)), which is NULL for a NULL input or pattern. * test(cayenne): count -0.0 as equal to 0.0 in the boundary-values test DataFusion 55 compares floats in SQL by IEEE 754 value, so float_val = 0.0 matches the -0.0 row too (4 rows, not 3). A MemTable in the same session returns the same rows, so this is the expectation, not Cayenne. * build: repin the forks and guard their new patches - Ballista 8bf37ad0: the executor poll loop accepts a vcore semaphore with no permits yet. Upstream's assert panicked the loop at startup, so no Spice executor registered and every cluster query failed. - DataFusion d98ae645 (MSRV-clean unparser), federation 9ca84a39, table-providers 661d16d1, iceberg 575d81a1: their own tests now run against the Spice DataFusion (and, for table-providers, ADBC) forks. - Guard tests for the unparser rescope, eager aggregation through StatisticsContext, SchemaCastScanExec statistics forwarding and Ballista TaskInfo fields 11/12; the ledger names them and records the Ballista fix. - The distributed EXPLAIN snapshot now reads Ballista 55's UnknownPartitioning(1) where it printed None. * build: point the fork pins at the spiceai-*-patches-2 branches The DataFusion 55 lines of arrow-rs, datafusion, datafusion-ballista, datafusion-federation, datafusion-table-providers, duckdb-rs, snowflake-rs and spark-connect-rs are reviewed on new spiceai-*-patches-2 branches, leaving the earlier spiceai-*-patches branches untouched. Same revisions; only the branch names in the pin comments, the fork-patch ledger and the table-providers pin check change. * fix: re-root a delegated projection on the inner plan before swapping DataFusion 55's projection pushdown reads a projection's input as the node being swapped with: `ProjectionExec` collapses the chain starting at `projection.input()`, `CoalescePartitionsExec` rebuilds from its child. A wrapper that handed its inner plan the projection unchanged (its input being the wrapper) got it back still above the wrapper, and the pushdown then nested another copy on every step: the Cayenne maintained-aggregate soundness test grew to gigabytes and never finished. CayenneAccelerationExec, the DuckDB aggregate pushdown marker and PartitionedUnionExec now rebuild the projection over their inner plan (whose schema they share) before delegating. Regression tests cover the Cayenne wrapper and the DuckDB marker. * fix(cayenne): read legacy inline map data under Arrow 59 and keep the type policy - Inline data written under a nullable map-entries declaration is read through arrow_tools::map_entries, since Arrow 59 refuses that declaration at decode. - The primary-key fast-path test's control query no longer uses a bare SUM, which DataFusion 55 answers from statistics without a scan. - RunEndEncoded columns stay refused: the Vortex pin can store them, but accepting them is a product change, so the drift test records the exception. * fix: bring the remaining tests and connectors in line with DataFusion 55 and Arrow 59 - Flight: repair a server's nullable map-entries declaration before decoding (Arrow 59 refuses it), as the Flight SQL path does. - Federation: SQLite and MySQL refuse a filter on a volatile projection output (DataFusion fork unparser), Turso forwards it, MSSQL opts out. - Declared types: parse the decimal form directly so an out-of-range precision/scale still names the problem under arrow-rs's stricter parser. - Expectations that DataFusion 55 changed: predicate reordering (reorder_q8), binary literals rendered as X'..', sorts moved below projections (vector scans), the NSQL context's version and Spark function list, and the null-aware anti join corrected in Ballista's distributed planner. - json_get tests follow datafusion-functions-json 0.55 (negative array indices, json_as_text for a string cast, integral floats in json_get_int). * fix(vortex): port the merged point-lookup and key-block code to DataFusion 55 and Vortex 0.86 Trunk's point-lookup and key-block changes were written against DataFusion 54 and the earlier Vortex pin. The scan now binds its projection and filter before handing them to Vortex (the reader takes bound expressions), the point read keeps trunk's optimize-then-bind of its conjuncts, dynamic filters are detected through DynamicFilterTracking, and TableSchema is built with From. Two merged tests move to the current Vortex Arrow conversion and DataFusion's non-zero batch_size. * fix(arrow_tools): keep the producer's dictionary ids when repairing a map declaration Repairing a nullable map-entries declaration re-encoded the IPC schema, and re-encoding renumbers dictionaries from zero while the dictionary batches that follow keep the producer's ids. A producer that numbers them differently then failed to decode, or decoded each dictionary column against another's dictionary and returned wrong values with no error (reproduced on the Flight SQL read, the Flight read and the Flight subscribe paths). The repair now patches the schema message in place: one byte per map field (the entries nullability, or the Map-to-List type tag for the decodable form), verified with the flatbuffers verifier and checked against the expected schema; an unexpected layout is refused rather than passed through. * fix(runtime): keep the built-in semantics of functions datafusion-spark 55 would shadow datafusion-spark 55 registers Spark versions of power/pow, atan2 and concat_ws over the built-ins (infinity instead of an error for zero to a negative power, Float64 widening, array flattening) and adds hypot, monthname, quote and weekday. None of that is a deliberate change, so those functions join trunc, date_trunc and date_part in the withheld list, and a test pins exactly which names resolve to a Spark implementation. The NSQL context now lists the functions the session actually registers, so it no longer advertises the withheld ones. * fix(duckdb): keep a binary literal out of SQL sent to DuckDB DataFusion 55's unparser renders a binary literal as X'ff', which DuckDB 1.4.4 reads as the text 'xff': through DuckDB federation, b = X'ff' returned the row holding the bytes 'xff' instead of 0xFF (and X'' missed the empty blob), with no error. Expressions with a binary literal now stay local; a NULL still pushes down. Covered end to end by duckdb_binary_literal_filters_agree_with_local_evaluation. * build: pin DataFusion at 006f3d21, which drops input sums from join statistics DataFusion 55 answers a bare SUM from exact statistics, and inner/outer join statistics carried each input's sum through, so SUM over a join returned the whole input table's sum (silent wrong results; caught by the Cayenne result-correctness parity tests). The fork fix is spiceai/datafusion#239; the ledger records it with those parity tests as its guard. * test(runtime-datafusion): drop a redundant clone in the binary-literal test * test: fix the Spark built-ins table width and report the pinned DataFusion revision - the_built_session_keeps_the_built_in_math_and_string_functions: the expected table padded one column a character wider than the rendered output; the values were already right. - spice-substrait-compliance reports DATAFUSION_FORK_REV, which still named the pre-upgrade revision; it now matches the workspace pin, as its test requires. * test(oracle): re-record the pushdown plan for DataFusion 55's decimal literal display DataFusion 55 prints a decimal literal as Decimal128(123.45,10,2) where 54 printed Decimal128(Some(12345),10,2). The SQL sent to Oracle is unchanged. * build: repin the forks to their merged fix PRs and carry spiceai-54 forward - DataFusion f22d1a74: spiceai-54 forward-merged into the 55 line (spiceai/datafusion#241), bringing spiceai#237 (a Date32 literal's cast uses the dialect's date type; SQLite read CAST('…' AS DATE) as a number, so date ranges pushed to SQLite matched nothing) and metadata-column listing pruning. - arrow-rs be4918e7, Ballista 7c4d54c2, table-providers 3045960a, iceberg 50e982dc, arrow-adbc 674d6184: the merged CI fix PRs. - Guard for spiceai#237: a Date32 range unparsed for SQLite keeps the rows it selects (runs the SQL on SQLite with --features sqlite); it fails on the previous pin. - Ledger rows for both; the substrait harness reports the new DataFusion pin. * build: pin DataFusion at ae719330, the merge of spiceai/datafusion#241 Same tree as f22d1a74 (spiceai-54 forward-merged into the 55 line); now the head of spiceai-55-patches-2 rather than a temporary branch. * style: rustfmt the SQLite date-range federation test * build: pin arrow-rs at its spiceai-59 merge and DataFusion at spiceai-55-patches-3 - arrow-rs: 2e2cc330, the merge of spiceai/arrow-rs#29 into spiceai-59 (same tree as the previous pin). - DataFusion: fc84f2d1 on spiceai-55-patches-3, which carries the Date32 literal fix (spiceai/datafusion#237) but not the metadata-column listing pruning (spiceai#229); the ledger records that pruning as not carried on 55 until it has a guard here. - Correct the spark-connect-rs branch annotation to spiceai-59-patches-2. * build: pin DataFusion at the spiceai-55 merge, and iceberg-rust and snowflake-rs at their arrow-rs bumps - DataFusion: 02550cf9, the merge of spiceai/datafusion#243 into spiceai-55 (same tree as the previous pin); no longer a temporary pin. - iceberg-rust: 0284be3d and snowflake-rs: b95a74e8, which move each fork's own arrow-rs pin to 2e2cc330 (Cargo manifests only; Spice builds the same code). * build: repin ballista, table-providers, iceberg-rust and spark-connect-rs to their branch heads - datafusion-ballista 557ae4ea: arrow-rs/DataFusion landed pins, and the Rust 1.99 clippy fixes (spiceai#70, spiceai#71, spiceai#72). - datafusion-table-providers 5f50cfd5 and iceberg-rust a0bd6b04: arrow-rs/DataFusion landed pins (spiceai#82, spiceai#53). - spark-connect-rs 3bf6421a: regenerated lockfile. datafusion-federation and duckdb-rs stay at 9ca84a39 and 8ee43073: table-providers still pins exactly those revisions, and a second revision of either would put two copies of the crate in the graph. Their newer heads only change their own Cargo manifests and lockfile. * build: pin iceberg-rust, snowflake-rs, spark-connect-rs, spice-rs and spicebench at their merges The fork PRs these pins pointed at have merged, mostly as squash merges, so several pinned commits are no longer on their base branches. Each pin now points at the merge on the fork's base branch: - iceberg-rust spiceai-0.10.1-df-55 @ 3e2a14e (spiceai/iceberg-rust#50, spiceai#54) - snowflake-rs spiceai-59 @ f555738 (spiceai/snowflake-rs#13, spiceai#15) - spark-connect-rs spiceai-59-2 @ 18ae9bd (spiceai/spark-connect-rs#15, spiceai#16) - spice-rs trunk @ 4429f39 (spiceai/spice-rs#103) - spicebench trunk @ 8a90555 (spiceai/spicebench#284) datafusion-federation stays at 9ca84a3, which is on spiceai-55 because spiceai/datafusion-federation#88 was a merge commit; only its branch note changes. What changes for Spice: lint-only rewrites in snowflake-api, and tests in iceberg-datafusion and spark-connect-core. The other differences are in the forks' own [patch.crates-io] sections, which cargo ignores for a dependency. spice-rs's tree is identical. * build: pin the forks at their landed version-branch revisions - arrow-adbc 6e4119ac (spiceai-24, spiceai#7), iceberg-rust 3e2a14e8 (spiceai-0.10.1-df-55, spiceai#50), datafusion-table-providers 465926a3 (spiceai-55, spiceai#80), snowflake-rs f5557381 (spiceai-59, spiceai#13), spark-connect-rs 18ae9bd3 (spiceai-59-2, spiceai#15), spice-rs 4429f395 and spicebench 8a90555d (trunk). - datafusion-federation stays at 9ca84a39, now on spiceai-55 (spiceai#88): datafusion-table-providers pins exactly that revision. - duckdb-rs stays at 8ee43073 and temporary: spiceai#49 was squash-merged, so that revision is not on spiceai-1.4.4, and datafusion-table-providers pins it, so both have to move together. - datafusion-ballista is unchanged; its version PR is still open. * build: pin datafusion-ballista at upstream 55.0.0-rc1 spiceai/datafusion-ballista#73 merged upstream's 55.0.0-rc1 tag into spiceai-55-patches-2, the head of spiceai/datafusion-ballista#68. spiceai#68 is still open, waiting on license-compliance decisions. Pin 718429f, the spiceai#73 merge, until spiceai#68 lands on spiceai-55. ballista-core, -executor, -scheduler, -history and -api-types move from 54.0.0 to 55.0.0; no packages are added or removed in Cargo.lock. `cargo check -p spiced --locked` passes on rustc 1.98.1. * docs(fork_patches): record the datafusion-ballista pin at 718429f4 (upstream 55.0.0-rc1, spiceai#73) * build: pin datafusion-ballista at a7c4c585, the merge of spiceai/datafusion-ballista#68 into spiceai-55 * build: pin datafusion-ballista at the spiceai-55 merge of spiceai#68 spiceai/datafusion-ballista#68 merged into spiceai-55 as a7c4c58, a merge commit whose tree is identical to 718429f, the previous pin (`git diff 718429f a7c4c58` is empty). Repoint the pin at the landed revision and drop the TEMPORARY marker from the fork-patch ledger, as scripts/check_fork_patches.py asks. The ledger's ballista section now names spiceai-55 and upstream 55.0.0-rc1 as its base. Cargo.lock changes only the source revision of the five ballista packages. * perf(cayenne): keep small Vortex scans whole on the main query path DataFusion 55 lowered `repartition_file_min_size` from 10 MiB to 1 MiB (apache/datafusion#22439), so the main Cayenne scan byte-range-splits a 1-10 MiB dimension table `target_partitions` ways and pays a Vortex footer open per range. The main scan now opts small scans out below 10 MiB, as the internal and protected-snapshot scans already do below 256 MiB; Parquet listing scans keep the new default. `small_snapshot_groups_opt_out_of_repartitioning` now asserts the main scan's opt-out and fails with the threshold back at 0. * docs(fork_patches): record duckdb-rs on the protected spiceai-1.4.4-patches-2 head, which table-providers pins * test(duckdb): build CreateExternalTable with DataFusion 55's locations field * ci: validate the datafusion-table-providers pin against spiceai-55, where spiceai#80 landed * test(cayenne): size the point-lookup fan-out control above the main scan's 10 MiB opt-out * fix(data-accelerator-api): implement DataFusion 55's apply_expressions for KeepFirstExec * test: port trunk's new tests to DataFusion 55 (CreateExternalTable locations, common::TableReference) * fix(cayenne): keep -0.0 and 0.0 equal in pushed-down float comparisons and memory-mode indexes DataFusion 55 compares floats with the two zeros equal (apache/datafusion#22835), while Vortex and the Arrow row encoding order them apart. Spell out both zeros when a float comparison or IN list with a zero literal is pushed into a Vortex scan, keep float column-vs-column comparisons above the scan, and hash memory-mode index keys with the file-mode KeyEncoder. * build: move to iceberg-rust 0.11.0 (RC) and opendal 0.58 Points the iceberg crates at spiceai/iceberg-rust#55, which merges the upstream 0.11.x release branch onto the DataFusion 55 fork line. opendal moves to 0.58 to match iceberg-storage-opendal, so one opendal stack is built instead of two. 0.58 removes the OperatorBuilder step: Operator::new returns the operator with the same default layers that finish() applied. * Bump datafusion to 55.2-rc1 * build: pin iceberg-rust at spiceai-0.11.0-df-55 (6317d03e) spiceai/iceberg-rust#55 merged into spiceai-0.11.0-df-55. Same Rust code as the previous pin (e02547cb); the branch adds only a Python CI pin and a doc-comment fix. fork_patches.md records the new branch and that SigV4 is now an AuthManager in the fork's REST catalog. * build: keep the lockfile's unrelated resolutions from the DataFusion bump The previous commit's lockfile update also re-resolved five crates with range requirements (hashbrown, heck, base64). Restore them so the iceberg repin changes only the iceberg-rust source lines. * build: pin iceberg-rust at spiceai-0.11.0-df-55 (7735caa2) Picks up spiceai/iceberg-rust#49 (IcebergTableProvider::catalog and ::table_ident) and spiceai#56 (the fork's own DataFusion pin, which does not affect this build). spiceai#49's fork_patches row and its compile guard land with spiceai#14588, which first calls the getters. * build: pin datafusion at spiceai/datafusion#250 (55.2.0-rc1 and the 54 backports) Moves the DataFusion fork from 02550cf9 (55.1.0) to 1749ce4e, the head of spiceai/datafusion#250, which merges into spiceai#249: `spiceai-55` at 55.2.0-rc1 (spiceai/datafusion#248) plus the `spiceai-54` patches spiceai#233, spiceai#234 and spiceai#245 carried onto 55, and spiceai#250's review fix. The guards for those follow in the next commit. The pin names a pull request branch, so the ledger row is marked TEMPORARY and `scripts/check_fork_patches.py` refuses to let it land until spiceai#250 and spiceai#249 merge and the pin moves to the `spiceai-55` merge commit. The 55.2.0-rc1 merge changes no Spice patch. Outside `Cargo.lock`, the fork's 02550cf9..27351ab0 diff has the same patch-id as upstream's 55.1.0..55.2.0-rc1 diff, except `datafusion-spark`'s `quote.rs`, which the fork's backport had already made identical to 55.2.0-rc1's. That backport's ledger row is dropped; upstream carries the fix as apache/datafusion#25277. `Cargo.lock` changes only the 36 DataFusion entries. * test: guard the DataFusion unparser and footer-cache fixes backported from 54 Each backported fix gets a guard here, so the next re-cut of the fork cannot drop it silently, plus a ledger row naming the guard. - spiceai/datafusion#233 (spiceai#14373): a join that is another join's right input stays a parenthesised joined table on that join's right. `a_join_that_is_another_joins_right_input_stays_on_its_right` checks that `b` is introduced before the outer `ON` names it and that a LEFT JOIN keeps a filter from inside it out of `WHERE`. With the `sqlite` feature it runs the SQL on SQLite and compares the rows with DataFusion's for the same plan. - spiceai/datafusion#234 (spiceai#14375): a `Limit` that is a join input bounds that input, not the join. `a_limit_on_a_join_input_bounds_that_input_rather_than_the_join`, with the same row comparison. - spiceai/datafusion#250: a projection over a join that passes up a mark join's mark is refused rather than unparsed with the mark unbound. `a_mark_a_projection_over_a_join_passes_up_is_refused`. - spiceai/datafusion#245 (spiceai#12952): the file metadata cache frees an evicted footer's hit counter. 55's `DefaultCache` already does, so 55 carries no patch; `crates/cayenne/tests/footer_cache_hit_counter_test.rs` pins the behaviour by counting the live heap across 20,000 cold footer reads into a full cache. At 02550cf9 the three unparser guards fail. SQLite rejects the nested join's SQL ("ON clause references tables to its right"), returns 1 row where the plan returns 2 for `b FULL JOIN (c LIMIT 1)`, and the mark join renders as `… LEFT OUTER JOIN (b) ON a.id = b.id WHERE NOT c.mark`. Against a local DataFusion build with the hit-counter prune removed, the footer-cache guard measures 2,646,176 bytes of heap growth. At 1749ce4e all four pass, and the footer-cache growth is 0 bytes. * build: pin datafusion at spiceai/datafusion#249's merge commit Moves the fork pin from spiceai#250's branch head to bebc4d58, the merge of spiceai#249 into `spiceai-55`. Its tree is spiceai#249's head, cbda233a: 55.2.0-rc1 plus the backported spiceai#233, spiceai#234 and spiceai#245. The pin names a landed revision, so the TEMPORARY marker comes off and `scripts/check_fork_patches.py` passes. spiceai#250 is not part of this revision, so its guard and ledger row are removed. * Update datafusion version to 55.2 in Cargo.toml * docs(fork_patches): record the unparser frame refactor spiceai#249 carries `cbda233a6` is the fourth Spice commit at the pinned revision. It moves the join-input bookkeeping that spiceai#233 and spiceai#234 add out of `select_to_sql_recursively_inner`, so an unoptimised build's frame stays where it was. It had no ledger row, so a re-cut could drop it without the audit noticing. Its row is a GAP, with an Open gaps entry. What it buys is stack headroom in a debug build: without it, the fork's `roundtrip_statement` needs 2,112 KiB and overflows a 2 MiB test thread. This repo runs tests on 8 MiB threads, so no test here sees the difference deterministically. The entry names the fork-side check to run at a re-cut instead. * Fix tests * build: pin DataFusion at spiceai/datafusion#251, so MAX of a partition column skips empty partitions (spiceai#14801) * build: pin datafusion at spiceai/datafusion#251's merge commit Moves the DataFusion fork pin from spiceai#249's merge commit to the head of `spiceai-55`, which adds two patches on top of it: - spiceai#251: a file that holds no rows gets no partition-column min/max. Without it, `MIN`/`MAX` of a partition column is answered from the listing's statistics with the value of a partition whose files are all empty: `max(p)` answers '99' for a `p=99` directory holding one empty file, where the rows give '3'. - spiceai#240: the listing prunes files by metadata-column predicates before opening them and reports those filters `Exact` (spiceai#229 on `spiceai-54`). Neither changes a manifest, so `Cargo.lock` only moves its source revision. Each patch gets a ledger row and a guard in the listing connector's tests, which the gate runs as library tests. * fix(listing): keep every predicate residual on the `_location` fast path The pin now carries spiceai/datafusion#240, so the inner listing reports a metadata-column predicate such as `_size < 50` `Exact`. `MetadataPruningListingTable` handed pushdown classification to the inner listing whenever a `_location` predicate was present, but its `_location` fast path opens the named objects and applies no other predicate, so `_location = x AND _size < 50` returned the rows of a file the `_size` predicate excludes. Report every predicate `Inexact` on that path as well, so a residual `FilterExec` re-applies it. This is spiceai#14790's change, which makes the same fix for partition predicates, taken byte for byte so that whichever of the two lands second merges cleanly. The new test goes through `create_listing_table` with `_location` and `_size` enabled. * test(metadata): expect the listing to answer `_size = 2319` exactly With spiceai/datafusion#240 the listing prunes files by a metadata-column predicate and reports it `Exact`, so the plan for `met_size`, which enables only `_size`, has no residual `Filter`. The scan lists only the four 2319-byte files; `data_0.parquet` is 2317 bytes, as its S3 `Content-Length` shows. Regenerated by running `metadata::s3_metadata_columns` and accepted with `cargo insta accept`. The rest of the test passes unchanged at this pin, including the `_location` plans and the returned rows. * fix(data-connector-api): satisfy pedantic clippy in the metadata-pruning listing test `make lint-rust`'s `--tests` pass rejected two lints in `listing_table_prunes_files_by_a_metadata_column_predicate`: `redundant_closure_for_method_calls` (`|group| group.len()`) and `format_collect` (`map(format!).collect::<String>()`). Use `FileGroup::len` and build the file contents with `writeln!` into one `String`. The test is unchanged in what it checks. --------- Co-authored-by: Luke Kim <80174+lukekim@users.noreply.github.com> Co-authored-by: Sergei Grebnov <sergei.grebnov@gmail.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.

Merges upstream DataFusion
55.1.0into the Spice line (spiceai-55, cut fromspiceai-54); conflicts are resolved in the merge commit and recorded in its message.On top of the merge:
sum_values through inner and outer joins, soSUM(col)over a join is not answered from statistics with the input table's sum (fix(physical-plan): drop input sums from inner and outer join statistics #239).datafusion-sparkquote.rsimports its expression types fromdatafusion_expr, so the crate builds without itscorefeature (55.1.0 has the broken import).array_any_valueempty-element guard is dropped: 55.1.0 carries the same guard.if letmatch guard, which broke the workspace MSRV (1.94).db730b42.The metadata-column listing pruning from
spiceai-54(#229) is not carried on this line yet, because Spice has no test that fails without it.Please merge with a merge commit, not squash, so each patch stays its own commit.
Pinned by spiceai/spiceai#14612 (tracking: spiceai/spiceai#13570). Supersedes #238.