Repository navigation
Conversation
…lter (apache#23901) ## Which issue does this PR close? - Closes apache#23900. ## Rationale for this change `push_down_filter` infers equi-key predicates across a join's ON keys and pushes them to the opposite side. For a null-aware join (the `LeftAnti` join produced by `NOT IN` with a nullable subquery), an outer predicate on the left key like `outer.id > 5` is rewritten to `sub.id > 5` and pushed onto the subquery input. Since the inferred predicate must be null-rejecting to be pushed, this drops the subquery's NULL rows and breaks the three-valued `NOT IN` semantics — a NULL in the subquery key must reach the join so the result is empty. Same class of bug as apache#23848, in a different rule. ## What changes are included in this PR? - Skip predicate inference in `infer_join_predicates` when `join.null_aware` is set (mirrors the apache#23848 guard on `FilterNullJoinKeys`). - A `push_down_filter` unit test asserting no predicate is inferred onto the subquery side of a null-aware `LeftAnti` join. - SLT coverage for the failing query, plus a `prefer_hash_join = false` / multi-partition variant. ## Are these changes tested? Yes — new unit test (verified it fails without the guard) and SLT cases. The full optimizer lib suite passes. ## Are there any user-facing changes? No, aside from the correctness fix.
## Which issue does this PR close? - Part of apache#19241. - Stacked on [apache#23311](apache#23311). - Next in stack: apache#23015. - Extracted from apache#19390. ## Rationale for this change For very small `IN` lists, building or probing a hash table can be more work than just comparing the input value with each constant. For example, for `x IN (10, 20, 30)`, the fast path can behave like: ```text x == 10 OR x == 20 OR x == 30 ``` Because the list is tiny, those comparisons are cheap. The implementation stores the constants in a fixed-size array and checks them with a compact comparison chain. “Branchless” here means the comparisons are combined without stopping at the first match. That can be faster for these small fixed-width lists because the CPU gets a predictable sequence of simple operations instead of hash-table setup and probe logic. For primitive values that are not already plain unsigned integers, this PR keeps the logical Arrow type explicit and uses a matching same-width comparison representation only inside the branchless filter. For example, `Float16` uses `UInt16` storage, `Float32` uses `UInt32` storage, and `TimestampNanosecond` uses `UInt64` storage. `Decimal128` and `IntervalMonthDayNano` use their own 16-byte native representation. This preserves bit-pattern equality while relying on Arrow's native primitive compatibility rules: timestamp timezone metadata and Decimal128 precision/scale metadata may differ, while incompatible primitive representations remain rejected. ## What changes are included in this PR? - Adds a const-generic `BranchlessFilter` for small primitive `IN` lists. - Adds thresholds for when this path is used: - up to 16 values for 1-byte types - up to 8 values for 2-byte types - up to 32 values for 4-byte types - up to 16 values for 8-byte types - up to 4 values for 16-byte types - Keeps dispatch concrete and explicit in `strategy.rs`. - Maps each optimized logical type to the comparison representation used by the branchless filter: - `Int8` -> `UInt8` - `Int16`, `Float16` -> `UInt16` - `Int32`, `Float32`, `Date32`, `Time32` -> `UInt32` - `Int64`, `Float64`, `Date64`, `Time64`, `Timestamp`, `Duration` -> `UInt64` - `Decimal128`, `IntervalMonthDayNano` -> their native 16-byte representation - Leaves larger 1-byte and 2-byte lists on the existing bitmap filters. - Leaves larger 4-byte and 8-byte lists on the existing hash/generic paths. - Leaves wider primitive types such as `Decimal256` and unsupported complex types on the generic path. - Keeps the same `IN` / `NOT IN` null behavior as the rest of the stack. - Adds focused coverage for branchless null handling, signed boundary values, slices, Float16/Float32/Float64 bit patterns, compatible timestamp/Decimal128 metadata, incompatible timestamp units, IntervalMonthDayNano values, and same-width wrong-type probe rejection. ## Are these changes tested? Yes. - `cargo fmt --all -- --check` - `cargo test -p datafusion-physical-expr expressions::in_list --lib` - `cargo test -p datafusion-physical-expr --bench in_list_strategy --no-run` - `cargo clippy --all-targets --all-features -- -D warnings` ## Are there any user-facing changes? No. This is an internal performance optimization only. ## Local benchmark snapshot Built and run with `release-nonlto`, filtered to the relevant small primitive-list rows: ```bash cargo bench -p datafusion-physical-expr --profile release-nonlto --bench in_list_strategy -- <filter> --save-baseline <baseline> ``` Filters used: `narrow_integer`, `primitive/i32/small_list`, `primitive/i64/small_list`, `f32/small_list`, `timestamp_ns/small_list`, and `interval_month_day_nano/small_list`. Method: directly compared Criterion's raw sample minima (`min(time / iterations)`) from `sample.json`. Lower is better; changes within +/-5% are treated as noise. Compared baselines: [apache#23311](apache#23311) -> [apache#23014](apache#23014) Relevant scope: small primitive-list rows. Summary: 39 relevant rows, 28 faster, 0 slower, 11 within +/-5%. Largest relevant deltas: | Benchmark | Before | After | Change | |---|---:|---:|---:| | `timestamp_ns/small_list/list=4/match=50%` | 46.55 us | 3.17 us | -93.2% (14.69x faster) | | `f32/small_list/list=4/match=50%` | 33.93 us | 3.04 us | -91.0% (11.15x faster) | | `primitive/i32/small_list/list=4/match=50%` | 32.63 us | 3.08 us | -90.5% (10.58x faster) | | `primitive/i64/small_list/list=4/match=50%` | 33.55 us | 3.18 us | -90.5% (10.54x faster) | | `timestamp_ns/small_list/list=4/match=0%` | 19.57 us | 3.18 us | -83.8% (6.16x faster) | | `f32/small_list/list=4/match=0%` | 18.14 us | 3.05 us | -83.2% (5.95x faster) | | `primitive/i32/small_list/list=4/match=0%` | 17.00 us | 3.04 us | -82.1% (5.59x faster) | | `primitive/i64/small_list/list=4/match=0%` | 17.12 us | 3.22 us | -81.2% (5.31x faster) | | `primitive/i32/small_list/list=16/match=50%/NOT_IN` | 31.98 us | 7.26 us | -77.3% (4.41x faster) | | `nulls/primitive/i32/small_list/list=16/match=50%/nulls=20%` | 29.35 us | 7.32 us | -75.1% (4.01x faster) | | `timestamp_ns/small_list/list=16/match=50%` | 45.32 us | 11.79 us | -74.0% (3.84x faster) | | `nulls/primitive/i32/small_list/list=16/match=50%/nulls=50%` | 25.89 us | 7.31 us | -71.8% (3.54x faster) | | `nulls/primitive/i32/small_list/list=16/match=50%/nulls=20%/NOT_IN` | 26.05 us | 7.42 us | -71.5% (3.51x faster) | | `interval_month_day_nano/small_list/list=4/match=50%` | 52.94 us | 15.52 us | -70.7% (3.41x faster) | | `f32/small_list/list=32/match=50%` | 38.78 us | 13.27 us | -65.8% (2.92x faster) | | `primitive/i64/small_list/list=16/match=50%` | 29.46 us | 11.76 us | -60.1% (2.50x faster) | <details> <summary>Full relevant table (39 rows)</summary> | Benchmark | Before | After | Change | |---|---:|---:|---:| | `narrow_integer/u8/list=4/match=0%` | 3.86 us | 2.79 us | -27.8% (1.38x faster) | | `narrow_integer/u8/list=4/match=50%` | 3.84 us | 2.78 us | -27.7% (1.38x faster) | | `narrow_integer/u8/list=16/match=0%` | 3.88 us | 3.85 us | -0.8% (within noise) | | `narrow_integer/u8/list=16/match=50%` | 3.84 us | 3.86 us | +0.5% (within noise) | | `narrow_integer/i16/list=4/match=0%` | 3.93 us | 3.18 us | -19.1% (1.24x faster) | | `narrow_integer/i16/list=4/match=50%` | 3.92 us | 3.16 us | -19.5% (1.24x faster) | | `narrow_integer/i16/list=64/match=0%` | 3.96 us | 3.82 us | -3.5% (within noise) | | `narrow_integer/i16/list=64/match=50%` | 3.91 us | 3.80 us | -2.9% (within noise) | | `narrow_integer/i16/list=256/match=0%` | 3.90 us | 3.81 us | -2.5% (within noise) | | `narrow_integer/i16/list=256/match=50%` | 3.97 us | 3.81 us | -4.1% (within noise) | | `narrow_integer/f16/list=4/match=0%` | 3.87 us | 3.16 us | -18.5% (1.23x faster) | | `narrow_integer/f16/list=4/match=50%` | 3.94 us | 3.15 us | -20.2% (1.25x faster) | | `narrow_integer/f16/list=64/match=0%` | 3.87 us | 3.84 us | -0.6% (within noise) | | `narrow_integer/f16/list=64/match=50%` | 3.93 us | 3.85 us | -1.9% (within noise) | | `narrow_integer/f16/list=256/match=0%` | 3.90 us | 3.84 us | -1.5% (within noise) | | `narrow_integer/f16/list=256/match=50%` | 3.87 us | 3.91 us | +1.2% (within noise) | | `nulls/narrow_integer/u8/list=16/match=50%/nulls=20%` | 3.92 us | 4.02 us | +2.5% (within noise) | | `primitive/i32/small_list/list=4/match=0%` | 17.00 us | 3.04 us | -82.1% (5.59x faster) | | `primitive/i32/small_list/list=4/match=50%` | 32.63 us | 3.08 us | -90.5% (10.58x faster) | | `primitive/i32/small_list/list=32/match=0%` | 16.34 us | 13.33 us | -18.5% (1.23x faster) | | `primitive/i32/small_list/list=32/match=50%` | 31.17 us | 13.31 us | -57.3% (2.34x faster) | | `primitive/i32/small_list/list=16/match=50%/NOT_IN` | 31.98 us | 7.26 us | -77.3% (4.41x faster) | | `nulls/primitive/i32/small_list/list=16/match=50%/nulls=20%` | 29.35 us | 7.32 us | -75.1% (4.01x faster) | | `nulls/primitive/i32/small_list/list=16/match=50%/nulls=20%/NOT_IN` | 26.05 us | 7.42 us | -71.5% (3.51x faster) | | `nulls/primitive/i32/small_list/list=16/match=50%/nulls=50%` | 25.89 us | 7.31 us | -71.8% (3.54x faster) | | `primitive/i64/small_list/list=4/match=0%` | 17.12 us | 3.22 us | -81.2% (5.31x faster) | | `primitive/i64/small_list/list=4/match=50%` | 33.55 us | 3.18 us | -90.5% (10.54x faster) | | `primitive/i64/small_list/list=16/match=0%` | 16.34 us | 11.93 us | -27.0% (1.37x faster) | | `primitive/i64/small_list/list=16/match=50%` | 29.46 us | 11.76 us | -60.1% (2.50x faster) | | `f32/small_list/list=4/match=0%` | 18.14 us | 3.05 us | -83.2% (5.95x faster) | | `f32/small_list/list=4/match=50%` | 33.93 us | 3.04 us | -91.0% (11.15x faster) | | `f32/small_list/list=32/match=0%` | 22.05 us | 13.43 us | -39.1% (1.64x faster) | | `f32/small_list/list=32/match=50%` | 38.78 us | 13.27 us | -65.8% (2.92x faster) | | `timestamp_ns/small_list/list=4/match=0%` | 19.57 us | 3.18 us | -83.8% (6.16x faster) | | `timestamp_ns/small_list/list=4/match=50%` | 46.55 us | 3.17 us | -93.2% (14.69x faster) | | `timestamp_ns/small_list/list=16/match=0%` | 19.73 us | 12.07 us | -38.8% (1.63x faster) | | `timestamp_ns/small_list/list=16/match=50%` | 45.32 us | 11.79 us | -74.0% (3.84x faster) | | `interval_month_day_nano/small_list/list=4/match=0%` | 20.12 us | 13.20 us | -34.4% (1.52x faster) | | `interval_month_day_nano/small_list/list=4/match=50%` | 52.94 us | 15.52 us | -70.7% (3.41x faster) | </details> --------- Co-authored-by: Andrew Lamb <andrew@nerdnetworks.org>
## Which issue does this PR close? - Closes apache#23770 - Part of apache#15914 ## Rationale for this change Spark provides [`hypot(expr1, expr2)`](https://spark.apache.org/docs/latest/api/sql/#hypot), which returns `sqrt(expr1^2 + expr2^2)` computed without intermediate overflow or underflow. It was not yet implemented in `datafusion-spark` — only an auto-generated test stub existed at `spark/math/hypot.slt` with its query commented out. ## What changes are included in this PR? - Add `SparkHypot` (implementing `ScalarUDFImpl`) in `datafusion/spark/src/function/math/hypot.rs`, backed by Rust's `f64::hypot` — the same overflow-safe algorithm as Java/Spark's `Math.hypot`. - Register it in `datafusion/spark/src/function/math/mod.rs`. - Enable the `hypot.slt` sqllogictest. The signature is `exact(Float64, Float64) -> Float64`, following the `datafusion-spark` convention of only accepting types Spark supports. Computation uses the Arrow `binary` kernel so NULL in either argument propagates to a NULL result, matching Spark. ## Are these changes tested? Yes — `datafusion/sqllogictest/test_files/spark/math/hypot.slt` covers: - scalar Pythagorean triples (`hypot(3, 4)` → 5, `hypot(5, 12)` → 13), - double inputs, - NULL propagation when either argument is NULL, - the array path (including a NULL row), - overflow-safety: `hypot(3e200, 4e200)` stays finite, whereas a naive `sqrt(a^2 + b^2)` would overflow to `Infinity`. ## Are there any user-facing changes? Yes — adds the Spark-compatible `hypot` scalar function to `datafusion-spark`. No breaking changes to public APIs.
## Which issue does this PR close? <!-- We generally require a GitHub issue to be filed for all bug fixes and enhancements and this helps us generate change logs for our releases. You can link an issue to this PR using the GitHub syntax. For example `Closes #123` indicates that this PR will close issue #123. --> - Part of apache#23393 . ## Rationale for this change `SlidingMinAccumulator::size` and `SlidingMaxAccumulator::size` only reported the stack size of their `ScalarValue` field plus its heap payload, ignoring the memory held by the underlying `MovingMin` / `MovingMax` sliding-window buffers. For windowed `MIN`/`MAX` over string or list data, the two per-element stacks can hold megabytes of `ScalarValue` payload that the memory pool was never told about, understating accumulator memory usage. ## What changes are included in this PR? - Add a private `heap_size(elem_heap)` method to `MovingMin<T>` and `MovingMax<T>` that reports the two stack buffers' capacity in bytes plus each stored element's heap payload. - Factor the shared implementation into a `moving_stacks_heap_size` free helper so the two types stay in sync. - Include the buffer bytes in `SlidingMinAccumulator::size` and `SlidingMaxAccumulator::size` via the new method. ## Are these changes tested? Yes. Two new unit tests in `datafusion/functions-aggregate/src/min_max.rs`: - `moving_min_max_heap_size_i32` — fixed-width `T`, verifies buffer-only accounting with and without pushed elements. - `moving_min_max_heap_size_counts_elems` — `T = String`, verifies each of the two slots in a `(T, T)` pair contributes independently to the heap payload (mirroring the two independent `Clone`s made by `push`).
## Which issue does this PR close? - Related to apache#21882 (comment). ## Rationale for this change PR apache#21882 adds custom spill file support. Having an example showing how a downstream application can use this new API will help make sure the API is good enough for our needs ## What changes are included in this PR? This PR adds a small example showing how users can back spill files with an `ObjectStore`, using a local object store for a runnable example while keeping the implementation applicable to remote stores. ## Are these changes tested? Yes by CI ## Are there any user-facing changes? Yes. This adds a new example for configuring ObjectStore-backed spill files.
## Which issue does this PR close? <!-- We generally require a GitHub issue to be filed for all bug fixes and enhancements and this helps us generate change logs for our releases. You can link an issue to this PR using the GitHub syntax. For example `Closes #123` indicates that this PR will close issue #123. --> - Closes #N/A. ## Rationale for this change <!-- Why are you proposing this change? If this is already explained clearly in the issue then this section is not needed. Explaining clearly why changes are proposed helps reviewers understand your changes and offer better suggestions for fixes. --> Since DataFusion doesn't typically guartantee wire format compatibility, I remove the backward compatiblity shim that introduced in PR apache#23189 FYI: apache#23189 (comment) ## What changes are included in this PR? <!-- There is no need to duplicate the description in the issue here but it is sometimes worth providing a summary of the individual changes in this PR. --> DataFusion does not guarantee serialized plans across versions. Keeping `partitioned_by_file_group` therefore leaves a dead schema field and decoder path after `output_partitioning` became the source of truth. Reserve the old field number and name to prevent future reuse. ## Are these changes tested? <!-- We typically require tests for all PRs in order to: 1. Prevent the code from being accidentally broken by subsequent changes 2. Serve as another way to document the expected behavior of the code If tests are not included in your PR, please explain why (for example, are they covered by existing tests)? --> Yes ## Are there any user-facing changes? <!-- If there are user-facing changes then we may require documentation to be updated before approving the PR. --> <!-- If there are any breaking changes to public APIs, please add the `api change` label. --> Yes Signed-off-by: Jiawei Zhao <Phoenix500526@163.com> Co-authored-by: Andrew Lamb <andrew@nerdnetworks.org>
…llable build key (apache#23173) ## Which issue does this close? - Closes apache#23126. ## Rationale for this change `x NOT IN (subquery)` plans to a null-aware `LeftAnti` hash join (build = outer `x`, probe = subquery). Join dynamic filter pushdown pushes a bounds + membership filter, built from the build keys, onto the probe scan. That filter can prune every probe row. A null-aware `LeftAnti` reads an empty probe as a genuinely-empty subquery, so it emits build-side NULL rows that should drop: `NULL NOT IN (non-empty)` is UNKNOWN, not TRUE. The result is scan-dependent, so it's a silent correctness bug. A `VALUES` scan ignores the pushed filter and stays correct; a parquet scan applies it and is wrong. apache#23103 (the probe-side NULL drop) is orthogonal; this is the build-side NULL. ## What changes are included in this PR? Skip join dynamic filter pushdown for a null-aware anti join when the build key can be NULL. The build-side NULL emission depends on whether the probe is truly empty, which the pushed filter can change by emptying it. A NOT NULL build key has no such NULL, so it keeps the pushdown. The check is static: a schema-nullable build key disables the pushdown even when the data contains no NULLs. A runtime alternative (keep the pushdown and neutralize the filter only when the build actually holds a NULL key) would restore the optimization for those cases. I'd leave that as a follow-up. ## Are these changes tested? Yes. A `push_down_filter_parquet.slt` case reproduces it (build-side NULL, a non-matching parquet probe) and asserts the single correct row. Without the change it returns the extra NULL. In addition, unit tests pin both directions of the guard: a nullable build key rejects the pushdown and a NOT NULL build key keeps it. ## Are there any user-facing changes? `NOT IN` over a parquet (or otherwise prunable) scan with a nullable outer key now returns correct results. Such joins lose the dynamic filter pushdown.
…cimal) (apache#23631) ## Which issue does this PR close? <!-- We generally require a GitHub issue to be filed for all bug fixes and enhancements and this helps us generate change logs for our releases. You can link an issue to this PR using the GitHub syntax. For example `Closes #123` indicates that this PR will close issue #123. --> N/A ## Rationale for this change <!-- Why are you proposing this change? If this is already explained clearly in the issue then this section is not needed. Explaining clearly why changes are proposed helps reviewers understand your changes and offer better suggestions for fixes. --> I noticed there were some subtle errors with how decimals were handled in scalar values, and also opportunity to remove power calls in favour of precomputed constant tables. Also filling out some other missing support. ## What changes are included in this PR? <!-- There is no need to duplicate the description in the issue here but it is sometimes worth providing a summary of the individual changes in this PR. --> I recommend looking at the commits as they are self contained with detailed messages for each. ## Are these changes tested? <!-- We typically require tests for all PRs in order to: 1. Prevent the code from being accidentally broken by subsequent changes 2. Serve as another way to document the expected behavior of the code If tests are not included in your PR, please explain why (for example, are they covered by existing tests)? --> Yes ## Are there any user-facing changes? <!-- If there are user-facing changes then we may require documentation to be updated before approving the PR. --> No <!-- If there are any breaking changes to public APIs, please add the `api change` label. -->
…lter pushdown enabled (apache#23638) ## Which issue does this PR close? - Closes apache#23531. ## Rationale for this change `input_file_name()` is `FileSource` dependent like `file_row_index()`, and therefore shouldn't be pushed down into a filter. ## What changes are included in this PR? `PushdownChecker` now handles both UDFs consistently. If we keep adding this sort of metadata functions, we might want a better API to detect them, but for now I think this is a reasonable approach that isn't very invasive. ## Are these changes tested? Additional SLT test that verifies that both function behave correctly with pushdown either enabled or disabled. ## Are there any user-facing changes? None --------- Signed-off-by: Adam Gutglick <adamgsal@gmail.com>
…xpr::check_bigger_cast (apache#23808) (apache#23809) ## Which issue does this PR close? - Closes apache#23808. ## Rationale for this change ref. apache#23807 `CastExpr::check_bigger_cast` is used to determine whether a cast is a widening (order-preserving) conversion. Currently, it classifies `Int32 -> Float32`, `UInt32 -> Float32`, `Int64 -> Float64`, and `UInt64 -> Float64` as widening casts. However, integer-to-float conversions for 32-bit and 64-bit integers lose precision when values exceed the mantissa bit limit (24 bits for `Float32`, 53 bits for `Float64`). For example: `16_777_216_i32 as f32 == 16_777_217_i32 as f32` (both yield 16777216.0f32). Because distinct integer inputs can collapse to the same float output, these casts are not strictly 1-to-1 (injective) and can break suffix sort key ordering in multi-column ordering analysis (e.g. `[CAST(a AS Float32), b]`). ## What changes are included in this PR? - Updated `CastExpr::check_bigger_cast` to exclude precision-losing integer-to-float conversions (`Int32/UInt32 -> Float32` and `Int64/UInt64 -> Float64`). - Added unit test `test_check_bigger_cast_precision_loss` to verify precision-losing casts return `false` while exact conversions (`Int16 -> Float32`, `Int32 -> Float64`, etc.) continue to return `true`. ## Are these changes tested? Yes, new unit test `test_check_bigger_cast_precision_loss` in `cast.rs`. ## Are there any user-facing changes? No breaking API changes. Internal optimizer behavior fix.
…ossible (apache#23702) ## Which issue does this PR close? N/A ## Rationale for this change `SortPreservingMergeStream` is a little complex, so add some guiding comments and make it as textbook-like as possible ## What changes are included in this PR? Added comments, reorder code While this was done this also fixed couple of bugs due to how it work: 1. leftover drain was not counted in the `elapsed_compute` 2. `limit(0)` returns 0 rows and not 1 ## Are these changes tested? Existing tests ## Are there any user-facing changes? Not API ones. `limit(0)` now returns 0 rows
## Which issue does this PR close? - Closes apache#21428. ## Rationale for this change This PR adds `BuildHasher`-based variants for `hash_utils` so callers can compute row hashes with a caller-provided hash builder instead of always using DataFusion's default `RandomState`. The main constraint is performance: `with_hashes` is a hot path, especially for string, dictionary, and nested array hashing. A previous version in apache#21429 caused measurable regressions in the default `RandomState` path, for example `large_utf8: single, no nulls` regressed from roughly `26.7us` to `36.3us`, and `large_utf8: multiple, no nulls` from roughly `112us` to `127us`. This version keeps the default path performance-oriented by avoiding a fully generic `BuildHasher` rewrite of the existing hot loops. ## What changes are included in this PR? This PR adds: - `with_hashes_with_hasher` - `create_hashes_with_hasher` - custom-hasher implementations for primitive, string, binary, byte-view, dictionary, and nested arrays - tests covering custom hashers, multi-column hashing, and dictionary equivalence The implementation intentionally uses a hybrid design: - Default `RandomState` leaf hot paths remain specialized. - Custom `BuildHasher` leaf paths live separately in `hash_utils/build_hasher.rs`. - Nested/structural logic is shared through an internal child-hashing adapter, so struct/list/map/union/run/dictionary behavior does not need to be broadly duplicated. The trade-off is that there is still some duplication for primitive/string/binary leaf loops. That duplication is intentional: those are the hottest loops, and keeping them separate prevents the existing `RandomState` path from becoming generic over `BuildHasher` or being perturbed by the custom-hasher implementation. ## Are these changes tested? Yes. --------- Co-authored-by: Dmitrii Blaginin <dmitrii@blaginin.me> Co-authored-by: Andrew Lamb <andrew@nerdnetworks.org>
…t whether they keep the same ordering of the input (apache#23807) ## Which issue does this PR close? - Closes apache#23798 Related to: - apache#16217 ## Rationale for this change To be able to keep the same sorting order allowing for more optimizations Now comet own `cast` implementation can be recognized as not modifying sort order in the same cases that datafusion cast does. ## What changes are included in this PR? added `strictly_order_preserving` property to `ExprProperties` + varius other places and replaced the hard coded logic for cast (`substitute_cast_ordering`) about keeping input order with more generic approach that now any expression can implement and have the same advantage of sort elimination also marked from_unixtime as keeping ordering to show a case of this optimization ## Are these changes tested? yes ## Are there any user-facing changes? yes, breaking change, added `strictly_order_preserving` property to `ExprProperties` and to `FFI_ExprProperties` this property means that given expression `f` and 2 values from the input column `a` and `b` the following variants are kept: 1. `a.cmp(b) == f(a).cmp(f(b))` 2. nulls maps to nulls Example of satisfying expression: `cast(col_a as BIGINT)` where `col_a` is `INT` it is keeping the properties Example of not satisfying: `floor` - floor can not `array_repeat(my_col, 2)` which might look like at first glance as keeping the property as well but in fact it does not. the reason is that `array_repeat(null, 2)` will output list of 2 nulls which breaks the 2nd property that nulls must be kept as nulls how to migrate: Option 1 - keeping the old behavior (safest but least performant) set `strictly_order_preserving` to false Option 2 - Using the new optimization that this opens up: set `strictly_order_preserving` only if the expression keep both variants
## Which issue does this PR close? Part of apache#22330. ## Rationale for this change `ForeignScalarUDF` inherits the default `preserves_lex_ordering`, so producer overrides are lost across the FFI boundary. ## What changes are included in this PR? - Forward `preserves_lex_ordering` through `FFI_ScalarUDF`. - Reuse the existing placement UDF for unit and dynamic-library coverage. ## Are these changes tested? - `cargo test -p datafusion-ffi --features integration-tests` - `cargo clippy --all-targets --all-features -- -D warnings` ## Are there any user-facing changes? The FFI ABI changes. Foreign libraries must rebuild against the new DataFusion version. --------- Signed-off-by: Amogh Ramesh <ramogh2404@gmail.com>
…tonic deques (apache#23826) (apache#23827) ## Which issue does this PR close? - Closes apache#23826 ### What changes are included in this PR? This PR optimizes sliding window `MIN`/`MAX` aggregate functions using a **Sequence-Numbered Monotonic Deque** instead of a Two-Stack Queue. **Revised Design:** - We store `(sequence_number, value)` pairs in a single `VecDeque`, kept in strictly monotonic order. - **`push(val)`**: Evicts dominated elements from the back, then pushes the new value with the current `push_seq` and increments `push_seq`. - **`pop()`**: Increments `pop_seq`. If the front elements sequence number equals the old `pop_seq`, it is expired and popped. - **Benefits**: No secondary FIFO queue (lower memory) and no `clone()` overhead. ### Are these changes tested? Yes, existing tests pass. Added tests for empty-window `pop()` and duplicate-heavy scenarios. ### Are there any user-facing changes? **API Change**: `MovingMin` and `MovingMax` were changed to `pub(crate)` visibility, and their `pop()` methods now return `()` instead of returning a value. Performance is significantly improved (2x-3.5x throughput). --------- Co-authored-by: Pavan <pavan@Lakshmis-MacBook-Air.local>
## Which issue does this PR close? NA ## Rationale for this change Adds [datapress](https://docs.datap-rs.org) to the list of known users in the documentation. ## What changes are included in this PR? NA ## Are these changes tested? NA ## Are there any user-facing changes? NA
## Which issue does this PR close? - Closes apache#11748 ## Rationale for this change `EliminateGroupByConstant` removes GROUP BY expressions that are constants. If all of the GROUP BY expressions are constants, eliminating all of them results in converting a grouped aggregate into an ungrouped (global) aggregate. This changes the semantics of the query: a grouped aggregate query on an empty input returns zero rows, whereas an ungrouped aggregate query produces a single row. ## What changes are included in this PR? `EliminateGroupByConstant` now declines to eliminate constant GROUP BY expressions, if doing so would result in removing all of the grouping expressions. ## Are these changes tested? Yes. Sqllogictest reproducer for the original issue, updated unit test and optimizer SLT expectations. ## Are there any user-facing changes? Queries with all-constant GROUP BY may now return fewer (correct) rows.
…pache#23874) ## Which issue does this PR close? - Closes apache#23872 ## Rationale for this change `SlidingMinAccumulator::update_batch` skipped NULL values. This meant that if all the non-NULL values in a window frame were retracted, the window frame would not be empty but the `MovingMin` data structure by the `SlidingMinAccumulator` would not contain any values. This resulted in incorrectly returning a stale non-NULL value for a sliding `min()` over a window frame consisting of only NULL values. Along the way, optimize the min and max sliding window accumulators to make them both more efficient and more symmetric with one another. In the original coding, `SlidingMaxAccumulator` included NULL values but `SlidingMinAccumulator` omitted them, in part because omitting NULLs made the original `retract_batch` implementation more expensive. This PR optimizes `retract_batch`, so we can now use the same scheme for both the min and max sliding accumulators: * Omit NULLs on `update_batch` (this improves on the prior behavior of `max`) * Efficiently account for NULLs in `retract_batch` (this improves on the prior behavior of `min`) * Ensure correct results for all-NULL window frames (this fixes the prior bug in `min`). ## What changes are included in this PR? * Fix bug in `min()` over all-NULL window frames * Optimize `SlidingMinAccumulator::retract_batch` (avoid materializing values just to count NULLs) * Optimize `SlidingMinAccumulator::update_batch` (omit NULLs), also improving symmetry with `min` * Optimize both accumulators to stop caching the current `min` / `max`; this saves a few clones, but perhaps more importantly it is simpler and avoids the risk of inconsistency between the cached value and the underlying `MovingMin` / `MovingMax` data structure ## Are these changes tested? Yes, new tests added. ## Are there any user-facing changes? No, aside from the bug fix.
## Which issue does this PR close?
- N/A
## Rationale for this change
```
$ cargo test -p datafusion-functions-aggregate
[...]
warning: associated function `from_parts` is never used
--> datafusion/physical-expr/src/expressions/dynamic_filters/mod.rs:464:8
[...]
```
Fix this by appropriately gating the compilation of `from_parts`.
## What changes are included in this PR?
* Squelch unused code warning
## Are these changes tested?
Yes.
## Are there any user-facing changes?
No.
…w retract (apache#23913) ## Which issue does this PR close? - Closes apache#23912. ## Rationale for this change `DistinctPercentileContAccumulator` reused the shared set-based `GenericDistinctBuffer`, which doesn't fit it: (1) the buffer asserts a single input column but `percentile_cont` passes two (value + percentile), so every `percentile_cont(DISTINCT ...)` panicked; (2) the buffer is a plain `HashSet` with no multiplicity, so sliding-window `retract_batch` dropped a value while duplicates were still in the frame. ## What changes are included in this PR? - Replace the shared buffer in this accumulator with a per-accumulator `HashMap<Hashable, usize>` count map: `update_batch` reads only the value column and increments; `retract_batch` decrements and removes a key only at zero; `state`/`merge_batch` keep the same List state shape. Other `GenericDistinctBuffer` users are untouched. - Regression tests in `aggregate.slt` for the plain distinct query and the sliding-window duplicate-retract case. ## Are these changes tested? Yes — new regression tests; the full `aggregate.slt` suite passes. ## Are there any user-facing changes? `percentile_cont(DISTINCT ...)` now works instead of panicking, and returns correct results in sliding windows.
## Which issue does this PR close? - Closes apache#23717. ## Rationale for this change Spark and Java format negative numeric values with parentheses when the `(` flag is present. The decimal formatting path always emitted a minus sign because it ignored `negative_in_parentheses`, while the floating-point path already handled the flag. This made `format_string` inconsistent across numeric input types. ## What changes are included in this PR? - Add a closing suffix when a negative decimal uses parentheses formatting. - Include the suffix when calculating width, left alignment, and zero padding. - Add regression coverage for grouped negative decimals with and without an explicit width. ## Are these changes tested? Yes. The following checks pass locally: - `cargo fmt --all -- --check` - `cargo clippy --all-targets --all-features -- -D warnings` - `cargo test -p datafusion-spark --lib` - The full workspace test command from the contributor guide with `avro,json,backtrace,extended_tests,recursive_protection,parquet_encryption` enabled ## Are there any user-facing changes? Yes. Spark-compatible formatting of negative decimal values now honors the parentheses flag. There are no public API or breaking changes.
## Which issue does this PR close? <!-- We generally require a GitHub issue to be filed for all bug fixes and enhancements and this helps us generate change logs for our releases. You can link an issue to this PR using the GitHub syntax. For example `Closes #123` indicates that this PR will close issue #123. --> - Closes #. ## Rationale for this change <!-- Why are you proposing this change? If this is already explained clearly in the issue then this section is not needed. Explaining clearly why changes are proposed helps reviewers understand your changes and offer better suggestions for fixes. --> Precedes apache#23720. While displaying a plan with metrics, allow filtering them by name, not just by category or type. This is very useful when writing snapshot tests where only a specific metrics needs to be asserted, without other unrelated metrics polluting the snapshot assertion. Regardless of what happens with apache#23720, I think this PR is still worth it, as it allows creating some really nice `insta` tests asserting runtime properties reliably by just cherry picking the runtime metrics relevant for that specific test. This is relevant not only within DataFusion codebase, but also for other people's codebases using DataFusion and relying on `insta` for snapshot testing, See an example of this here: https://github.com/apache/datafusion/pull/23720/changes#diff-8281405e117428077c19c07afbbe59fed303a3460096e958f761d34f0899b618 ## What changes are included in this PR? <!-- There is no need to duplicate the description in the issue here but it is sometimes worth providing a summary of the individual changes in this PR. --> While displaying metrics, allows filtering them by name ## Are these changes tested? <!-- We typically require tests for all PRs in order to: 1. Prevent the code from being accidentally broken by subsequent changes 2. Serve as another way to document the expected behavior of the code If tests are not included in your PR, please explain why (for example, are they covered by existing tests)? --> Yes, by a new small test. ## Are there any user-facing changes? <!-- If there are user-facing changes then we may require documentation to be updated before approving the PR. --> <!-- If there are any breaking changes to public APIs, please add the `api change` label. --> People using the metrics display API will be able to filter metrics by name.
## Which issue does this PR close? N/A ## Rationale for this change I saw that `empty` implementation is inefficient. I wrote review comments for more why inefficient. there are some optimizations that do not require benchmark since they are obvious once you understand, this is one of them ## What changes are included in this PR? rewrote the function to be fast ## Are these changes tested? existing tests ## Are there any user-facing changes? no
Bumps the codeql-actions group with 2 updates: [github/codeql-action/init](https://github.com/github/codeql-action) and [github/codeql-action/analyze](https://github.com/github/codeql-action). Updates `github/codeql-action/init` from 4.37.1 to 4.37.3 <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/github/codeql-action/releases">github/codeql-action/init's releases</a>.</em></p> <blockquote> <h2>v4.37.3</h2> <p>No user facing changes.</p> <h2>v4.37.2</h2> <ul> <li>The new address format for the <code>config-file</code> input that was introduced in CodeQL Action 4.37.0 is now enabled by default. In addition to the format described there, the <code>remote=</code> prefix can now be used to explicitly indicate that the input refers to a remote file. All previous input formats continue to be accepted as well. <a href="https://redirect.github.com/github/codeql-action/pull/4023">#4023</a></li> <li>The CodeQL Action can now make use of <a href="https://docs.github.com/en/code-security/how-tos/secure-at-scale/configure-organization-security/manage-usage-and-access/giving-org-access-private-registries">configured private registries</a> in Default Setup to retrieve CodeQL configuration files from remote repositories that require authentication. This will allow customers to store their CodeQL configuration in a single repository that can then be referenced by Default Setup workflows in other repositories. We expect to roll this and other, related changes out to everyone in July. <a href="https://redirect.github.com/github/codeql-action/pull/4007">#4007</a></li> </ul> </blockquote> </details> <details> <summary>Changelog</summary> <p><em>Sourced from <a href="https://github.com/github/codeql-action/blob/main/CHANGELOG.md">github/codeql-action/init's changelog</a>.</em></p> <blockquote> <h1>CodeQL Action Changelog</h1> <p>See the <a href="https://github.com/github/codeql-action/releases">releases page</a> for the relevant changes to the CodeQL CLI and language packs.</p> <h2>[UNRELEASED]</h2> <ul> <li>This version of the CodeQL Action adds support for the <code>tools</code> input for the <code>codeql-action/init</code> step to be specified using a <code>github-codeql-tools</code> <a href="https://docs.github.com/en/organizations/managing-organization-settings/managing-custom-properties-for-repositories-in-your-organization">repository property</a>. This feature will gradually be rolled out following the release of this version. Once rolled out, this allows for the CodeQL CLI version that is used in GitHub-managed workflows, such as Default Setup, to be set to a custom value. For example, customers who run into issues with rate limits when a new CodeQL CLI version is released can set the value to <code>toolcache</code> to always use the CodeQL CLI version that is available in the runner toolcache. For Advanced Setup workflows, the value provided for <code>tools</code> in the workflow definition always takes precedence unless the value of the repository property starts with <code>!</code>. <a href="https://redirect.github.com/github/codeql-action/pull/4037">#4037</a></li> </ul> <h2>4.37.3 - 22 Jul 2026</h2> <p>No user facing changes.</p> <h2>4.37.2 - 21 Jul 2026</h2> <ul> <li>The new address format for the <code>config-file</code> input that was introduced in CodeQL Action 4.37.0 is now enabled by default. In addition to the format described there, the <code>remote=</code> prefix can now be used to explicitly indicate that the input refers to a remote file. All previous input formats continue to be accepted as well. <a href="https://redirect.github.com/github/codeql-action/pull/4023">#4023</a></li> <li>The CodeQL Action can now make use of <a href="https://docs.github.com/en/code-security/how-tos/secure-at-scale/configure-organization-security/manage-usage-and-access/giving-org-access-private-registries">configured private registries</a> in Default Setup to retrieve CodeQL configuration files from remote repositories that require authentication. This will allow customers to store their CodeQL configuration in a single repository that can then be referenced by Default Setup workflows in other repositories. We expect to roll this and other, related changes out to everyone in July. <a href="https://redirect.github.com/github/codeql-action/pull/4007">#4007</a></li> </ul> <h2>4.37.1 - 16 Jul 2026</h2> <ul> <li><em>Upcoming breaking change</em>: Add a deprecation warning for customers using CodeQL version 2.20.6 and earlier. These versions of CodeQL were discontinued on 1 July 2026 alongside GitHub Enterprise Server 3.16, and will be unsupported by the next minor release of the CodeQL Action. <a href="https://redirect.github.com/github/codeql-action/pull/3956">#3956</a></li> <li>Update default CodeQL bundle version to <a href="https://github.com/github/codeql-action/releases/tag/codeql-bundle-v2.26.1">2.26.1</a>. <a href="https://redirect.github.com/github/codeql-action/pull/4019">#4019</a></li> </ul> <h2>4.37.0 - 08 Jul 2026</h2> <ul> <li>Update default CodeQL bundle version to <a href="https://github.com/github/codeql-action/releases/tag/codeql-bundle-v2.26.0">2.26.0</a>. <a href="https://redirect.github.com/github/codeql-action/pull/3995">#3995</a></li> <li>In addition to the existing input format, the <code>config-file</code> input for the <code>codeql-action/init</code> step will soon support a new <code>[owner/]repo[@ref][:path]</code> format. All components except the repository name are optional. If omitted, <code>owner</code> defaults to the same owner as the repository the analysis is running for, <code>ref</code> to <code>main</code>, and <code>path</code> to <code>.github/codeql-action.yaml</code>. Support for this format ships in this version of the CodeQL Action, but will only be enabled over the coming weeks. <a href="https://redirect.github.com/github/codeql-action/pull/3973">#3973</a></li> </ul> <h2>4.36.3 - 01 Jul 2026</h2> <p>No user facing changes.</p> <h2>4.36.2 - 04 Jun 2026</h2> <ul> <li>Cache CodeQL CLI version information across Actions steps. <a href="https://redirect.github.com/github/codeql-action/pull/3943">#3943</a></li> <li>Reduce requests while waiting for analysis processing by using exponential backoff when polling SARIF processing status. <a href="https://redirect.github.com/github/codeql-action/pull/3937">#3937</a></li> <li>Update default CodeQL bundle version to <a href="https://github.com/github/codeql-action/releases/tag/codeql-bundle-v2.25.6">2.25.6</a>. <a href="https://redirect.github.com/github/codeql-action/pull/3948">#3948</a></li> </ul> <h2>4.36.1 - 02 Jun 2026</h2> <p>No user facing changes.</p> <h2>4.36.0 - 22 May 2026</h2> <ul> <li><em>Breaking change</em>: Bump the minimum required CodeQL bundle version to 2.19.4. <a href="https://redirect.github.com/github/codeql-action/pull/3894">#3894</a></li> <li>Add support for SHA-256 Git object IDs. <a href="https://redirect.github.com/github/codeql-action/pull/3893">#3893</a></li> <li>Update default CodeQL bundle version to <a href="https://github.com/github/codeql-action/releases/tag/codeql-bundle-v2.25.5">2.25.5</a>. <a href="https://redirect.github.com/github/codeql-action/pull/3926">#3926</a></li> </ul> <h2>4.35.5 - 15 May 2026</h2> <ul> <li>We have improved how the JavaScript bundles for the CodeQL Action are generated to avoid duplication across bundles and reduce the size of the repository by around 70%. This should have no effect on the runtime behaviour of the CodeQL Action. <a href="https://redirect.github.com/github/codeql-action/pull/3899">#3899</a></li> </ul> <!-- raw HTML omitted --> </blockquote> <p>... (truncated)</p> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/github/codeql-action/commit/e4fba868fa4b1b91e1fdab776edc8cfbe6e9fb81"><code>e4fba86</code></a> Merge pull request <a href="https://redirect.github.com/github/codeql-action/issues/4031">#4031</a> from github/update-v4.37.3-72f6a9da0</li> <li><a href="https://github.com/github/codeql-action/commit/fb50ab5d62a274adf3ef3e22cfe750ae87a0ede7"><code>fb50ab5</code></a> Update changelog for v4.37.3</li> <li><a href="https://github.com/github/codeql-action/commit/72f6a9da0def52d9193d6a758f0378b65091f8d1"><code>72f6a9d</code></a> Merge pull request <a href="https://redirect.github.com/github/codeql-action/issues/4030">#4030</a> from github/mbg/fix/no-proxy</li> <li><a href="https://github.com/github/codeql-action/commit/3b5ee58597653d9cc6785f3f1277f796d81f3646"><code>3b5ee58</code></a> Use default <code>request</code> options instead of <code>undefined</code></li> <li><a href="https://github.com/github/codeql-action/commit/bfb6be4b5ecd3650f02f530571453e8c64ef0778"><code>bfb6be4</code></a> Merge pull request <a href="https://redirect.github.com/github/codeql-action/issues/4028">#4028</a> from github/mergeback/v4.37.2-to-main-e0647621</li> <li><a href="https://github.com/github/codeql-action/commit/526ab84f9858816d9cf5f7b9df4dd5e2235f0eba"><code>526ab84</code></a> Rebuild</li> <li><a href="https://github.com/github/codeql-action/commit/d6217b9b8c14166e4851db94c11155d03bd13c07"><code>d6217b9</code></a> Update changelog and version after v4.37.2</li> <li><a href="https://github.com/github/codeql-action/commit/e0647621c2984b5ed2f768cb892365bf2a616ad1"><code>e064762</code></a> Merge pull request <a href="https://redirect.github.com/github/codeql-action/issues/4027">#4027</a> from github/update-v4.37.2-385bcdc5a</li> <li><a href="https://github.com/github/codeql-action/commit/e0faed839190caa67a5cd42f1cc16246028ca3df"><code>e0faed8</code></a> Add a couple of change notes</li> <li><a href="https://github.com/github/codeql-action/commit/73aad0eaa9df172668665a150d17b8bc5a650c20"><code>73aad0e</code></a> Update changelog for v4.37.2</li> <li>Additional commits viewable in <a href="https://github.com/github/codeql-action/compare/7188fc363630916deb702c7fdcf4e481b751f97a...e4fba868fa4b1b91e1fdab776edc8cfbe6e9fb81">compare view</a></li> </ul> </details> <br /> Updates `github/codeql-action/analyze` from 4.37.1 to 4.37.3 <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/github/codeql-action/releases">github/codeql-action/analyze's releases</a>.</em></p> <blockquote> <h2>v4.37.3</h2> <p>No user facing changes.</p> <h2>v4.37.2</h2> <ul> <li>The new address format for the <code>config-file</code> input that was introduced in CodeQL Action 4.37.0 is now enabled by default. In addition to the format described there, the <code>remote=</code> prefix can now be used to explicitly indicate that the input refers to a remote file. All previous input formats continue to be accepted as well. <a href="https://redirect.github.com/github/codeql-action/pull/4023">#4023</a></li> <li>The CodeQL Action can now make use of <a href="https://docs.github.com/en/code-security/how-tos/secure-at-scale/configure-organization-security/manage-usage-and-access/giving-org-access-private-registries">configured private registries</a> in Default Setup to retrieve CodeQL configuration files from remote repositories that require authentication. This will allow customers to store their CodeQL configuration in a single repository that can then be referenced by Default Setup workflows in other repositories. We expect to roll this and other, related changes out to everyone in July. <a href="https://redirect.github.com/github/codeql-action/pull/4007">#4007</a></li> </ul> </blockquote> </details> <details> <summary>Changelog</summary> <p><em>Sourced from <a href="https://github.com/github/codeql-action/blob/main/CHANGELOG.md">github/codeql-action/analyze's changelog</a>.</em></p> <blockquote> <h1>CodeQL Action Changelog</h1> <p>See the <a href="https://github.com/github/codeql-action/releases">releases page</a> for the relevant changes to the CodeQL CLI and language packs.</p> <h2>[UNRELEASED]</h2> <ul> <li>This version of the CodeQL Action adds support for the <code>tools</code> input for the <code>codeql-action/init</code> step to be specified using a <code>github-codeql-tools</code> <a href="https://docs.github.com/en/organizations/managing-organization-settings/managing-custom-properties-for-repositories-in-your-organization">repository property</a>. This feature will gradually be rolled out following the release of this version. Once rolled out, this allows for the CodeQL CLI version that is used in GitHub-managed workflows, such as Default Setup, to be set to a custom value. For example, customers who run into issues with rate limits when a new CodeQL CLI version is released can set the value to <code>toolcache</code> to always use the CodeQL CLI version that is available in the runner toolcache. For Advanced Setup workflows, the value provided for <code>tools</code> in the workflow definition always takes precedence unless the value of the repository property starts with <code>!</code>. <a href="https://redirect.github.com/github/codeql-action/pull/4037">#4037</a></li> </ul> <h2>4.37.3 - 22 Jul 2026</h2> <p>No user facing changes.</p> <h2>4.37.2 - 21 Jul 2026</h2> <ul> <li>The new address format for the <code>config-file</code> input that was introduced in CodeQL Action 4.37.0 is now enabled by default. In addition to the format described there, the <code>remote=</code> prefix can now be used to explicitly indicate that the input refers to a remote file. All previous input formats continue to be accepted as well. <a href="https://redirect.github.com/github/codeql-action/pull/4023">#4023</a></li> <li>The CodeQL Action can now make use of <a href="https://docs.github.com/en/code-security/how-tos/secure-at-scale/configure-organization-security/manage-usage-and-access/giving-org-access-private-registries">configured private registries</a> in Default Setup to retrieve CodeQL configuration files from remote repositories that require authentication. This will allow customers to store their CodeQL configuration in a single repository that can then be referenced by Default Setup workflows in other repositories. We expect to roll this and other, related changes out to everyone in July. <a href="https://redirect.github.com/github/codeql-action/pull/4007">#4007</a></li> </ul> <h2>4.37.1 - 16 Jul 2026</h2> <ul> <li><em>Upcoming breaking change</em>: Add a deprecation warning for customers using CodeQL version 2.20.6 and earlier. These versions of CodeQL were discontinued on 1 July 2026 alongside GitHub Enterprise Server 3.16, and will be unsupported by the next minor release of the CodeQL Action. <a href="https://redirect.github.com/github/codeql-action/pull/3956">#3956</a></li> <li>Update default CodeQL bundle version to <a href="https://github.com/github/codeql-action/releases/tag/codeql-bundle-v2.26.1">2.26.1</a>. <a href="https://redirect.github.com/github/codeql-action/pull/4019">#4019</a></li> </ul> <h2>4.37.0 - 08 Jul 2026</h2> <ul> <li>Update default CodeQL bundle version to <a href="https://github.com/github/codeql-action/releases/tag/codeql-bundle-v2.26.0">2.26.0</a>. <a href="https://redirect.github.com/github/codeql-action/pull/3995">#3995</a></li> <li>In addition to the existing input format, the <code>config-file</code> input for the <code>codeql-action/init</code> step will soon support a new <code>[owner/]repo[@ref][:path]</code> format. All components except the repository name are optional. If omitted, <code>owner</code> defaults to the same owner as the repository the analysis is running for, <code>ref</code> to <code>main</code>, and <code>path</code> to <code>.github/codeql-action.yaml</code>. Support for this format ships in this version of the CodeQL Action, but will only be enabled over the coming weeks. <a href="https://redirect.github.com/github/codeql-action/pull/3973">#3973</a></li> </ul> <h2>4.36.3 - 01 Jul 2026</h2> <p>No user facing changes.</p> <h2>4.36.2 - 04 Jun 2026</h2> <ul> <li>Cache CodeQL CLI version information across Actions steps. <a href="https://redirect.github.com/github/codeql-action/pull/3943">#3943</a></li> <li>Reduce requests while waiting for analysis processing by using exponential backoff when polling SARIF processing status. <a href="https://redirect.github.com/github/codeql-action/pull/3937">#3937</a></li> <li>Update default CodeQL bundle version to <a href="https://github.com/github/codeql-action/releases/tag/codeql-bundle-v2.25.6">2.25.6</a>. <a href="https://redirect.github.com/github/codeql-action/pull/3948">#3948</a></li> </ul> <h2>4.36.1 - 02 Jun 2026</h2> <p>No user facing changes.</p> <h2>4.36.0 - 22 May 2026</h2> <ul> <li><em>Breaking change</em>: Bump the minimum required CodeQL bundle version to 2.19.4. <a href="https://redirect.github.com/github/codeql-action/pull/3894">#3894</a></li> <li>Add support for SHA-256 Git object IDs. <a href="https://redirect.github.com/github/codeql-action/pull/3893">#3893</a></li> <li>Update default CodeQL bundle version to <a href="https://github.com/github/codeql-action/releases/tag/codeql-bundle-v2.25.5">2.25.5</a>. <a href="https://redirect.github.com/github/codeql-action/pull/3926">#3926</a></li> </ul> <h2>4.35.5 - 15 May 2026</h2> <ul> <li>We have improved how the JavaScript bundles for the CodeQL Action are generated to avoid duplication across bundles and reduce the size of the repository by around 70%. This should have no effect on the runtime behaviour of the CodeQL Action. <a href="https://redirect.github.com/github/codeql-action/pull/3899">#3899</a></li> </ul> <!-- raw HTML omitted --> </blockquote> <p>... (truncated)</p> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/github/codeql-action/commit/e4fba868fa4b1b91e1fdab776edc8cfbe6e9fb81"><code>e4fba86</code></a> Merge pull request <a href="https://redirect.github.com/github/codeql-action/issues/4031">#4031</a> from github/update-v4.37.3-72f6a9da0</li> <li><a href="https://github.com/github/codeql-action/commit/fb50ab5d62a274adf3ef3e22cfe750ae87a0ede7"><code>fb50ab5</code></a> Update changelog for v4.37.3</li> <li><a href="https://github.com/github/codeql-action/commit/72f6a9da0def52d9193d6a758f0378b65091f8d1"><code>72f6a9d</code></a> Merge pull request <a href="https://redirect.github.com/github/codeql-action/issues/4030">#4030</a> from github/mbg/fix/no-proxy</li> <li><a href="https://github.com/github/codeql-action/commit/3b5ee58597653d9cc6785f3f1277f796d81f3646"><code>3b5ee58</code></a> Use default <code>request</code> options instead of <code>undefined</code></li> <li><a href="https://github.com/github/codeql-action/commit/bfb6be4b5ecd3650f02f530571453e8c64ef0778"><code>bfb6be4</code></a> Merge pull request <a href="https://redirect.github.com/github/codeql-action/issues/4028">#4028</a> from github/mergeback/v4.37.2-to-main-e0647621</li> <li><a href="https://github.com/github/codeql-action/commit/526ab84f9858816d9cf5f7b9df4dd5e2235f0eba"><code>526ab84</code></a> Rebuild</li> <li><a href="https://github.com/github/codeql-action/commit/d6217b9b8c14166e4851db94c11155d03bd13c07"><code>d6217b9</code></a> Update changelog and version after v4.37.2</li> <li><a href="https://github.com/github/codeql-action/commit/e0647621c2984b5ed2f768cb892365bf2a616ad1"><code>e064762</code></a> Merge pull request <a href="https://redirect.github.com/github/codeql-action/issues/4027">#4027</a> from github/update-v4.37.2-385bcdc5a</li> <li><a href="https://github.com/github/codeql-action/commit/e0faed839190caa67a5cd42f1cc16246028ca3df"><code>e0faed8</code></a> Add a couple of change notes</li> <li><a href="https://github.com/github/codeql-action/commit/73aad0eaa9df172668665a150d17b8bc5a650c20"><code>73aad0e</code></a> Update changelog for v4.37.2</li> <li>Additional commits viewable in <a href="https://github.com/github/codeql-action/compare/7188fc363630916deb702c7fdcf4e481b751f97a...e4fba868fa4b1b91e1fdab776edc8cfbe6e9fb81">compare view</a></li> </ul> </details> <br /> Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting `@dependabot rebase`. [//]: # (dependabot-automerge-start) [//]: # (dependabot-automerge-end) --- <details> <summary>Dependabot commands and options</summary> <br /> You can trigger Dependabot actions by commenting on this PR: - `@dependabot rebase` will rebase this PR - `@dependabot recreate` will recreate this PR, overwriting any edits that have been made to it - `@dependabot show <dependency name> ignore conditions` will show all of the ignore conditions of the specified dependency - `@dependabot ignore <dependency name> major version` will close this group update PR and stop Dependabot creating any more for the specific dependency's major version (unless you unignore this specific dependency's major version or upgrade to it yourself) - `@dependabot ignore <dependency name> minor version` will close this group update PR and stop Dependabot creating any more for the specific dependency's minor version (unless you unignore this specific dependency's minor version or upgrade to it yourself) - `@dependabot ignore <dependency name>` will close this group update PR and stop Dependabot creating any more for the specific dependency (unless you unignore this specific dependency or upgrade to it yourself) - `@dependabot unignore <dependency name>` will remove all of the ignore conditions of the specified dependency - `@dependabot unignore <dependency name> <ignore condition>` will remove the ignore condition of the specified dependency and ignore conditions </details> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
…e#23941) Bumps [taiki-e/install-action](https://github.com/taiki-e/install-action) from 2.84.0 to 2.85.2. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/taiki-e/install-action/releases">taiki-e/install-action's releases</a>.</em></p> <blockquote> <h2>2.85.2</h2> <ul> <li> <p>Update <code>prek@latest</code> to 0.4.11.</p> </li> <li> <p>Update <code>mise@latest</code> to 2026.7.13.</p> </li> <li> <p>Update <code>kingfisher@latest</code> to 1.109.0.</p> </li> </ul> <h2>2.85.1</h2> <ul> <li> <p>Update <code>vacuum@latest</code> to 0.30.0.</p> </li> <li> <p>Update <code>uv@latest</code> to 0.11.32.</p> </li> <li> <p>Update <code>mise@latest</code> to 2026.7.12.</p> </li> <li> <p>Update <code>cyclonedx@latest</code> to 0.33.1.</p> </li> <li> <p>Update <code>cargo-neat@latest</code> to 0.5.2.</p> </li> </ul> <h2>2.85.0</h2> <ul> <li> <p>Support <code>wild</code> (alias: <code>wild-linker</code>). (<a href="https://redirect.github.com/taiki-e/install-action/pull/1949">#1949</a>)</p> </li> <li> <p>Support <code>bpf-linker</code>. (<a href="https://redirect.github.com/taiki-e/install-action/pull/1950">#1950</a>)</p> </li> <li> <p>Support <code>rafn</code>. (<a href="https://redirect.github.com/taiki-e/install-action/pull/1935">#1935</a>, thanks <a href="https://github.com/DarkWanderer"><code>@DarkWanderer</code></a>)</p> </li> <li> <p>Update <code>cargo-neat@latest</code> to 0.5.1.</p> </li> <li> <p>Update <code>zizmor@latest</code> to 1.28.0.</p> </li> <li> <p>Update <code>wasmtime@latest</code> to 47.0.2.</p> </li> <li> <p>Update <code>uv@latest</code> to 0.11.31.</p> </li> <li> <p>Update <code>syft@latest</code> to 1.49.0.</p> </li> </ul> <h2>2.84.1</h2> <ul> <li> <p>Update <code>wasmtime@latest</code> to 47.0.1.</p> </li> <li> <p>Update <code>wasm-tools@latest</code> to 1.254.0.</p> </li> <li> <p>Update <code>uv@latest</code> to 0.11.30.</p> </li> <li> <p>Update <code>mise@latest</code> to 2026.7.11.</p> </li> <li> <p>Update <code>cargo-neat@latest</code> to 0.5.0.</p> </li> <li> <p>Update <code>cargo-crap@latest</code> to 0.3.1.</p> </li> <li> <p>Update <code>biome@latest</code> to 2.5.5.</p> </li> </ul> </blockquote> </details> <details> <summary>Changelog</summary> <p><em>Sourced from <a href="https://github.com/taiki-e/install-action/blob/main/CHANGELOG.md">taiki-e/install-action's changelog</a>.</em></p> <blockquote> <h1>Changelog</h1> <p>All notable changes to this project will be documented in this file.</p> <p>This project adheres to <a href="https://semver.org">Semantic Versioning</a>.</p> <!-- raw HTML omitted --> <h2>[Unreleased]</h2> <h2>[2.85.2] - 2026-07-26</h2> <ul> <li> <p>Update <code>prek@latest</code> to 0.4.11.</p> </li> <li> <p>Update <code>mise@latest</code> to 2026.7.13.</p> </li> <li> <p>Update <code>kingfisher@latest</code> to 1.109.0.</p> </li> </ul> <h2>[2.85.1] - 2026-07-25</h2> <ul> <li> <p>Update <code>vacuum@latest</code> to 0.30.0.</p> </li> <li> <p>Update <code>uv@latest</code> to 0.11.32.</p> </li> <li> <p>Update <code>mise@latest</code> to 2026.7.12.</p> </li> <li> <p>Update <code>cyclonedx@latest</code> to 0.33.1.</p> </li> <li> <p>Update <code>cargo-neat@latest</code> to 0.5.2.</p> </li> </ul> <h2>[2.85.0] - 2026-07-23</h2> <ul> <li> <p>Support <code>wild</code> (alias: <code>wild-linker</code>). (<a href="https://redirect.github.com/taiki-e/install-action/pull/1949">#1949</a>)</p> </li> <li> <p>Support <code>bpf-linker</code>. (<a href="https://redirect.github.com/taiki-e/install-action/pull/1950">#1950</a>)</p> </li> <li> <p>Support <code>rafn</code>. (<a href="https://redirect.github.com/taiki-e/install-action/pull/1935">#1935</a>, thanks <a href="https://github.com/DarkWanderer"><code>@DarkWanderer</code></a>)</p> </li> <li> <p>Update <code>cargo-neat@latest</code> to 0.5.1.</p> </li> <li> <p>Update <code>zizmor@latest</code> to 1.28.0.</p> </li> <li> <p>Update <code>wasmtime@latest</code> to 47.0.2.</p> </li> <li> <p>Update <code>uv@latest</code> to 0.11.31.</p> </li> <li> <p>Update <code>syft@latest</code> to 1.49.0.</p> </li> </ul> <!-- raw HTML omitted --> </blockquote> <p>... (truncated)</p> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/taiki-e/install-action/commit/41049aa56687c35e0afa74eed4f09cec4f9afabf"><code>41049aa</code></a> Release 2.85.2</li> <li><a href="https://github.com/taiki-e/install-action/commit/dfcf36552b1b743e910b9b9b6b114a08e56a1e5e"><code>dfcf365</code></a> Update <code>prek@latest</code> to 0.4.11</li> <li><a href="https://github.com/taiki-e/install-action/commit/eea03ccfa855b965a301aa52424915fddc32f855"><code>eea03cc</code></a> Update <code>mise@latest</code> to 2026.7.13</li> <li><a href="https://github.com/taiki-e/install-action/commit/81ca2feb84748ed3fa514707ee1004f4f3ad715a"><code>81ca2fe</code></a> Update martin manifest</li> <li><a href="https://github.com/taiki-e/install-action/commit/cc90ed04bc1a258ea90a83789f24b759a77c9671"><code>cc90ed0</code></a> Update <code>kingfisher@latest</code> to 1.109.0</li> <li><a href="https://github.com/taiki-e/install-action/commit/55639a3362f508fda5154d3451072205f5c32ca6"><code>55639a3</code></a> Update cargo-shear manifest</li> <li><a href="https://github.com/taiki-e/install-action/commit/3d7d7cd5ac7f994c1892ae0c06165095b9139094"><code>3d7d7cd</code></a> Release 2.85.1</li> <li><a href="https://github.com/taiki-e/install-action/commit/d09ccb4fe2105ac9ead6376f7a423ff88defeb57"><code>d09ccb4</code></a> Update <code>vacuum@latest</code> to 0.30.0</li> <li><a href="https://github.com/taiki-e/install-action/commit/ac43dee1a92e482c0e244bdb0bb126908159e245"><code>ac43dee</code></a> Update <code>uv@latest</code> to 0.11.32</li> <li><a href="https://github.com/taiki-e/install-action/commit/49b16979f38f30e44b1f68af9115be9a6ffc5215"><code>49b1697</code></a> Update prek manifest</li> <li>Additional commits viewable in <a href="https://github.com/taiki-e/install-action/compare/a6b2e2dcd845ddd7f509ce4f3ed3d922b80cc5d9...41049aa56687c35e0afa74eed4f09cec4f9afabf">compare view</a></li> </ul> </details> <br /> [](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores) Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting `@dependabot rebase`. [//]: # (dependabot-automerge-start) [//]: # (dependabot-automerge-end) --- <details> <summary>Dependabot commands and options</summary> <br /> You can trigger Dependabot actions by commenting on this PR: - `@dependabot rebase` will rebase this PR - `@dependabot recreate` will recreate this PR, overwriting any edits that have been made to it - `@dependabot show <dependency name> ignore conditions` will show all of the ignore conditions of the specified dependency - `@dependabot ignore this major version` will close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this minor version` will close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this dependency` will close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself) </details> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Bumps [actions/stale](https://github.com/actions/stale) from 10.4.0 to 11.0.0. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/actions/stale/releases">actions/stale's releases</a>.</em></p> <blockquote> <h2>v11.0.0</h2> <h2>What's Changed</h2> <h3>Enhancement</h3> <ul> <li>Migrate to ESM and update dependencies by <a href="https://github-grid.enterprise.slack.com/team/U08CVLQ4JKE"><code>@chiranjib-swain</code></a> in <a href="https://redirect.github.com/actions/stale/pull/1350">actions/stale#1350</a></li> </ul> <h3>Dependency Update</h3> <ul> <li>Override brace-expansion to 5.0.8 to address 24 high-severity dependency vulnerabilities by <a href="https://github.com/dependabot"><code>@dependabot</code></a> in <a href="https://redirect.github.com/actions/stale/pull/1351">actions/stale#1351</a></li> </ul> <p><strong>Full Changelog</strong>: <a href="https://github.com/actions/stale/compare/v10...v11.0.0">https://github.com/actions/stale/compare/v10...v11.0.0</a></p> </blockquote> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/actions/stale/commit/4391f3da665fdf50b6810c1a66712fb9ba21aa93"><code>4391f3d</code></a> Fix 24 high severity vulnerabilities by overriding brace-expansion to 5.0.8 (...</li> <li><a href="https://github.com/actions/stale/commit/eaf9131fae5eafd0c31a64ebe3a2e183266fec48"><code>eaf9131</code></a> refactor: update imports to use ES module syntax and improve test structure (...</li> <li>See full diff in <a href="https://github.com/actions/stale/compare/1e223db275d687790206a7acac4d1a11bd6fe629...4391f3da665fdf50b6810c1a66712fb9ba21aa93">compare view</a></li> </ul> </details> <br /> [](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores) Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting `@dependabot rebase`. [//]: # (dependabot-automerge-start) [//]: # (dependabot-automerge-end) --- <details> <summary>Dependabot commands and options</summary> <br /> You can trigger Dependabot actions by commenting on this PR: - `@dependabot rebase` will rebase this PR - `@dependabot recreate` will recreate this PR, overwriting any edits that have been made to it - `@dependabot show <dependency name> ignore conditions` will show all of the ignore conditions of the specified dependency - `@dependabot ignore this major version` will close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this minor version` will close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this dependency` will close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself) </details> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Bumps [base64](https://github.com/marshallpierce/rust-base64) from 0.22.1 to 0.23.0. <details> <summary>Changelog</summary> <p><em>Sourced from <a href="https://github.com/marshallpierce/rust-base64/blob/master/RELEASE-NOTES.md">base64's changelog</a>.</em></p> <blockquote> <h1>0.23.0</h1> <ul> <li>Added more consts for preconfigured configs and engines</li> <li>Make DecodeError::InvalidLastSymbol more clear by including the decoded value</li> <li>Added SIMD-accelerated engines behind the default-on <code>simd-unsafe</code> feature: <code>Simd</code> picks the best instruction set at runtime (AVX2 on <code>x86_64</code>, NEON on <code>aarch64</code>) and falls back to the scalar <code>GeneralPurpose</code> engine, while <code>Avx2</code> and <code>Neon</code> target one instruction set with no runtime detection and work in <code>no_std</code>. The engines support the standard and URL-safe alphabets.</li> <li>Update MSRV to 1.71.0</li> <li>Add support for custom padding symbols</li> </ul> </blockquote> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/marshallpierce/rust-base64/commit/9e9220a4166f628de7c8803289e120ae1e944f78"><code>9e9220a</code></a> v0.23.0</li> <li><a href="https://github.com/marshallpierce/rust-base64/commit/870326ec592eebde9d6bfe4c5d8130c591273e9c"><code>870326e</code></a> Merge pull request <a href="https://redirect.github.com/marshallpierce/rust-base64/issues/306">#306</a> from marshallpierce/mp/trailing-bits-docs</li> <li><a href="https://github.com/marshallpierce/rust-base64/commit/fbec5f1050f9fc16e6a826ebabaa2b7b0644bd67"><code>fbec5f1</code></a> Document no trailing trailing bits</li> <li><a href="https://github.com/marshallpierce/rust-base64/commit/0a23549968f059b53cf39e96eba8f46779f322a7"><code>0a23549</code></a> Merge pull request <a href="https://redirect.github.com/marshallpierce/rust-base64/issues/305">#305</a> from marshallpierce/mp/edition-2021</li> <li><a href="https://github.com/marshallpierce/rust-base64/commit/f10b7e20614135aa61289140683fc93e5a45d338"><code>f10b7e2</code></a> Update deps & edition</li> <li><a href="https://github.com/marshallpierce/rust-base64/commit/9d21a598860645cb6290940e7a43033bc43ebd74"><code>9d21a59</code></a> Merge pull request <a href="https://redirect.github.com/marshallpierce/rust-base64/issues/304">#304</a> from marshallpierce/mp/custom-padding-rebase</li> <li><a href="https://github.com/marshallpierce/rust-base64/commit/f70bad2caaa85350b95d988bfb9a0997e824bfd8"><code>f70bad2</code></a> Support custom padding symbols</li> <li><a href="https://github.com/marshallpierce/rust-base64/commit/684d79cd3deb8dfd5323619634c75bc0ff6edfd9"><code>684d79c</code></a> Merge pull request <a href="https://redirect.github.com/marshallpierce/rust-base64/issues/301">#301</a> from marshallpierce/mp/simd-gardening</li> <li><a href="https://github.com/marshallpierce/rust-base64/commit/5bf66f2646c6fd99e1b18bf4e2bb0a47d34e1eaa"><code>5bf66f2</code></a> Merge pull request <a href="https://redirect.github.com/marshallpierce/rust-base64/issues/284">#284</a> from AbeZbm/add-tests</li> <li><a href="https://github.com/marshallpierce/rust-base64/commit/d3831cfbf7dafe226a8383450c3410e2f67c826a"><code>d3831cf</code></a> Followups to SIMD work</li> <li>Additional commits viewable in <a href="https://github.com/marshallpierce/rust-base64/compare/v0.22.1...v0.23.0">compare view</a></li> </ul> </details> <br /> [](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores) Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting `@dependabot rebase`. [//]: # (dependabot-automerge-start) [//]: # (dependabot-automerge-end) --- <details> <summary>Dependabot commands and options</summary> <br /> You can trigger Dependabot actions by commenting on this PR: - `@dependabot rebase` will rebase this PR - `@dependabot recreate` will recreate this PR, overwriting any edits that have been made to it - `@dependabot show <dependency name> ignore conditions` will show all of the ignore conditions of the specified dependency - `@dependabot ignore this major version` will close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this minor version` will close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this dependency` will close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself) </details> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Bumps [astral-sh/setup-uv](https://github.com/astral-sh/setup-uv) from 8.3.2 to 9.0.0. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/astral-sh/setup-uv/releases">astral-sh/setup-uv's releases</a>.</em></p> <blockquote> <h2>v9.0.0 🌈 Change <code>prune-cache</code> default to <code>false</code></h2> <h2>Changes</h2> <p>This release disables the default cache cache pruning to ease the load on the PyPi infrastructure. Since users might experience more GitHub Actions cache usage which might result in higher costs this is marked as a breaking change. To read more on why we did this (now) you can read the detailed analysis and reasoning in <a href="https://redirect.github.com/astral-sh/setup-uv/issues/967">#967</a></p> <p>Besides this big breaking change we also have a small bugfix while building caches for linux distributions that behave a big different than the "big ones" and a speed up in version resolution by only reading the version manifest until a matching version is found saving runtime and network bandwith.</p> <h2>🚨 Breaking changes</h2> <ul> <li>Change <code>prune-cache</code> default to <code>false</code> <a href="https://github.com/charliermarsh"><code>@charliermarsh</code></a> (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/967">#967</a>)</li> </ul> <h2>🐛 Bug fixes</h2> <ul> <li>fix: fall back to distribution ID when os-release has no version field <a href="https://github.com/cxzhong"><code>@cxzhong</code></a> (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/961">#961</a>)</li> </ul> <h2>🚀 Enhancements</h2> <ul> <li>Speed up version client by partial response reads <a href="https://github.com/eifinger"><code>@eifinger</code></a> (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/807">#807</a>)</li> </ul> <h2>🧰 Maintenance</h2> <ul> <li>chore: update known checksums for 0.11.30 @<a href="https://github.com/apps/github-actions">github-actions[bot]</a> (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/968">#968</a>)</li> <li>chore: update known checksums for 0.11.29 @<a href="https://github.com/apps/github-actions">github-actions[bot]</a> (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/960">#960</a>)</li> </ul> <h2>📚 Documentation</h2> <ul> <li>docs: update version references to v8.3.2 @<a href="https://github.com/apps/github-actions">github-actions[bot]</a> (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/949">#949</a>)</li> </ul> <h2>⬆️ Dependency updates</h2> <ul> <li>chore(deps): roll up Dependabot updates <a href="https://github.com/eifinger"><code>@eifinger</code></a> (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/970">#970</a>)</li> <li>chore(deps): roll up Dependabot updates <a href="https://github.com/eifinger"><code>@eifinger</code></a> (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/962">#962</a>)</li> </ul> </blockquote> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/astral-sh/setup-uv/commit/c771a70e6277c0a99b617c7a806ffedaca235ff9"><code>c771a70</code></a> chore(deps): roll up Dependabot updates (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/970">#970</a>)</li> <li><a href="https://github.com/astral-sh/setup-uv/commit/2f537ca87c1ffa233ca2a1b84815388e3e42d845"><code>2f537ca</code></a> chore: update known checksums for 0.11.30 (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/968">#968</a>)</li> <li><a href="https://github.com/astral-sh/setup-uv/commit/2269552d547df6f50e57442326930d30d943afe3"><code>2269552</code></a> Speed up version client by partial response reads (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/807">#807</a>)</li> <li><a href="https://github.com/astral-sh/setup-uv/commit/47a7f4fb2e900d6c33a5b5f231fa21dbfaeba52f"><code>47a7f4f</code></a> Change <code>prune-cache</code> default to <code>false</code> (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/967">#967</a>)</li> <li><a href="https://github.com/astral-sh/setup-uv/commit/71966eff34a27b0a62ed4b9f6f6e383e071b1bb5"><code>71966ef</code></a> chore(deps): roll up Dependabot updates (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/962">#962</a>)</li> <li><a href="https://github.com/astral-sh/setup-uv/commit/f12b1f0a84bd6dc2331b36b2bbdbb1d1e617dbcc"><code>f12b1f0</code></a> fix: fall back to distribution ID when os-release has no version field (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/961">#961</a>)</li> <li><a href="https://github.com/astral-sh/setup-uv/commit/ecd24dd710f2fb0dca1693a67af11fc4a5c5ec84"><code>ecd24dd</code></a> chore: update known checksums for 0.11.29 (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/960">#960</a>)</li> <li><a href="https://github.com/astral-sh/setup-uv/commit/6a191366842ac1502ba6c07e9b5acd5c2d9d8db3"><code>6a19136</code></a> docs: update version references to v8.3.2 (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/949">#949</a>)</li> <li>See full diff in <a href="https://github.com/astral-sh/setup-uv/compare/11f9893b081a58869d3b5fccaea48c9e9e46f990...c771a70e6277c0a99b617c7a806ffedaca235ff9">compare view</a></li> </ul> </details> <br /> [](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores) Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting `@dependabot rebase`. [//]: # (dependabot-automerge-start) [//]: # (dependabot-automerge-end) --- <details> <summary>Dependabot commands and options</summary> <br /> You can trigger Dependabot actions by commenting on this PR: - `@dependabot rebase` will rebase this PR - `@dependabot recreate` will recreate this PR, overwriting any edits that have been made to it - `@dependabot show <dependency name> ignore conditions` will show all of the ignore conditions of the specified dependency - `@dependabot ignore this major version` will close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this minor version` will close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this dependency` will close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself) </details> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
…fy to be textbook like as possible (apache#23761) ## Which issue does this PR close? N/A ## Rationale for this change SortMergeJoin bitwise stream implementation is very complex and hard to understand while on paper it should be pretty simple. the reason for that is we have to store state between polls (we had `boundary`) and handle the case where both can get `Poll::Pending` from child and `Poll::Ready` from child which further complicate the code ## What changes are included in this PR? 1. Move to async generators 2. Rewrote main loop to be textbook like as possible (this was entirely written by Claude Fable, sorry, I tried manually but the code was too complex to hold in my head 😅 ) ## Are these changes tested? existing tests ## Are there any user-facing changes? The join_time now includes the time to read from the async spill stream between pending which is arguable more correct since this time is part of the operator, although long waits between pending calls will be counted in the op `join_time` while the alternative is not counting the read from file and decoding...
…lator (apache#23946) ## Rationale for this change Follow-up to post-merge review feedback from @neilconway on apache#23913. ## What changes are included in this PR? - Use the fast foldhash `RandomState` for the distinct-value count map instead of the standard library's default SipHash (the shared `GenericDistinctBuffer` already does this; the merged fix regressed to SipHash on this hot path). - Use `estimate_memory_size` in `size()` instead of a hand-rolled capacity calculation. - Adopt the null-free fast path in `update_batch`/`retract_batch` (skip per-element validity checks when the input has no nulls), mirroring `GenericDistinctBuffer`. - Add `ORDER BY` to the sliding-window regression test so its row order is deterministic. ## Are these changes tested? Yes — existing percentile unit tests and the full `aggregate.slt` pass. ## Are there any user-facing changes? No.
…pstream now carries DataFusion 55.1.0 adds the same start == end guard (with try_extend_nulls) immediately above ours, so the Spice copy was unreachable and its deprecated extend_nulls call warned. The regression tests stay.
- tests use StatisticsContext and PruningPredicateBuilder instead of the deprecated partition_statistics / PruningPredicate::try_new - proto: expect(unused_variables) instead of allow (clippy::allow_attributes) - rustdoc: fix unresolved and private intra-doc links - rustfmt and taplo formatting
The aggregate-unprojection arm used an `if let` match guard, which is not stable on the workspace's rust-version (1.94.0): `cargo +1.94.1 check` fails with E0658. Move the check into the unqualified-column arm as a let-chain; same behaviour, and the workspace checks on 1.94.1 again.
estimate_join_cardinality carried each input's column statistics through an inner, left, right or full join unchanged, sum_value included. A join repeats a row once per match and drops the rows that match nothing, so the input's sum says nothing about the output's; kept exact, it let AggregateStatistics answer SUM over the join with the input table's total (a wrong result with no error). Semi and anti joins already drop it, and the cross join scales it.
ListingTable materializes metadata columns (_last_modified, _size, _location) from each object's ObjectMeta, but pruning only used partition columns. A predicate such as `_last_modified > <watermark>` fell through to a row-level FilterExec, so every file was opened and read before its rows were dropped. For a compressed row format (e.g. jsonl.gz) this decompressed the whole object on every scan. Prune the listing by these predicates before any file is opened: - `pruned_partition_list_with_metadata` filters the listed ObjectMeta set by evaluating the metadata-only predicate against each object. The existing `pruned_partition_list` delegates to it with empty metadata arguments, so its signature and behavior are unchanged. - `filter_by_metadata` evaluates the predicate via create_physical_expr against a one-row batch built from MetadataColumn::to_scalar_value, giving correct >, >=, <, <=, =, BETWEEN, IN, LIKE, cast and NULL semantics. A NULL (unknown) result prunes the file, matching SQL WHERE semantics. - `supports_filters_pushdown` reports a metadata-only conjunct as Exact, so the redundant FilterExec is dropped. A mixed metadata+data predicate, or one under OR/NOT, stays Inexact and keeps a residual filter (correct, not optimal). Metadata pruning is orthogonal to partition pruning and applies to both partitioned and unpartitioned tables, and to all file formats. Reproduces spiceai/spiceai#14264.
A caller that obtains an ObjectMeta from a HEAD on a known key (rather than a listing) can apply the identical metadata-column prune before opening the file, and report the same predicates as Exact.
Compile the metadata predicate once per listing instead of per file (MetadataPredicate), threading the caller's session ExecutionProps through so filters using session variables still evaluate correctly. Deduplicate metadata-column-name extraction between scan_with_args and supports_filters_pushdown, and add test coverage: unit-level predicate tests, a listing-helper integration test, and a ListingTable plan-level test asserting the Exact pushdown marking and scan wiring actually prune file groups and residual filters as expected. Also documents the Exact-pushdown contract so a TableProvider wrapping ListingTable with its own scan path doesn't inherit it without enforcing it.
|
CI note: this repo's GitHub workflows never start (dispatches stay queued with 0 jobs on spiceai-54 too), so only the license check runs. The CI commands were run locally on this head: |
fix(physical-plan): drop input sums from inner and outer join statistics
fix(unparser): spell a Date32 literal's cast with the dialect's date type (refs spiceai/spiceai#14491)
Brings across the Spice patches on the 54 line that were missing here: the unparser spelling a Date32 literal's cast with the dialect's date type (#237), and pruning ListingTable file listings by metadata-column predicates (`_last_modified`, `_size`, `_location`), which `supports_filters_pushdown` reports `Exact`. table.rs conflicted with DataFusion 55's declared output partitioning. Both behaviors are kept. `scan_with_args` separates out the metadata filters and passes them to `list_files_for_scan_with_metadata`, which dispatches the way 55's `list_files_for_scan` does. The regular scan prunes by metadata while listing, before the file limit. The declared partitioning path assigns files to partitions before any filter runs, then removes files that fail the metadata predicate from their assigned group, the same way it handles partition filters. A file's partition therefore never depends on the query. Limits are still not pushed to listing when output partitioning is declared or residual filters remain. `MetadataPredicate` now passes the `PhysicalPlanningContext` argument that 55's `create_physical_expr` requires. Adds a test for a metadata filter on a table with declared output partitioning. It checks the filter is `Exact`, that no FilterExec is planned, that the partition count is kept, that the excluded file is pruned, and that the correct rows come back.
Merge spiceai-54 into spiceai-55-patches-2 (#237, metadata-column listing pruning)
|
This looks superseded — worth a decision before any more work goes into it.
So merging this as-is would re-litigate #240 through a side door, and its stated purpose — bringing 55.1.0 in — is already done by the 55.2.0-rc1 merge. Recommend closing this in favour of #240, or retargeting it to just the piece that is genuinely still missing. Either way it is an author/maintainer call, not one to make from a babysitting pass — leaving it untouched. |

Merges upstream DataFusion
55.1.0into the Spice line (spiceai-55, cut fromspiceai-54); conflicts are resolved in the merge commit and recorded in its message.On top of the merge:
datafusion-sparkquote.rsimports its expression types fromdatafusion_expr, so the crate builds without itscorefeature (backport; 55.1.0 has the broken import).array_any_valueempty-element guard is dropped: 55.1.0 carries the same guard.if letmatch guard, which broke the workspace MSRV (1.94).db730b42.Pinned by spiceai/spiceai#14612 (tracking: spiceai/spiceai#13570). Supersedes #236.