Skip to content

Merge upstream DataFusion 55.1.0 into spiceai-55 - #238

Closed
krinart wants to merge 906 commits into
spiceai-55from
spiceai-55-patches-2
Closed

krinart wants to merge 906 commits into
spiceai-55from
spiceai-55-patches-2

Conversation

@krinart

@krinart krinart commented Sep 30, 2026

Copy link
Copy Markdown

Merges upstream DataFusion 55.1.0 into the Spice line (spiceai-55, cut from spiceai-54); conflicts are resolved in the merge commit and recorded in its message.

On top of the merge:

  • datafusion-spark quote.rs imports its expression types from datafusion_expr, so the crate builds without its core feature (backport; 55.1.0 has the broken import).
  • The array_any_value empty-element guard is dropped: 55.1.0 carries the same guard.
  • The ORDER BY unprojection no longer uses an if let match guard, which broke the workspace MSRV (1.94).
  • CI lint (clippy, rustdoc, rustfmt, taplo) is clean; builds against the arrow-rs fork at db730b42.

Pinned by spiceai/spiceai#14612 (tracking: spiceai/spiceai#13570). Supersedes #236.

viirya and others added 30 commits July 26, 2026 00:06
…lter (apache#23901)

## Which issue does this PR close?

- Closes apache#23900.

## Rationale for this change

`push_down_filter` infers equi-key predicates across a join's ON keys
and pushes them to the opposite side. For a null-aware join (the
`LeftAnti` join produced by `NOT IN` with a nullable subquery), an outer
predicate on the left key like `outer.id > 5` is rewritten to `sub.id >
5` and pushed onto the subquery input. Since the inferred predicate must
be null-rejecting to be pushed, this drops the subquery's NULL rows and
breaks the three-valued `NOT IN` semantics — a NULL in the subquery key
must reach the join so the result is empty.

Same class of bug as apache#23848, in a different rule.

## What changes are included in this PR?

- Skip predicate inference in `infer_join_predicates` when
`join.null_aware` is set (mirrors the apache#23848 guard on
`FilterNullJoinKeys`).
- A `push_down_filter` unit test asserting no predicate is inferred onto
the subquery side of a null-aware `LeftAnti` join.
- SLT coverage for the failing query, plus a `prefer_hash_join = false`
/ multi-partition variant.

## Are these changes tested?

Yes — new unit test (verified it fails without the guard) and SLT cases.
The full optimizer lib suite passes.

## Are there any user-facing changes?

No, aside from the correctness fix.
## Which issue does this PR close?

- Part of apache#19241.
- Stacked on [apache#23311](apache#23311).
- Next in stack: apache#23015.
- Extracted from apache#19390.

## Rationale for this change

For very small `IN` lists, building or probing a hash table can be more
work than just comparing the input value with each constant.

For example, for `x IN (10, 20, 30)`, the fast path can behave like:

```text
x == 10 OR x == 20 OR x == 30
```

Because the list is tiny, those comparisons are cheap. The
implementation stores the constants in a fixed-size array and checks
them with a compact comparison chain.

“Branchless” here means the comparisons are combined without stopping at
the first match. That can be faster for these small fixed-width lists
because the CPU gets a predictable sequence of simple operations instead
of hash-table setup and probe logic.

For primitive values that are not already plain unsigned integers, this
PR keeps the logical Arrow type explicit and uses a matching same-width
comparison representation only inside the branchless filter. For
example, `Float16` uses `UInt16` storage, `Float32` uses `UInt32`
storage, and `TimestampNanosecond` uses `UInt64` storage. `Decimal128`
and `IntervalMonthDayNano` use their own 16-byte native representation.
This preserves bit-pattern equality while relying on Arrow's native
primitive compatibility rules: timestamp timezone metadata and
Decimal128 precision/scale metadata may differ, while incompatible
primitive representations remain rejected.

## What changes are included in this PR?

- Adds a const-generic `BranchlessFilter` for small primitive `IN`
lists.
- Adds thresholds for when this path is used:
  - up to 16 values for 1-byte types
  - up to 8 values for 2-byte types
  - up to 32 values for 4-byte types
  - up to 16 values for 8-byte types
  - up to 4 values for 16-byte types
- Keeps dispatch concrete and explicit in `strategy.rs`.
- Maps each optimized logical type to the comparison representation used
by the branchless filter:
  - `Int8` -> `UInt8`
  - `Int16`, `Float16` -> `UInt16`
  - `Int32`, `Float32`, `Date32`, `Time32` -> `UInt32`
- `Int64`, `Float64`, `Date64`, `Time64`, `Timestamp`, `Duration` ->
`UInt64`
- `Decimal128`, `IntervalMonthDayNano` -> their native 16-byte
representation
- Leaves larger 1-byte and 2-byte lists on the existing bitmap filters.
- Leaves larger 4-byte and 8-byte lists on the existing hash/generic
paths.
- Leaves wider primitive types such as `Decimal256` and unsupported
complex types on the generic path.
- Keeps the same `IN` / `NOT IN` null behavior as the rest of the stack.
- Adds focused coverage for branchless null handling, signed boundary
values, slices, Float16/Float32/Float64 bit patterns, compatible
timestamp/Decimal128 metadata, incompatible timestamp units,
IntervalMonthDayNano values, and same-width wrong-type probe rejection.

## Are these changes tested?

Yes.

- `cargo fmt --all -- --check`
- `cargo test -p datafusion-physical-expr expressions::in_list --lib`
- `cargo test -p datafusion-physical-expr --bench in_list_strategy
--no-run`
- `cargo clippy --all-targets --all-features -- -D warnings`

## Are there any user-facing changes?

No. This is an internal performance optimization only.

## Local benchmark snapshot

Built and run with `release-nonlto`, filtered to the relevant small
primitive-list rows:

```bash
cargo bench -p datafusion-physical-expr --profile release-nonlto --bench in_list_strategy -- <filter> --save-baseline <baseline>
```

Filters used: `narrow_integer`, `primitive/i32/small_list`,
`primitive/i64/small_list`, `f32/small_list`, `timestamp_ns/small_list`,
and `interval_month_day_nano/small_list`.

Method: directly compared Criterion's raw sample minima (`min(time /
iterations)`) from `sample.json`. Lower is better; changes within +/-5%
are treated as noise.

Compared baselines:
[apache#23311](apache#23311) ->
[apache#23014](apache#23014)

Relevant scope: small primitive-list rows.

Summary: 39 relevant rows, 28 faster, 0 slower, 11 within +/-5%.

Largest relevant deltas:

| Benchmark | Before | After | Change |
|---|---:|---:|---:|
| `timestamp_ns/small_list/list=4/match=50%` | 46.55 us | 3.17 us |
-93.2% (14.69x faster) |
| `f32/small_list/list=4/match=50%` | 33.93 us | 3.04 us | -91.0%
(11.15x faster) |
| `primitive/i32/small_list/list=4/match=50%` | 32.63 us | 3.08 us |
-90.5% (10.58x faster) |
| `primitive/i64/small_list/list=4/match=50%` | 33.55 us | 3.18 us |
-90.5% (10.54x faster) |
| `timestamp_ns/small_list/list=4/match=0%` | 19.57 us | 3.18 us |
-83.8% (6.16x faster) |
| `f32/small_list/list=4/match=0%` | 18.14 us | 3.05 us | -83.2% (5.95x
faster) |
| `primitive/i32/small_list/list=4/match=0%` | 17.00 us | 3.04 us |
-82.1% (5.59x faster) |
| `primitive/i64/small_list/list=4/match=0%` | 17.12 us | 3.22 us |
-81.2% (5.31x faster) |
| `primitive/i32/small_list/list=16/match=50%/NOT_IN` | 31.98 us | 7.26
us | -77.3% (4.41x faster) |
| `nulls/primitive/i32/small_list/list=16/match=50%/nulls=20%` | 29.35
us | 7.32 us | -75.1% (4.01x faster) |
| `timestamp_ns/small_list/list=16/match=50%` | 45.32 us | 11.79 us |
-74.0% (3.84x faster) |
| `nulls/primitive/i32/small_list/list=16/match=50%/nulls=50%` | 25.89
us | 7.31 us | -71.8% (3.54x faster) |
| `nulls/primitive/i32/small_list/list=16/match=50%/nulls=20%/NOT_IN` |
26.05 us | 7.42 us | -71.5% (3.51x faster) |
| `interval_month_day_nano/small_list/list=4/match=50%` | 52.94 us |
15.52 us | -70.7% (3.41x faster) |
| `f32/small_list/list=32/match=50%` | 38.78 us | 13.27 us | -65.8%
(2.92x faster) |
| `primitive/i64/small_list/list=16/match=50%` | 29.46 us | 11.76 us |
-60.1% (2.50x faster) |

<details>
<summary>Full relevant table (39 rows)</summary>

| Benchmark | Before | After | Change |
|---|---:|---:|---:|
| `narrow_integer/u8/list=4/match=0%` | 3.86 us | 2.79 us | -27.8%
(1.38x faster) |
| `narrow_integer/u8/list=4/match=50%` | 3.84 us | 2.78 us | -27.7%
(1.38x faster) |
| `narrow_integer/u8/list=16/match=0%` | 3.88 us | 3.85 us | -0.8%
(within noise) |
| `narrow_integer/u8/list=16/match=50%` | 3.84 us | 3.86 us | +0.5%
(within noise) |
| `narrow_integer/i16/list=4/match=0%` | 3.93 us | 3.18 us | -19.1%
(1.24x faster) |
| `narrow_integer/i16/list=4/match=50%` | 3.92 us | 3.16 us | -19.5%
(1.24x faster) |
| `narrow_integer/i16/list=64/match=0%` | 3.96 us | 3.82 us | -3.5%
(within noise) |
| `narrow_integer/i16/list=64/match=50%` | 3.91 us | 3.80 us | -2.9%
(within noise) |
| `narrow_integer/i16/list=256/match=0%` | 3.90 us | 3.81 us | -2.5%
(within noise) |
| `narrow_integer/i16/list=256/match=50%` | 3.97 us | 3.81 us | -4.1%
(within noise) |
| `narrow_integer/f16/list=4/match=0%` | 3.87 us | 3.16 us | -18.5%
(1.23x faster) |
| `narrow_integer/f16/list=4/match=50%` | 3.94 us | 3.15 us | -20.2%
(1.25x faster) |
| `narrow_integer/f16/list=64/match=0%` | 3.87 us | 3.84 us | -0.6%
(within noise) |
| `narrow_integer/f16/list=64/match=50%` | 3.93 us | 3.85 us | -1.9%
(within noise) |
| `narrow_integer/f16/list=256/match=0%` | 3.90 us | 3.84 us | -1.5%
(within noise) |
| `narrow_integer/f16/list=256/match=50%` | 3.87 us | 3.91 us | +1.2%
(within noise) |
| `nulls/narrow_integer/u8/list=16/match=50%/nulls=20%` | 3.92 us | 4.02
us | +2.5% (within noise) |
| `primitive/i32/small_list/list=4/match=0%` | 17.00 us | 3.04 us |
-82.1% (5.59x faster) |
| `primitive/i32/small_list/list=4/match=50%` | 32.63 us | 3.08 us |
-90.5% (10.58x faster) |
| `primitive/i32/small_list/list=32/match=0%` | 16.34 us | 13.33 us |
-18.5% (1.23x faster) |
| `primitive/i32/small_list/list=32/match=50%` | 31.17 us | 13.31 us |
-57.3% (2.34x faster) |
| `primitive/i32/small_list/list=16/match=50%/NOT_IN` | 31.98 us | 7.26
us | -77.3% (4.41x faster) |
| `nulls/primitive/i32/small_list/list=16/match=50%/nulls=20%` | 29.35
us | 7.32 us | -75.1% (4.01x faster) |
| `nulls/primitive/i32/small_list/list=16/match=50%/nulls=20%/NOT_IN` |
26.05 us | 7.42 us | -71.5% (3.51x faster) |
| `nulls/primitive/i32/small_list/list=16/match=50%/nulls=50%` | 25.89
us | 7.31 us | -71.8% (3.54x faster) |
| `primitive/i64/small_list/list=4/match=0%` | 17.12 us | 3.22 us |
-81.2% (5.31x faster) |
| `primitive/i64/small_list/list=4/match=50%` | 33.55 us | 3.18 us |
-90.5% (10.54x faster) |
| `primitive/i64/small_list/list=16/match=0%` | 16.34 us | 11.93 us |
-27.0% (1.37x faster) |
| `primitive/i64/small_list/list=16/match=50%` | 29.46 us | 11.76 us |
-60.1% (2.50x faster) |
| `f32/small_list/list=4/match=0%` | 18.14 us | 3.05 us | -83.2% (5.95x
faster) |
| `f32/small_list/list=4/match=50%` | 33.93 us | 3.04 us | -91.0%
(11.15x faster) |
| `f32/small_list/list=32/match=0%` | 22.05 us | 13.43 us | -39.1%
(1.64x faster) |
| `f32/small_list/list=32/match=50%` | 38.78 us | 13.27 us | -65.8%
(2.92x faster) |
| `timestamp_ns/small_list/list=4/match=0%` | 19.57 us | 3.18 us |
-83.8% (6.16x faster) |
| `timestamp_ns/small_list/list=4/match=50%` | 46.55 us | 3.17 us |
-93.2% (14.69x faster) |
| `timestamp_ns/small_list/list=16/match=0%` | 19.73 us | 12.07 us |
-38.8% (1.63x faster) |
| `timestamp_ns/small_list/list=16/match=50%` | 45.32 us | 11.79 us |
-74.0% (3.84x faster) |
| `interval_month_day_nano/small_list/list=4/match=0%` | 20.12 us |
13.20 us | -34.4% (1.52x faster) |
| `interval_month_day_nano/small_list/list=4/match=50%` | 52.94 us |
15.52 us | -70.7% (3.41x faster) |

</details>

---------

Co-authored-by: Andrew Lamb <andrew@nerdnetworks.org>
## Which issue does this PR close?

- Closes apache#23770
- Part of apache#15914

## Rationale for this change

Spark provides [`hypot(expr1,
expr2)`](https://spark.apache.org/docs/latest/api/sql/#hypot), which
returns `sqrt(expr1^2 + expr2^2)` computed without intermediate overflow
or underflow. It was not yet implemented in `datafusion-spark` — only an
auto-generated test stub existed at `spark/math/hypot.slt` with its
query commented out.

## What changes are included in this PR?

- Add `SparkHypot` (implementing `ScalarUDFImpl`) in
`datafusion/spark/src/function/math/hypot.rs`, backed by Rust's
`f64::hypot` — the same overflow-safe algorithm as Java/Spark's
`Math.hypot`.
- Register it in `datafusion/spark/src/function/math/mod.rs`.
- Enable the `hypot.slt` sqllogictest.

The signature is `exact(Float64, Float64) -> Float64`, following the
`datafusion-spark` convention of only accepting types Spark supports.
Computation uses the Arrow `binary` kernel so NULL in either argument
propagates to a NULL result, matching Spark.

## Are these changes tested?

Yes — `datafusion/sqllogictest/test_files/spark/math/hypot.slt` covers:
- scalar Pythagorean triples (`hypot(3, 4)` → 5, `hypot(5, 12)` → 13),
- double inputs,
- NULL propagation when either argument is NULL,
- the array path (including a NULL row),
- overflow-safety: `hypot(3e200, 4e200)` stays finite, whereas a naive
`sqrt(a^2 + b^2)` would overflow to `Infinity`.

## Are there any user-facing changes?

Yes — adds the Spark-compatible `hypot` scalar function to
`datafusion-spark`. No breaking changes to public APIs.
## Which issue does this PR close?

<!--
We generally require a GitHub issue to be filed for all bug fixes and
enhancements and this helps us generate change logs for our releases.
You can link an issue to this PR using the GitHub syntax. For example
`Closes #123` indicates that this PR will close issue #123.
-->

- Part of apache#23393 .

   ## Rationale for this change
  
`SlidingMinAccumulator::size` and `SlidingMaxAccumulator::size` only
reported
the stack size of their `ScalarValue` field plus its heap payload,
ignoring the
memory held by the underlying `MovingMin` / `MovingMax` sliding-window
buffers.
For windowed `MIN`/`MAX` over string or list data, the two per-element
stacks
can hold megabytes of `ScalarValue` payload that the memory pool was
never
  told about, understating accumulator memory usage.

  ## What changes are included in this PR?

- Add a private `heap_size(elem_heap)` method to `MovingMin<T>` and
`MovingMax<T>`
that reports the two stack buffers' capacity in bytes plus each stored
    element's heap payload.
- Factor the shared implementation into a `moving_stacks_heap_size` free
helper
    so the two types stay in sync.
  - Include the buffer bytes in `SlidingMinAccumulator::size` and
    `SlidingMaxAccumulator::size` via the new method.

  ## Are these changes tested?

Yes. Two new unit tests in
`datafusion/functions-aggregate/src/min_max.rs`:

- `moving_min_max_heap_size_i32` — fixed-width `T`, verifies buffer-only
    accounting with and without pushed elements.
- `moving_min_max_heap_size_counts_elems` — `T = String`, verifies each
of the
two slots in a `(T, T)` pair contributes independently to the heap
payload
    (mirroring the two independent `Clone`s made by `push`).
## Which issue does this PR close?

- Related to
apache#21882 (comment).

## Rationale for this change

PR apache#21882 adds custom spill file support. 

Having an example showing how a downstream application can use this new
API will help make sure the API is good enough for our needs

## What changes are included in this PR?

This PR adds a small example showing how users can back spill files with
an `ObjectStore`, using a local object store for a runnable example
while keeping the implementation applicable to remote stores.

## Are these changes tested?

Yes by CI

## Are there any user-facing changes?

Yes. This adds a new example for configuring ObjectStore-backed spill
files.
## Which issue does this PR close?

<!--
We generally require a GitHub issue to be filed for all bug fixes and
enhancements and this helps us generate change logs for our releases.
You can link an issue to this PR using the GitHub syntax. For example
`Closes #123` indicates that this PR will close issue #123.
-->

- Closes #N/A.

## Rationale for this change

<!--
Why are you proposing this change? If this is already explained clearly
in the issue then this section is not needed.
Explaining clearly why changes are proposed helps reviewers understand
your changes and offer better suggestions for fixes.
-->

Since DataFusion doesn't typically guartantee wire format compatibility,
I remove the backward compatiblity shim that introduced in PR apache#23189

FYI:
apache#23189 (comment)


## What changes are included in this PR?

<!--
There is no need to duplicate the description in the issue here but it
is sometimes worth providing a summary of the individual changes in this
PR.
-->

DataFusion does not guarantee serialized plans across versions. Keeping
`partitioned_by_file_group` therefore leaves a dead schema field and
decoder path after `output_partitioning` became the source of truth.
Reserve the old field number and name to prevent future reuse.

## Are these changes tested?

<!--
We typically require tests for all PRs in order to:
1. Prevent the code from being accidentally broken by subsequent changes
2. Serve as another way to document the expected behavior of the code

If tests are not included in your PR, please explain why (for example,
are they covered by existing tests)?
-->

Yes

## Are there any user-facing changes?

<!--
If there are user-facing changes then we may require documentation to be
updated before approving the PR.
-->

<!--
If there are any breaking changes to public APIs, please add the `api
change` label.
-->
Yes

Signed-off-by: Jiawei Zhao <Phoenix500526@163.com>
Co-authored-by: Andrew Lamb <andrew@nerdnetworks.org>
…llable build key (apache#23173)

## Which issue does this close?

- Closes apache#23126.

## Rationale for this change

`x NOT IN (subquery)` plans to a null-aware `LeftAnti` hash join (build
= outer `x`, probe = subquery). Join dynamic filter pushdown pushes a
bounds + membership filter, built from the build keys, onto the probe
scan. That filter can prune every probe row. A null-aware `LeftAnti`
reads an empty probe as a genuinely-empty subquery, so it emits
build-side NULL rows that should drop: `NULL NOT IN (non-empty)` is
UNKNOWN, not TRUE.

The result is scan-dependent, so it's a silent correctness bug. A
`VALUES` scan ignores the pushed filter and stays correct; a parquet
scan applies it and is wrong.

apache#23103 (the probe-side NULL drop) is orthogonal; this is the build-side
NULL.

## What changes are included in this PR?

Skip join dynamic filter pushdown for a null-aware anti join when the
build key can be NULL. The build-side NULL emission depends on whether
the probe is truly empty, which the pushed filter can change by emptying
it. A NOT NULL build key has no such NULL, so it keeps the pushdown.

The check is static: a schema-nullable build key disables the pushdown
even when the data contains no NULLs. A runtime alternative (keep the
pushdown and neutralize the filter only when the build actually holds a
NULL key) would restore the optimization for those cases. I'd leave that
as a follow-up.

## Are these changes tested?

Yes. A `push_down_filter_parquet.slt` case reproduces it (build-side
NULL, a non-matching parquet probe) and asserts the single correct row.
Without the change it returns the extra NULL. In addition, unit tests
pin both directions of the guard: a nullable build key rejects the
pushdown and a NOT NULL build key keeps it.

## Are there any user-facing changes?

`NOT IN` over a parquet (or otherwise prunable) scan with a nullable
outer key now returns correct results. Such joins lose the dynamic
filter pushdown.
…cimal) (apache#23631)

## Which issue does this PR close?

<!--
We generally require a GitHub issue to be filed for all bug fixes and
enhancements and this helps us generate change logs for our releases.
You can link an issue to this PR using the GitHub syntax. For example
`Closes #123` indicates that this PR will close issue #123.
-->

N/A

## Rationale for this change

<!--
Why are you proposing this change? If this is already explained clearly
in the issue then this section is not needed.
Explaining clearly why changes are proposed helps reviewers understand
your changes and offer better suggestions for fixes.
-->

I noticed there were some subtle errors with how decimals were handled
in scalar values, and also opportunity to remove power calls in favour
of precomputed constant tables. Also filling out some other missing
support.

## What changes are included in this PR?

<!--
There is no need to duplicate the description in the issue here but it
is sometimes worth providing a summary of the individual changes in this
PR.
-->

I recommend looking at the commits as they are self contained with
detailed messages for each.

## Are these changes tested?

<!--
We typically require tests for all PRs in order to:
1. Prevent the code from being accidentally broken by subsequent changes
2. Serve as another way to document the expected behavior of the code

If tests are not included in your PR, please explain why (for example,
are they covered by existing tests)?
-->

Yes

## Are there any user-facing changes?

<!--
If there are user-facing changes then we may require documentation to be
updated before approving the PR.
-->

No

<!--
If there are any breaking changes to public APIs, please add the `api
change` label.
-->
…lter pushdown enabled (apache#23638)

## Which issue does this PR close?

- Closes apache#23531.

## Rationale for this change

`input_file_name()` is `FileSource` dependent like `file_row_index()`,
and therefore shouldn't be pushed down into a filter.

## What changes are included in this PR?

`PushdownChecker` now handles both UDFs consistently.

If we keep adding this sort of metadata functions, we might want a
better API to detect them, but for now I think this is a reasonable
approach that isn't very invasive.

## Are these changes tested?

Additional SLT test that verifies that both function behave correctly
with pushdown either enabled or disabled.

## Are there any user-facing changes?

None

---------

Signed-off-by: Adam Gutglick <adamgsal@gmail.com>
…xpr::check_bigger_cast (apache#23808) (apache#23809)

## Which issue does this PR close?

- Closes apache#23808.

## Rationale for this change
 ref. apache#23807

`CastExpr::check_bigger_cast` is used to determine whether a cast is a
widening (order-preserving) conversion. Currently, it classifies `Int32
-> Float32`, `UInt32 -> Float32`, `Int64 -> Float64`, and `UInt64 ->
Float64` as widening casts.

However, integer-to-float conversions for 32-bit and 64-bit integers
lose precision when values exceed the mantissa bit limit (24 bits for
`Float32`, 53 bits for `Float64`). For example:
`16_777_216_i32 as f32 == 16_777_217_i32 as f32` (both yield
16777216.0f32).

Because distinct integer inputs can collapse to the same float output,
these casts are not strictly 1-to-1 (injective) and can break suffix
sort key ordering in multi-column ordering analysis (e.g. `[CAST(a AS
Float32), b]`).

## What changes are included in this PR?

- Updated `CastExpr::check_bigger_cast` to exclude precision-losing
integer-to-float conversions (`Int32/UInt32 -> Float32` and
`Int64/UInt64 -> Float64`).
- Added unit test `test_check_bigger_cast_precision_loss` to verify
precision-losing casts return `false` while exact conversions (`Int16 ->
Float32`, `Int32 -> Float64`, etc.) continue to return `true`.

## Are these changes tested?

Yes, new unit test `test_check_bigger_cast_precision_loss` in `cast.rs`.

## Are there any user-facing changes?

No breaking API changes. Internal optimizer behavior fix.
…ossible (apache#23702)

## Which issue does this PR close?

N/A

## Rationale for this change

`SortPreservingMergeStream` is a little complex, so add some guiding
comments and make it as textbook-like as possible

## What changes are included in this PR?

Added comments, reorder code

While this was done this also fixed couple of bugs due to how it work:
1. leftover drain was not counted in the `elapsed_compute`
2. `limit(0)` returns 0 rows and not 1 

## Are these changes tested?
Existing tests

## Are there any user-facing changes?
Not API ones.

`limit(0)` now returns 0 rows
## Which issue does this PR close?

- Closes apache#21428.

## Rationale for this change

This PR adds `BuildHasher`-based variants for `hash_utils` so callers
can compute row hashes with a caller-provided hash builder instead of
always using DataFusion's default `RandomState`.

The main constraint is performance: `with_hashes` is a hot path,
especially for string, dictionary, and nested array hashing. A previous
version in apache#21429 caused measurable regressions in the default
`RandomState` path, for example `large_utf8: single, no nulls` regressed
from roughly `26.7us` to `36.3us`, and `large_utf8: multiple, no nulls`
from roughly `112us` to `127us`.

This version keeps the default path performance-oriented by avoiding a
fully generic `BuildHasher` rewrite of the existing hot loops.

## What changes are included in this PR?

This PR adds:

- `with_hashes_with_hasher`
- `create_hashes_with_hasher`
- custom-hasher implementations for primitive, string, binary,
byte-view, dictionary, and nested arrays
- tests covering custom hashers, multi-column hashing, and dictionary
equivalence

The implementation intentionally uses a hybrid design:

- Default `RandomState` leaf hot paths remain specialized.
- Custom `BuildHasher` leaf paths live separately in
`hash_utils/build_hasher.rs`.
- Nested/structural logic is shared through an internal child-hashing
adapter, so struct/list/map/union/run/dictionary behavior does not need
to be broadly duplicated.

The trade-off is that there is still some duplication for
primitive/string/binary leaf loops. That duplication is intentional:
those are the hottest loops, and keeping them separate prevents the
existing `RandomState` path from becoming generic over `BuildHasher` or
being perturbed by the custom-hasher implementation.

## Are these changes tested?

Yes.

---------

Co-authored-by: Dmitrii Blaginin <dmitrii@blaginin.me>
Co-authored-by: Andrew Lamb <andrew@nerdnetworks.org>
…t whether they keep the same ordering of the input (apache#23807)

## Which issue does this PR close?

- Closes apache#23798

Related to:
- apache#16217

## Rationale for this change

To be able to keep the same sorting order allowing for more
optimizations

Now comet own `cast` implementation can be recognized as not modifying
sort order in the same cases that datafusion cast does.

## What changes are included in this PR?
added `strictly_order_preserving` property to `ExprProperties` + varius
other places
and replaced the hard coded logic for cast (`substitute_cast_ordering`)
about keeping input order with more generic approach that now any
expression can implement and have the same advantage of sort elimination

also marked from_unixtime as keeping ordering to show a case of this
optimization

## Are these changes tested?
yes

## Are there any user-facing changes?
yes, breaking change, added `strictly_order_preserving` property to
`ExprProperties` and to `FFI_ExprProperties`


this property means that given expression `f` and 2 values from the
input column `a` and `b` the following variants are kept:
1. `a.cmp(b) == f(a).cmp(f(b))`
2. nulls maps to nulls

Example of satisfying expression: 
`cast(col_a as BIGINT)` where `col_a` is `INT` it is keeping the
properties

Example of not satisfying:

`floor` - floor can not

`array_repeat(my_col, 2)` which might look like at first glance as
keeping the property as well but in fact it does not.

the reason is that `array_repeat(null, 2)` will output list of 2 nulls
which breaks the 2nd property that nulls must be kept as nulls


how to migrate:

Option 1 - keeping the old behavior (safest but least performant)
set `strictly_order_preserving` to false

Option 2 - Using the new optimization that this opens up:
set `strictly_order_preserving` only if the expression keep both
variants
## Which issue does this PR close?

Part of apache#22330.

## Rationale for this change

`ForeignScalarUDF` inherits the default `preserves_lex_ordering`, so
producer overrides are lost across the FFI boundary.

## What changes are included in this PR?

- Forward `preserves_lex_ordering` through `FFI_ScalarUDF`.
- Reuse the existing placement UDF for unit and dynamic-library
coverage.

## Are these changes tested?

- `cargo test -p datafusion-ffi --features integration-tests`
- `cargo clippy --all-targets --all-features -- -D warnings`

## Are there any user-facing changes?

The FFI ABI changes. Foreign libraries must rebuild against the new
DataFusion version.

---------

Signed-off-by: Amogh Ramesh <ramogh2404@gmail.com>
…tonic deques (apache#23826) (apache#23827)

## Which issue does this PR close?

- Closes apache#23826

### What changes are included in this PR?

This PR optimizes sliding window `MIN`/`MAX` aggregate functions using a
**Sequence-Numbered Monotonic Deque** instead of a Two-Stack Queue.

**Revised Design:**
- We store `(sequence_number, value)` pairs in a single `VecDeque`, kept
in strictly monotonic order.
- **`push(val)`**: Evicts dominated elements from the back, then pushes
the new value with the current `push_seq` and increments `push_seq`.
- **`pop()`**: Increments `pop_seq`. If the front elements sequence
number equals the old `pop_seq`, it is expired and popped.
- **Benefits**: No secondary FIFO queue (lower memory) and no `clone()`
overhead.

### Are these changes tested?
Yes, existing tests pass. Added tests for empty-window `pop()` and
duplicate-heavy scenarios.

### Are there any user-facing changes?
**API Change**: `MovingMin` and `MovingMax` were changed to `pub(crate)`
visibility, and their `pop()` methods now return `()` instead of
returning a value.
Performance is significantly improved (2x-3.5x throughput).

---------

Co-authored-by: Pavan <pavan@Lakshmis-MacBook-Air.local>
## Which issue does this PR close?

NA

## Rationale for this change

Adds [datapress](https://docs.datap-rs.org) to the list of known users
in the documentation.

## What changes are included in this PR?

NA

## Are these changes tested?

NA

## Are there any user-facing changes?

NA
## Which issue does this PR close?

- Closes apache#11748

## Rationale for this change

`EliminateGroupByConstant` removes GROUP BY expressions that are
constants. If all of the GROUP BY expressions are constants, eliminating
all of them results in converting a grouped aggregate into an ungrouped
(global) aggregate. This changes the semantics of the query: a grouped
aggregate query on an empty input returns zero rows, whereas an
ungrouped aggregate query produces a single row.

## What changes are included in this PR?

`EliminateGroupByConstant` now declines to eliminate constant GROUP BY
expressions, if doing so would result in removing all of the grouping
expressions.

## Are these changes tested?

Yes. Sqllogictest reproducer for the original issue, updated unit test
and optimizer SLT expectations.

## Are there any user-facing changes?

Queries with all-constant GROUP BY may now return fewer (correct) rows.
…pache#23874)

## Which issue does this PR close?

- Closes apache#23872 

## Rationale for this change

`SlidingMinAccumulator::update_batch` skipped NULL values. This meant
that if all the non-NULL values in a window frame were retracted, the
window frame would not be empty but the `MovingMin` data structure by
the `SlidingMinAccumulator` would not contain any values. This resulted
in incorrectly returning a stale non-NULL value for a sliding `min()`
over a window frame consisting of only NULL values.

Along the way, optimize the min and max sliding window accumulators to
make them both more efficient and more symmetric with one another. In
the original coding, `SlidingMaxAccumulator` included NULL values but
`SlidingMinAccumulator` omitted them, in part because omitting NULLs
made the original `retract_batch` implementation more expensive. This PR
optimizes `retract_batch`, so we can now use the same scheme for both
the min and max sliding accumulators:

* Omit NULLs on `update_batch` (this improves on the prior behavior of
`max`)
* Efficiently account for NULLs in `retract_batch` (this improves on the
prior behavior of `min`)
* Ensure correct results for all-NULL window frames (this fixes the
prior bug in `min`).

## What changes are included in this PR?

* Fix bug in `min()` over all-NULL window frames
* Optimize `SlidingMinAccumulator::retract_batch` (avoid materializing
values just to count NULLs)
* Optimize `SlidingMinAccumulator::update_batch` (omit NULLs), also
improving symmetry with `min`
* Optimize both accumulators to stop caching the current `min` / `max`;
this saves a few clones, but perhaps more importantly it is simpler and
avoids the risk of inconsistency between the cached value and the
underlying `MovingMin` / `MovingMax` data structure

## Are these changes tested?

Yes, new tests added.

## Are there any user-facing changes?

No, aside from the bug fix.
## Which issue does this PR close?

- N/A

## Rationale for this change

```
$ cargo test -p datafusion-functions-aggregate
[...]
  warning: associated function `from_parts` is never used
     --> datafusion/physical-expr/src/expressions/dynamic_filters/mod.rs:464:8
[...]
```

Fix this by appropriately gating the compilation of `from_parts`.

## What changes are included in this PR?

* Squelch unused code warning

## Are these changes tested?

Yes.

## Are there any user-facing changes?

No.
…w retract (apache#23913)

## Which issue does this PR close?

- Closes apache#23912.

## Rationale for this change

`DistinctPercentileContAccumulator` reused the shared set-based
`GenericDistinctBuffer`, which doesn't fit it: (1) the buffer asserts a
single input column but `percentile_cont` passes two (value +
percentile), so every `percentile_cont(DISTINCT ...)` panicked; (2) the
buffer is a plain `HashSet` with no multiplicity, so sliding-window
`retract_batch` dropped a value while duplicates were still in the
frame.

## What changes are included in this PR?

- Replace the shared buffer in this accumulator with a per-accumulator
`HashMap<Hashable, usize>` count map: `update_batch` reads only the
value column and increments; `retract_batch` decrements and removes a
key only at zero; `state`/`merge_batch` keep the same List state shape.
Other `GenericDistinctBuffer` users are untouched.
- Regression tests in `aggregate.slt` for the plain distinct query and
the sliding-window duplicate-retract case.

## Are these changes tested?

Yes — new regression tests; the full `aggregate.slt` suite passes.

## Are there any user-facing changes?

`percentile_cont(DISTINCT ...)` now works instead of panicking, and
returns correct results in sliding windows.
## Which issue does this PR close?

- Closes apache#23717.

## Rationale for this change

Spark and Java format negative numeric values with parentheses when the
`(` flag is present. The decimal formatting path always emitted a minus
sign because it ignored `negative_in_parentheses`, while the
floating-point path already handled the flag. This made `format_string`
inconsistent across numeric input types.

## What changes are included in this PR?

- Add a closing suffix when a negative decimal uses parentheses
formatting.
- Include the suffix when calculating width, left alignment, and zero
padding.
- Add regression coverage for grouped negative decimals with and without
an explicit width.

## Are these changes tested?

Yes. The following checks pass locally:

- `cargo fmt --all -- --check`
- `cargo clippy --all-targets --all-features -- -D warnings`
- `cargo test -p datafusion-spark --lib`
- The full workspace test command from the contributor guide with
`avro,json,backtrace,extended_tests,recursive_protection,parquet_encryption`
enabled

## Are there any user-facing changes?

Yes. Spark-compatible formatting of negative decimal values now honors
the parentheses flag. There are no public API or breaking changes.
## Which issue does this PR close?

<!--
We generally require a GitHub issue to be filed for all bug fixes and
enhancements and this helps us generate change logs for our releases.
You can link an issue to this PR using the GitHub syntax. For example
`Closes #123` indicates that this PR will close issue #123.
-->

- Closes #.

## Rationale for this change

<!--
Why are you proposing this change? If this is already explained clearly
in the issue then this section is not needed.
Explaining clearly why changes are proposed helps reviewers understand
your changes and offer better suggestions for fixes.
-->

Precedes apache#23720.

While displaying a plan with metrics, allow filtering them by name, not
just by category or type.

This is very useful when writing snapshot tests where only a specific
metrics needs to be asserted, without other unrelated metrics polluting
the snapshot assertion.

Regardless of what happens with
apache#23720, I think this PR is
still worth it, as it allows creating some really nice `insta` tests
asserting runtime properties reliably by just cherry picking the runtime
metrics relevant for that specific test. This is relevant not only
within DataFusion codebase, but also for other people's codebases using
DataFusion and relying on `insta` for snapshot testing, See an example
of this here:
https://github.com/apache/datafusion/pull/23720/changes#diff-8281405e117428077c19c07afbbe59fed303a3460096e958f761d34f0899b618

## What changes are included in this PR?

<!--
There is no need to duplicate the description in the issue here but it
is sometimes worth providing a summary of the individual changes in this
PR.
-->

While displaying metrics, allows filtering them by name

## Are these changes tested?

<!--
We typically require tests for all PRs in order to:
1. Prevent the code from being accidentally broken by subsequent changes
2. Serve as another way to document the expected behavior of the code

If tests are not included in your PR, please explain why (for example,
are they covered by existing tests)?
-->

Yes, by a new small test.

## Are there any user-facing changes?

<!--
If there are user-facing changes then we may require documentation to be
updated before approving the PR.
-->

<!--
If there are any breaking changes to public APIs, please add the `api
change` label.
-->

People using the metrics display API will be able to filter metrics by
name.
## Which issue does this PR close?

N/A

## Rationale for this change

I saw that `empty` implementation is inefficient.
I wrote review comments for more why inefficient.

there are some optimizations that do not require benchmark since they
are obvious once you understand, this is one of them

## What changes are included in this PR?
rewrote the function to be fast

## Are these changes tested?
existing tests

## Are there any user-facing changes?
no
Bumps the codeql-actions group with 2 updates:
[github/codeql-action/init](https://github.com/github/codeql-action) and
[github/codeql-action/analyze](https://github.com/github/codeql-action).

Updates `github/codeql-action/init` from 4.37.1 to 4.37.3
<details>
<summary>Release notes</summary>
<p><em>Sourced from <a
href="https://github.com/github/codeql-action/releases">github/codeql-action/init's
releases</a>.</em></p>
<blockquote>
<h2>v4.37.3</h2>
<p>No user facing changes.</p>
<h2>v4.37.2</h2>
<ul>
<li>The new address format for the <code>config-file</code> input that
was introduced in CodeQL Action 4.37.0 is now enabled by default. In
addition to the format described there, the <code>remote=</code> prefix
can now be used to explicitly indicate that the input refers to a remote
file. All previous input formats continue to be accepted as well. <a
href="https://redirect.github.com/github/codeql-action/pull/4023">#4023</a></li>
<li>The CodeQL Action can now make use of <a
href="https://docs.github.com/en/code-security/how-tos/secure-at-scale/configure-organization-security/manage-usage-and-access/giving-org-access-private-registries">configured
private registries</a> in Default Setup to retrieve CodeQL configuration
files from remote repositories that require authentication. This will
allow customers to store their CodeQL configuration in a single
repository that can then be referenced by Default Setup workflows in
other repositories. We expect to roll this and other, related changes
out to everyone in July. <a
href="https://redirect.github.com/github/codeql-action/pull/4007">#4007</a></li>
</ul>
</blockquote>
</details>
<details>
<summary>Changelog</summary>
<p><em>Sourced from <a
href="https://github.com/github/codeql-action/blob/main/CHANGELOG.md">github/codeql-action/init's
changelog</a>.</em></p>
<blockquote>
<h1>CodeQL Action Changelog</h1>
<p>See the <a
href="https://github.com/github/codeql-action/releases">releases
page</a> for the relevant changes to the CodeQL CLI and language
packs.</p>
<h2>[UNRELEASED]</h2>
<ul>
<li>This version of the CodeQL Action adds support for the
<code>tools</code> input for the <code>codeql-action/init</code> step to
be specified using a <code>github-codeql-tools</code> <a
href="https://docs.github.com/en/organizations/managing-organization-settings/managing-custom-properties-for-repositories-in-your-organization">repository
property</a>. This feature will gradually be rolled out following the
release of this version. Once rolled out, this allows for the CodeQL CLI
version that is used in GitHub-managed workflows, such as Default Setup,
to be set to a custom value. For example, customers who run into issues
with rate limits when a new CodeQL CLI version is released can set the
value to <code>toolcache</code> to always use the CodeQL CLI version
that is available in the runner toolcache. For Advanced Setup workflows,
the value provided for <code>tools</code> in the workflow definition
always takes precedence unless the value of the repository property
starts with <code>!</code>. <a
href="https://redirect.github.com/github/codeql-action/pull/4037">#4037</a></li>
</ul>
<h2>4.37.3 - 22 Jul 2026</h2>
<p>No user facing changes.</p>
<h2>4.37.2 - 21 Jul 2026</h2>
<ul>
<li>The new address format for the <code>config-file</code> input that
was introduced in CodeQL Action 4.37.0 is now enabled by default. In
addition to the format described there, the <code>remote=</code> prefix
can now be used to explicitly indicate that the input refers to a remote
file. All previous input formats continue to be accepted as well. <a
href="https://redirect.github.com/github/codeql-action/pull/4023">#4023</a></li>
<li>The CodeQL Action can now make use of <a
href="https://docs.github.com/en/code-security/how-tos/secure-at-scale/configure-organization-security/manage-usage-and-access/giving-org-access-private-registries">configured
private registries</a> in Default Setup to retrieve CodeQL configuration
files from remote repositories that require authentication. This will
allow customers to store their CodeQL configuration in a single
repository that can then be referenced by Default Setup workflows in
other repositories. We expect to roll this and other, related changes
out to everyone in July. <a
href="https://redirect.github.com/github/codeql-action/pull/4007">#4007</a></li>
</ul>
<h2>4.37.1 - 16 Jul 2026</h2>
<ul>
<li><em>Upcoming breaking change</em>: Add a deprecation warning for
customers using CodeQL version 2.20.6 and earlier. These versions of
CodeQL were discontinued on 1 July 2026 alongside GitHub Enterprise
Server 3.16, and will be unsupported by the next minor release of the
CodeQL Action. <a
href="https://redirect.github.com/github/codeql-action/pull/3956">#3956</a></li>
<li>Update default CodeQL bundle version to <a
href="https://github.com/github/codeql-action/releases/tag/codeql-bundle-v2.26.1">2.26.1</a>.
<a
href="https://redirect.github.com/github/codeql-action/pull/4019">#4019</a></li>
</ul>
<h2>4.37.0 - 08 Jul 2026</h2>
<ul>
<li>Update default CodeQL bundle version to <a
href="https://github.com/github/codeql-action/releases/tag/codeql-bundle-v2.26.0">2.26.0</a>.
<a
href="https://redirect.github.com/github/codeql-action/pull/3995">#3995</a></li>
<li>In addition to the existing input format, the
<code>config-file</code> input for the <code>codeql-action/init</code>
step will soon support a new <code>[owner/]repo[@ref][:path]</code>
format. All components except the repository name are optional. If
omitted, <code>owner</code> defaults to the same owner as the repository
the analysis is running for, <code>ref</code> to <code>main</code>, and
<code>path</code> to <code>.github/codeql-action.yaml</code>. Support
for this format ships in this version of the CodeQL Action, but will
only be enabled over the coming weeks. <a
href="https://redirect.github.com/github/codeql-action/pull/3973">#3973</a></li>
</ul>
<h2>4.36.3 - 01 Jul 2026</h2>
<p>No user facing changes.</p>
<h2>4.36.2 - 04 Jun 2026</h2>
<ul>
<li>Cache CodeQL CLI version information across Actions steps. <a
href="https://redirect.github.com/github/codeql-action/pull/3943">#3943</a></li>
<li>Reduce requests while waiting for analysis processing by using
exponential backoff when polling SARIF processing status. <a
href="https://redirect.github.com/github/codeql-action/pull/3937">#3937</a></li>
<li>Update default CodeQL bundle version to <a
href="https://github.com/github/codeql-action/releases/tag/codeql-bundle-v2.25.6">2.25.6</a>.
<a
href="https://redirect.github.com/github/codeql-action/pull/3948">#3948</a></li>
</ul>
<h2>4.36.1 - 02 Jun 2026</h2>
<p>No user facing changes.</p>
<h2>4.36.0 - 22 May 2026</h2>
<ul>
<li><em>Breaking change</em>: Bump the minimum required CodeQL bundle
version to 2.19.4. <a
href="https://redirect.github.com/github/codeql-action/pull/3894">#3894</a></li>
<li>Add support for SHA-256 Git object IDs. <a
href="https://redirect.github.com/github/codeql-action/pull/3893">#3893</a></li>
<li>Update default CodeQL bundle version to <a
href="https://github.com/github/codeql-action/releases/tag/codeql-bundle-v2.25.5">2.25.5</a>.
<a
href="https://redirect.github.com/github/codeql-action/pull/3926">#3926</a></li>
</ul>
<h2>4.35.5 - 15 May 2026</h2>
<ul>
<li>We have improved how the JavaScript bundles for the CodeQL Action
are generated to avoid duplication across bundles and reduce the size of
the repository by around 70%. This should have no effect on the runtime
behaviour of the CodeQL Action. <a
href="https://redirect.github.com/github/codeql-action/pull/3899">#3899</a></li>
</ul>
<!-- raw HTML omitted -->
</blockquote>
<p>... (truncated)</p>
</details>
<details>
<summary>Commits</summary>
<ul>
<li><a
href="https://github.com/github/codeql-action/commit/e4fba868fa4b1b91e1fdab776edc8cfbe6e9fb81"><code>e4fba86</code></a>
Merge pull request <a
href="https://redirect.github.com/github/codeql-action/issues/4031">#4031</a>
from github/update-v4.37.3-72f6a9da0</li>
<li><a
href="https://github.com/github/codeql-action/commit/fb50ab5d62a274adf3ef3e22cfe750ae87a0ede7"><code>fb50ab5</code></a>
Update changelog for v4.37.3</li>
<li><a
href="https://github.com/github/codeql-action/commit/72f6a9da0def52d9193d6a758f0378b65091f8d1"><code>72f6a9d</code></a>
Merge pull request <a
href="https://redirect.github.com/github/codeql-action/issues/4030">#4030</a>
from github/mbg/fix/no-proxy</li>
<li><a
href="https://github.com/github/codeql-action/commit/3b5ee58597653d9cc6785f3f1277f796d81f3646"><code>3b5ee58</code></a>
Use default <code>request</code> options instead of
<code>undefined</code></li>
<li><a
href="https://github.com/github/codeql-action/commit/bfb6be4b5ecd3650f02f530571453e8c64ef0778"><code>bfb6be4</code></a>
Merge pull request <a
href="https://redirect.github.com/github/codeql-action/issues/4028">#4028</a>
from github/mergeback/v4.37.2-to-main-e0647621</li>
<li><a
href="https://github.com/github/codeql-action/commit/526ab84f9858816d9cf5f7b9df4dd5e2235f0eba"><code>526ab84</code></a>
Rebuild</li>
<li><a
href="https://github.com/github/codeql-action/commit/d6217b9b8c14166e4851db94c11155d03bd13c07"><code>d6217b9</code></a>
Update changelog and version after v4.37.2</li>
<li><a
href="https://github.com/github/codeql-action/commit/e0647621c2984b5ed2f768cb892365bf2a616ad1"><code>e064762</code></a>
Merge pull request <a
href="https://redirect.github.com/github/codeql-action/issues/4027">#4027</a>
from github/update-v4.37.2-385bcdc5a</li>
<li><a
href="https://github.com/github/codeql-action/commit/e0faed839190caa67a5cd42f1cc16246028ca3df"><code>e0faed8</code></a>
Add a couple of change notes</li>
<li><a
href="https://github.com/github/codeql-action/commit/73aad0eaa9df172668665a150d17b8bc5a650c20"><code>73aad0e</code></a>
Update changelog for v4.37.2</li>
<li>Additional commits viewable in <a
href="https://github.com/github/codeql-action/compare/7188fc363630916deb702c7fdcf4e481b751f97a...e4fba868fa4b1b91e1fdab776edc8cfbe6e9fb81">compare
view</a></li>
</ul>
</details>
<br />

Updates `github/codeql-action/analyze` from 4.37.1 to 4.37.3
<details>
<summary>Release notes</summary>
<p><em>Sourced from <a
href="https://github.com/github/codeql-action/releases">github/codeql-action/analyze's
releases</a>.</em></p>
<blockquote>
<h2>v4.37.3</h2>
<p>No user facing changes.</p>
<h2>v4.37.2</h2>
<ul>
<li>The new address format for the <code>config-file</code> input that
was introduced in CodeQL Action 4.37.0 is now enabled by default. In
addition to the format described there, the <code>remote=</code> prefix
can now be used to explicitly indicate that the input refers to a remote
file. All previous input formats continue to be accepted as well. <a
href="https://redirect.github.com/github/codeql-action/pull/4023">#4023</a></li>
<li>The CodeQL Action can now make use of <a
href="https://docs.github.com/en/code-security/how-tos/secure-at-scale/configure-organization-security/manage-usage-and-access/giving-org-access-private-registries">configured
private registries</a> in Default Setup to retrieve CodeQL configuration
files from remote repositories that require authentication. This will
allow customers to store their CodeQL configuration in a single
repository that can then be referenced by Default Setup workflows in
other repositories. We expect to roll this and other, related changes
out to everyone in July. <a
href="https://redirect.github.com/github/codeql-action/pull/4007">#4007</a></li>
</ul>
</blockquote>
</details>
<details>
<summary>Changelog</summary>
<p><em>Sourced from <a
href="https://github.com/github/codeql-action/blob/main/CHANGELOG.md">github/codeql-action/analyze's
changelog</a>.</em></p>
<blockquote>
<h1>CodeQL Action Changelog</h1>
<p>See the <a
href="https://github.com/github/codeql-action/releases">releases
page</a> for the relevant changes to the CodeQL CLI and language
packs.</p>
<h2>[UNRELEASED]</h2>
<ul>
<li>This version of the CodeQL Action adds support for the
<code>tools</code> input for the <code>codeql-action/init</code> step to
be specified using a <code>github-codeql-tools</code> <a
href="https://docs.github.com/en/organizations/managing-organization-settings/managing-custom-properties-for-repositories-in-your-organization">repository
property</a>. This feature will gradually be rolled out following the
release of this version. Once rolled out, this allows for the CodeQL CLI
version that is used in GitHub-managed workflows, such as Default Setup,
to be set to a custom value. For example, customers who run into issues
with rate limits when a new CodeQL CLI version is released can set the
value to <code>toolcache</code> to always use the CodeQL CLI version
that is available in the runner toolcache. For Advanced Setup workflows,
the value provided for <code>tools</code> in the workflow definition
always takes precedence unless the value of the repository property
starts with <code>!</code>. <a
href="https://redirect.github.com/github/codeql-action/pull/4037">#4037</a></li>
</ul>
<h2>4.37.3 - 22 Jul 2026</h2>
<p>No user facing changes.</p>
<h2>4.37.2 - 21 Jul 2026</h2>
<ul>
<li>The new address format for the <code>config-file</code> input that
was introduced in CodeQL Action 4.37.0 is now enabled by default. In
addition to the format described there, the <code>remote=</code> prefix
can now be used to explicitly indicate that the input refers to a remote
file. All previous input formats continue to be accepted as well. <a
href="https://redirect.github.com/github/codeql-action/pull/4023">#4023</a></li>
<li>The CodeQL Action can now make use of <a
href="https://docs.github.com/en/code-security/how-tos/secure-at-scale/configure-organization-security/manage-usage-and-access/giving-org-access-private-registries">configured
private registries</a> in Default Setup to retrieve CodeQL configuration
files from remote repositories that require authentication. This will
allow customers to store their CodeQL configuration in a single
repository that can then be referenced by Default Setup workflows in
other repositories. We expect to roll this and other, related changes
out to everyone in July. <a
href="https://redirect.github.com/github/codeql-action/pull/4007">#4007</a></li>
</ul>
<h2>4.37.1 - 16 Jul 2026</h2>
<ul>
<li><em>Upcoming breaking change</em>: Add a deprecation warning for
customers using CodeQL version 2.20.6 and earlier. These versions of
CodeQL were discontinued on 1 July 2026 alongside GitHub Enterprise
Server 3.16, and will be unsupported by the next minor release of the
CodeQL Action. <a
href="https://redirect.github.com/github/codeql-action/pull/3956">#3956</a></li>
<li>Update default CodeQL bundle version to <a
href="https://github.com/github/codeql-action/releases/tag/codeql-bundle-v2.26.1">2.26.1</a>.
<a
href="https://redirect.github.com/github/codeql-action/pull/4019">#4019</a></li>
</ul>
<h2>4.37.0 - 08 Jul 2026</h2>
<ul>
<li>Update default CodeQL bundle version to <a
href="https://github.com/github/codeql-action/releases/tag/codeql-bundle-v2.26.0">2.26.0</a>.
<a
href="https://redirect.github.com/github/codeql-action/pull/3995">#3995</a></li>
<li>In addition to the existing input format, the
<code>config-file</code> input for the <code>codeql-action/init</code>
step will soon support a new <code>[owner/]repo[@ref][:path]</code>
format. All components except the repository name are optional. If
omitted, <code>owner</code> defaults to the same owner as the repository
the analysis is running for, <code>ref</code> to <code>main</code>, and
<code>path</code> to <code>.github/codeql-action.yaml</code>. Support
for this format ships in this version of the CodeQL Action, but will
only be enabled over the coming weeks. <a
href="https://redirect.github.com/github/codeql-action/pull/3973">#3973</a></li>
</ul>
<h2>4.36.3 - 01 Jul 2026</h2>
<p>No user facing changes.</p>
<h2>4.36.2 - 04 Jun 2026</h2>
<ul>
<li>Cache CodeQL CLI version information across Actions steps. <a
href="https://redirect.github.com/github/codeql-action/pull/3943">#3943</a></li>
<li>Reduce requests while waiting for analysis processing by using
exponential backoff when polling SARIF processing status. <a
href="https://redirect.github.com/github/codeql-action/pull/3937">#3937</a></li>
<li>Update default CodeQL bundle version to <a
href="https://github.com/github/codeql-action/releases/tag/codeql-bundle-v2.25.6">2.25.6</a>.
<a
href="https://redirect.github.com/github/codeql-action/pull/3948">#3948</a></li>
</ul>
<h2>4.36.1 - 02 Jun 2026</h2>
<p>No user facing changes.</p>
<h2>4.36.0 - 22 May 2026</h2>
<ul>
<li><em>Breaking change</em>: Bump the minimum required CodeQL bundle
version to 2.19.4. <a
href="https://redirect.github.com/github/codeql-action/pull/3894">#3894</a></li>
<li>Add support for SHA-256 Git object IDs. <a
href="https://redirect.github.com/github/codeql-action/pull/3893">#3893</a></li>
<li>Update default CodeQL bundle version to <a
href="https://github.com/github/codeql-action/releases/tag/codeql-bundle-v2.25.5">2.25.5</a>.
<a
href="https://redirect.github.com/github/codeql-action/pull/3926">#3926</a></li>
</ul>
<h2>4.35.5 - 15 May 2026</h2>
<ul>
<li>We have improved how the JavaScript bundles for the CodeQL Action
are generated to avoid duplication across bundles and reduce the size of
the repository by around 70%. This should have no effect on the runtime
behaviour of the CodeQL Action. <a
href="https://redirect.github.com/github/codeql-action/pull/3899">#3899</a></li>
</ul>
<!-- raw HTML omitted -->
</blockquote>
<p>... (truncated)</p>
</details>
<details>
<summary>Commits</summary>
<ul>
<li><a
href="https://github.com/github/codeql-action/commit/e4fba868fa4b1b91e1fdab776edc8cfbe6e9fb81"><code>e4fba86</code></a>
Merge pull request <a
href="https://redirect.github.com/github/codeql-action/issues/4031">#4031</a>
from github/update-v4.37.3-72f6a9da0</li>
<li><a
href="https://github.com/github/codeql-action/commit/fb50ab5d62a274adf3ef3e22cfe750ae87a0ede7"><code>fb50ab5</code></a>
Update changelog for v4.37.3</li>
<li><a
href="https://github.com/github/codeql-action/commit/72f6a9da0def52d9193d6a758f0378b65091f8d1"><code>72f6a9d</code></a>
Merge pull request <a
href="https://redirect.github.com/github/codeql-action/issues/4030">#4030</a>
from github/mbg/fix/no-proxy</li>
<li><a
href="https://github.com/github/codeql-action/commit/3b5ee58597653d9cc6785f3f1277f796d81f3646"><code>3b5ee58</code></a>
Use default <code>request</code> options instead of
<code>undefined</code></li>
<li><a
href="https://github.com/github/codeql-action/commit/bfb6be4b5ecd3650f02f530571453e8c64ef0778"><code>bfb6be4</code></a>
Merge pull request <a
href="https://redirect.github.com/github/codeql-action/issues/4028">#4028</a>
from github/mergeback/v4.37.2-to-main-e0647621</li>
<li><a
href="https://github.com/github/codeql-action/commit/526ab84f9858816d9cf5f7b9df4dd5e2235f0eba"><code>526ab84</code></a>
Rebuild</li>
<li><a
href="https://github.com/github/codeql-action/commit/d6217b9b8c14166e4851db94c11155d03bd13c07"><code>d6217b9</code></a>
Update changelog and version after v4.37.2</li>
<li><a
href="https://github.com/github/codeql-action/commit/e0647621c2984b5ed2f768cb892365bf2a616ad1"><code>e064762</code></a>
Merge pull request <a
href="https://redirect.github.com/github/codeql-action/issues/4027">#4027</a>
from github/update-v4.37.2-385bcdc5a</li>
<li><a
href="https://github.com/github/codeql-action/commit/e0faed839190caa67a5cd42f1cc16246028ca3df"><code>e0faed8</code></a>
Add a couple of change notes</li>
<li><a
href="https://github.com/github/codeql-action/commit/73aad0eaa9df172668665a150d17b8bc5a650c20"><code>73aad0e</code></a>
Update changelog for v4.37.2</li>
<li>Additional commits viewable in <a
href="https://github.com/github/codeql-action/compare/7188fc363630916deb702c7fdcf4e481b751f97a...e4fba868fa4b1b91e1fdab776edc8cfbe6e9fb81">compare
view</a></li>
</ul>
</details>
<br />


Dependabot will resolve any conflicts with this PR as long as you don't
alter it yourself. You can also trigger a rebase manually by commenting
`@dependabot rebase`.

[//]: # (dependabot-automerge-start)
[//]: # (dependabot-automerge-end)

---

<details>
<summary>Dependabot commands and options</summary>
<br />

You can trigger Dependabot actions by commenting on this PR:
- `@dependabot rebase` will rebase this PR
- `@dependabot recreate` will recreate this PR, overwriting any edits
that have been made to it
- `@dependabot show <dependency name> ignore conditions` will show all
of the ignore conditions of the specified dependency
- `@dependabot ignore <dependency name> major version` will close this
group update PR and stop Dependabot creating any more for the specific
dependency's major version (unless you unignore this specific
dependency's major version or upgrade to it yourself)
- `@dependabot ignore <dependency name> minor version` will close this
group update PR and stop Dependabot creating any more for the specific
dependency's minor version (unless you unignore this specific
dependency's minor version or upgrade to it yourself)
- `@dependabot ignore <dependency name>` will close this group update PR
and stop Dependabot creating any more for the specific dependency
(unless you unignore this specific dependency or upgrade to it yourself)
- `@dependabot unignore <dependency name>` will remove all of the ignore
conditions of the specified dependency
- `@dependabot unignore <dependency name> <ignore condition>` will
remove the ignore condition of the specified dependency and ignore
conditions


</details>

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
…e#23941)

Bumps
[taiki-e/install-action](https://github.com/taiki-e/install-action) from
2.84.0 to 2.85.2.
<details>
<summary>Release notes</summary>
<p><em>Sourced from <a
href="https://github.com/taiki-e/install-action/releases">taiki-e/install-action's
releases</a>.</em></p>
<blockquote>
<h2>2.85.2</h2>
<ul>
<li>
<p>Update <code>prek@latest</code> to 0.4.11.</p>
</li>
<li>
<p>Update <code>mise@latest</code> to 2026.7.13.</p>
</li>
<li>
<p>Update <code>kingfisher@latest</code> to 1.109.0.</p>
</li>
</ul>
<h2>2.85.1</h2>
<ul>
<li>
<p>Update <code>vacuum@latest</code> to 0.30.0.</p>
</li>
<li>
<p>Update <code>uv@latest</code> to 0.11.32.</p>
</li>
<li>
<p>Update <code>mise@latest</code> to 2026.7.12.</p>
</li>
<li>
<p>Update <code>cyclonedx@latest</code> to 0.33.1.</p>
</li>
<li>
<p>Update <code>cargo-neat@latest</code> to 0.5.2.</p>
</li>
</ul>
<h2>2.85.0</h2>
<ul>
<li>
<p>Support <code>wild</code> (alias: <code>wild-linker</code>). (<a
href="https://redirect.github.com/taiki-e/install-action/pull/1949">#1949</a>)</p>
</li>
<li>
<p>Support <code>bpf-linker</code>. (<a
href="https://redirect.github.com/taiki-e/install-action/pull/1950">#1950</a>)</p>
</li>
<li>
<p>Support <code>rafn</code>. (<a
href="https://redirect.github.com/taiki-e/install-action/pull/1935">#1935</a>,
thanks <a
href="https://github.com/DarkWanderer"><code>@​DarkWanderer</code></a>)</p>
</li>
<li>
<p>Update <code>cargo-neat@latest</code> to 0.5.1.</p>
</li>
<li>
<p>Update <code>zizmor@latest</code> to 1.28.0.</p>
</li>
<li>
<p>Update <code>wasmtime@latest</code> to 47.0.2.</p>
</li>
<li>
<p>Update <code>uv@latest</code> to 0.11.31.</p>
</li>
<li>
<p>Update <code>syft@latest</code> to 1.49.0.</p>
</li>
</ul>
<h2>2.84.1</h2>
<ul>
<li>
<p>Update <code>wasmtime@latest</code> to 47.0.1.</p>
</li>
<li>
<p>Update <code>wasm-tools@latest</code> to 1.254.0.</p>
</li>
<li>
<p>Update <code>uv@latest</code> to 0.11.30.</p>
</li>
<li>
<p>Update <code>mise@latest</code> to 2026.7.11.</p>
</li>
<li>
<p>Update <code>cargo-neat@latest</code> to 0.5.0.</p>
</li>
<li>
<p>Update <code>cargo-crap@latest</code> to 0.3.1.</p>
</li>
<li>
<p>Update <code>biome@latest</code> to 2.5.5.</p>
</li>
</ul>
</blockquote>
</details>
<details>
<summary>Changelog</summary>
<p><em>Sourced from <a
href="https://github.com/taiki-e/install-action/blob/main/CHANGELOG.md">taiki-e/install-action's
changelog</a>.</em></p>
<blockquote>
<h1>Changelog</h1>
<p>All notable changes to this project will be documented in this
file.</p>
<p>This project adheres to <a href="https://semver.org">Semantic
Versioning</a>.</p>
<!-- raw HTML omitted -->
<h2>[Unreleased]</h2>
<h2>[2.85.2] - 2026-07-26</h2>
<ul>
<li>
<p>Update <code>prek@latest</code> to 0.4.11.</p>
</li>
<li>
<p>Update <code>mise@latest</code> to 2026.7.13.</p>
</li>
<li>
<p>Update <code>kingfisher@latest</code> to 1.109.0.</p>
</li>
</ul>
<h2>[2.85.1] - 2026-07-25</h2>
<ul>
<li>
<p>Update <code>vacuum@latest</code> to 0.30.0.</p>
</li>
<li>
<p>Update <code>uv@latest</code> to 0.11.32.</p>
</li>
<li>
<p>Update <code>mise@latest</code> to 2026.7.12.</p>
</li>
<li>
<p>Update <code>cyclonedx@latest</code> to 0.33.1.</p>
</li>
<li>
<p>Update <code>cargo-neat@latest</code> to 0.5.2.</p>
</li>
</ul>
<h2>[2.85.0] - 2026-07-23</h2>
<ul>
<li>
<p>Support <code>wild</code> (alias: <code>wild-linker</code>). (<a
href="https://redirect.github.com/taiki-e/install-action/pull/1949">#1949</a>)</p>
</li>
<li>
<p>Support <code>bpf-linker</code>. (<a
href="https://redirect.github.com/taiki-e/install-action/pull/1950">#1950</a>)</p>
</li>
<li>
<p>Support <code>rafn</code>. (<a
href="https://redirect.github.com/taiki-e/install-action/pull/1935">#1935</a>,
thanks <a
href="https://github.com/DarkWanderer"><code>@​DarkWanderer</code></a>)</p>
</li>
<li>
<p>Update <code>cargo-neat@latest</code> to 0.5.1.</p>
</li>
<li>
<p>Update <code>zizmor@latest</code> to 1.28.0.</p>
</li>
<li>
<p>Update <code>wasmtime@latest</code> to 47.0.2.</p>
</li>
<li>
<p>Update <code>uv@latest</code> to 0.11.31.</p>
</li>
<li>
<p>Update <code>syft@latest</code> to 1.49.0.</p>
</li>
</ul>
<!-- raw HTML omitted -->
</blockquote>
<p>... (truncated)</p>
</details>
<details>
<summary>Commits</summary>
<ul>
<li><a
href="https://github.com/taiki-e/install-action/commit/41049aa56687c35e0afa74eed4f09cec4f9afabf"><code>41049aa</code></a>
Release 2.85.2</li>
<li><a
href="https://github.com/taiki-e/install-action/commit/dfcf36552b1b743e910b9b9b6b114a08e56a1e5e"><code>dfcf365</code></a>
Update <code>prek@latest</code> to 0.4.11</li>
<li><a
href="https://github.com/taiki-e/install-action/commit/eea03ccfa855b965a301aa52424915fddc32f855"><code>eea03cc</code></a>
Update <code>mise@latest</code> to 2026.7.13</li>
<li><a
href="https://github.com/taiki-e/install-action/commit/81ca2feb84748ed3fa514707ee1004f4f3ad715a"><code>81ca2fe</code></a>
Update martin manifest</li>
<li><a
href="https://github.com/taiki-e/install-action/commit/cc90ed04bc1a258ea90a83789f24b759a77c9671"><code>cc90ed0</code></a>
Update <code>kingfisher@latest</code> to 1.109.0</li>
<li><a
href="https://github.com/taiki-e/install-action/commit/55639a3362f508fda5154d3451072205f5c32ca6"><code>55639a3</code></a>
Update cargo-shear manifest</li>
<li><a
href="https://github.com/taiki-e/install-action/commit/3d7d7cd5ac7f994c1892ae0c06165095b9139094"><code>3d7d7cd</code></a>
Release 2.85.1</li>
<li><a
href="https://github.com/taiki-e/install-action/commit/d09ccb4fe2105ac9ead6376f7a423ff88defeb57"><code>d09ccb4</code></a>
Update <code>vacuum@latest</code> to 0.30.0</li>
<li><a
href="https://github.com/taiki-e/install-action/commit/ac43dee1a92e482c0e244bdb0bb126908159e245"><code>ac43dee</code></a>
Update <code>uv@latest</code> to 0.11.32</li>
<li><a
href="https://github.com/taiki-e/install-action/commit/49b16979f38f30e44b1f68af9115be9a6ffc5215"><code>49b1697</code></a>
Update prek manifest</li>
<li>Additional commits viewable in <a
href="https://github.com/taiki-e/install-action/compare/a6b2e2dcd845ddd7f509ce4f3ed3d922b80cc5d9...41049aa56687c35e0afa74eed4f09cec4f9afabf">compare
view</a></li>
</ul>
</details>
<br />


[![Dependabot compatibility
score](https://dependabot-badges.githubapp.com/badges/compatibility_score?dependency-name=taiki-e/install-action&package-manager=github_actions&previous-version=2.84.0&new-version=2.85.2)](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores)

Dependabot will resolve any conflicts with this PR as long as you don't
alter it yourself. You can also trigger a rebase manually by commenting
`@dependabot rebase`.

[//]: # (dependabot-automerge-start)
[//]: # (dependabot-automerge-end)

---

<details>
<summary>Dependabot commands and options</summary>
<br />

You can trigger Dependabot actions by commenting on this PR:
- `@dependabot rebase` will rebase this PR
- `@dependabot recreate` will recreate this PR, overwriting any edits
that have been made to it
- `@dependabot show <dependency name> ignore conditions` will show all
of the ignore conditions of the specified dependency
- `@dependabot ignore this major version` will close this PR and stop
Dependabot creating any more for this major version (unless you reopen
the PR or upgrade to it yourself)
- `@dependabot ignore this minor version` will close this PR and stop
Dependabot creating any more for this minor version (unless you reopen
the PR or upgrade to it yourself)
- `@dependabot ignore this dependency` will close this PR and stop
Dependabot creating any more for this dependency (unless you reopen the
PR or upgrade to it yourself)


</details>

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Bumps [actions/stale](https://github.com/actions/stale) from 10.4.0 to
11.0.0.
<details>
<summary>Release notes</summary>
<p><em>Sourced from <a
href="https://github.com/actions/stale/releases">actions/stale's
releases</a>.</em></p>
<blockquote>
<h2>v11.0.0</h2>
<h2>What's Changed</h2>
<h3>Enhancement</h3>
<ul>
<li>Migrate to ESM and update dependencies by <a
href="https://github-grid.enterprise.slack.com/team/U08CVLQ4JKE"><code>@​chiranjib-swain</code></a>
in <a
href="https://redirect.github.com/actions/stale/pull/1350">actions/stale#1350</a></li>
</ul>
<h3>Dependency Update</h3>
<ul>
<li>Override brace-expansion to 5.0.8 to address 24 high-severity
dependency vulnerabilities by <a
href="https://github.com/dependabot"><code>@​dependabot</code></a> in <a
href="https://redirect.github.com/actions/stale/pull/1351">actions/stale#1351</a></li>
</ul>
<p><strong>Full Changelog</strong>: <a
href="https://github.com/actions/stale/compare/v10...v11.0.0">https://github.com/actions/stale/compare/v10...v11.0.0</a></p>
</blockquote>
</details>
<details>
<summary>Commits</summary>
<ul>
<li><a
href="https://github.com/actions/stale/commit/4391f3da665fdf50b6810c1a66712fb9ba21aa93"><code>4391f3d</code></a>
Fix 24 high severity vulnerabilities by overriding brace-expansion to
5.0.8 (...</li>
<li><a
href="https://github.com/actions/stale/commit/eaf9131fae5eafd0c31a64ebe3a2e183266fec48"><code>eaf9131</code></a>
refactor: update imports to use ES module syntax and improve test
structure (...</li>
<li>See full diff in <a
href="https://github.com/actions/stale/compare/1e223db275d687790206a7acac4d1a11bd6fe629...4391f3da665fdf50b6810c1a66712fb9ba21aa93">compare
view</a></li>
</ul>
</details>
<br />


[![Dependabot compatibility
score](https://dependabot-badges.githubapp.com/badges/compatibility_score?dependency-name=actions/stale&package-manager=github_actions&previous-version=10.4.0&new-version=11.0.0)](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores)

Dependabot will resolve any conflicts with this PR as long as you don't
alter it yourself. You can also trigger a rebase manually by commenting
`@dependabot rebase`.

[//]: # (dependabot-automerge-start)
[//]: # (dependabot-automerge-end)

---

<details>
<summary>Dependabot commands and options</summary>
<br />

You can trigger Dependabot actions by commenting on this PR:
- `@dependabot rebase` will rebase this PR
- `@dependabot recreate` will recreate this PR, overwriting any edits
that have been made to it
- `@dependabot show <dependency name> ignore conditions` will show all
of the ignore conditions of the specified dependency
- `@dependabot ignore this major version` will close this PR and stop
Dependabot creating any more for this major version (unless you reopen
the PR or upgrade to it yourself)
- `@dependabot ignore this minor version` will close this PR and stop
Dependabot creating any more for this minor version (unless you reopen
the PR or upgrade to it yourself)
- `@dependabot ignore this dependency` will close this PR and stop
Dependabot creating any more for this dependency (unless you reopen the
PR or upgrade to it yourself)


</details>

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Bumps [base64](https://github.com/marshallpierce/rust-base64) from
0.22.1 to 0.23.0.
<details>
<summary>Changelog</summary>
<p><em>Sourced from <a
href="https://github.com/marshallpierce/rust-base64/blob/master/RELEASE-NOTES.md">base64's
changelog</a>.</em></p>
<blockquote>
<h1>0.23.0</h1>
<ul>
<li>Added more consts for preconfigured configs and engines</li>
<li>Make DecodeError::InvalidLastSymbol more clear by including the
decoded value</li>
<li>Added SIMD-accelerated engines behind the default-on
<code>simd-unsafe</code> feature: <code>Simd</code> picks the best
instruction set at runtime (AVX2 on <code>x86_64</code>, NEON on
<code>aarch64</code>) and falls back to the scalar
<code>GeneralPurpose</code> engine, while <code>Avx2</code> and
<code>Neon</code> target one instruction set with no runtime
detection and work in <code>no_std</code>. The engines support the
standard and URL-safe alphabets.</li>
<li>Update MSRV to 1.71.0</li>
<li>Add support for custom padding symbols</li>
</ul>
</blockquote>
</details>
<details>
<summary>Commits</summary>
<ul>
<li><a
href="https://github.com/marshallpierce/rust-base64/commit/9e9220a4166f628de7c8803289e120ae1e944f78"><code>9e9220a</code></a>
v0.23.0</li>
<li><a
href="https://github.com/marshallpierce/rust-base64/commit/870326ec592eebde9d6bfe4c5d8130c591273e9c"><code>870326e</code></a>
Merge pull request <a
href="https://redirect.github.com/marshallpierce/rust-base64/issues/306">#306</a>
from marshallpierce/mp/trailing-bits-docs</li>
<li><a
href="https://github.com/marshallpierce/rust-base64/commit/fbec5f1050f9fc16e6a826ebabaa2b7b0644bd67"><code>fbec5f1</code></a>
Document no trailing trailing bits</li>
<li><a
href="https://github.com/marshallpierce/rust-base64/commit/0a23549968f059b53cf39e96eba8f46779f322a7"><code>0a23549</code></a>
Merge pull request <a
href="https://redirect.github.com/marshallpierce/rust-base64/issues/305">#305</a>
from marshallpierce/mp/edition-2021</li>
<li><a
href="https://github.com/marshallpierce/rust-base64/commit/f10b7e20614135aa61289140683fc93e5a45d338"><code>f10b7e2</code></a>
Update deps &amp; edition</li>
<li><a
href="https://github.com/marshallpierce/rust-base64/commit/9d21a598860645cb6290940e7a43033bc43ebd74"><code>9d21a59</code></a>
Merge pull request <a
href="https://redirect.github.com/marshallpierce/rust-base64/issues/304">#304</a>
from marshallpierce/mp/custom-padding-rebase</li>
<li><a
href="https://github.com/marshallpierce/rust-base64/commit/f70bad2caaa85350b95d988bfb9a0997e824bfd8"><code>f70bad2</code></a>
Support custom padding symbols</li>
<li><a
href="https://github.com/marshallpierce/rust-base64/commit/684d79cd3deb8dfd5323619634c75bc0ff6edfd9"><code>684d79c</code></a>
Merge pull request <a
href="https://redirect.github.com/marshallpierce/rust-base64/issues/301">#301</a>
from marshallpierce/mp/simd-gardening</li>
<li><a
href="https://github.com/marshallpierce/rust-base64/commit/5bf66f2646c6fd99e1b18bf4e2bb0a47d34e1eaa"><code>5bf66f2</code></a>
Merge pull request <a
href="https://redirect.github.com/marshallpierce/rust-base64/issues/284">#284</a>
from AbeZbm/add-tests</li>
<li><a
href="https://github.com/marshallpierce/rust-base64/commit/d3831cfbf7dafe226a8383450c3410e2f67c826a"><code>d3831cf</code></a>
Followups to SIMD work</li>
<li>Additional commits viewable in <a
href="https://github.com/marshallpierce/rust-base64/compare/v0.22.1...v0.23.0">compare
view</a></li>
</ul>
</details>
<br />


[![Dependabot compatibility
score](https://dependabot-badges.githubapp.com/badges/compatibility_score?dependency-name=base64&package-manager=cargo&previous-version=0.22.1&new-version=0.23.0)](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores)

Dependabot will resolve any conflicts with this PR as long as you don't
alter it yourself. You can also trigger a rebase manually by commenting
`@dependabot rebase`.

[//]: # (dependabot-automerge-start)
[//]: # (dependabot-automerge-end)

---

<details>
<summary>Dependabot commands and options</summary>
<br />

You can trigger Dependabot actions by commenting on this PR:
- `@dependabot rebase` will rebase this PR
- `@dependabot recreate` will recreate this PR, overwriting any edits
that have been made to it
- `@dependabot show <dependency name> ignore conditions` will show all
of the ignore conditions of the specified dependency
- `@dependabot ignore this major version` will close this PR and stop
Dependabot creating any more for this major version (unless you reopen
the PR or upgrade to it yourself)
- `@dependabot ignore this minor version` will close this PR and stop
Dependabot creating any more for this minor version (unless you reopen
the PR or upgrade to it yourself)
- `@dependabot ignore this dependency` will close this PR and stop
Dependabot creating any more for this dependency (unless you reopen the
PR or upgrade to it yourself)


</details>

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Bumps [astral-sh/setup-uv](https://github.com/astral-sh/setup-uv) from
8.3.2 to 9.0.0.
<details>
<summary>Release notes</summary>
<p><em>Sourced from <a
href="https://github.com/astral-sh/setup-uv/releases">astral-sh/setup-uv's
releases</a>.</em></p>
<blockquote>
<h2>v9.0.0 🌈 Change <code>prune-cache</code> default to
<code>false</code></h2>
<h2>Changes</h2>
<p>This release disables the default cache cache pruning to ease the
load on the PyPi infrastructure.
Since users might experience more GitHub Actions cache usage which might
result in higher costs this is marked as a breaking change. To read more
on why we did this (now) you can read the detailed analysis and
reasoning in <a
href="https://redirect.github.com/astral-sh/setup-uv/issues/967">#967</a></p>
<p>Besides this big breaking change we also have a small bugfix while
building caches for linux distributions that behave a big different than
the &quot;big ones&quot; and a speed up in version resolution by only
reading the version manifest until a matching version is found saving
runtime and network bandwith.</p>
<h2>🚨 Breaking changes</h2>
<ul>
<li>Change <code>prune-cache</code> default to <code>false</code> <a
href="https://github.com/charliermarsh"><code>@​charliermarsh</code></a>
(<a
href="https://redirect.github.com/astral-sh/setup-uv/issues/967">#967</a>)</li>
</ul>
<h2>🐛 Bug fixes</h2>
<ul>
<li>fix: fall back to distribution ID when os-release has no version
field <a href="https://github.com/cxzhong"><code>@​cxzhong</code></a>
(<a
href="https://redirect.github.com/astral-sh/setup-uv/issues/961">#961</a>)</li>
</ul>
<h2>🚀 Enhancements</h2>
<ul>
<li>Speed up version client by partial response reads <a
href="https://github.com/eifinger"><code>@​eifinger</code></a> (<a
href="https://redirect.github.com/astral-sh/setup-uv/issues/807">#807</a>)</li>
</ul>
<h2>🧰 Maintenance</h2>
<ul>
<li>chore: update known checksums for 0.11.30 @<a
href="https://github.com/apps/github-actions">github-actions[bot]</a>
(<a
href="https://redirect.github.com/astral-sh/setup-uv/issues/968">#968</a>)</li>
<li>chore: update known checksums for 0.11.29 @<a
href="https://github.com/apps/github-actions">github-actions[bot]</a>
(<a
href="https://redirect.github.com/astral-sh/setup-uv/issues/960">#960</a>)</li>
</ul>
<h2>📚 Documentation</h2>
<ul>
<li>docs: update version references to v8.3.2 @<a
href="https://github.com/apps/github-actions">github-actions[bot]</a>
(<a
href="https://redirect.github.com/astral-sh/setup-uv/issues/949">#949</a>)</li>
</ul>
<h2>⬆️ Dependency updates</h2>
<ul>
<li>chore(deps): roll up Dependabot updates <a
href="https://github.com/eifinger"><code>@​eifinger</code></a> (<a
href="https://redirect.github.com/astral-sh/setup-uv/issues/970">#970</a>)</li>
<li>chore(deps): roll up Dependabot updates <a
href="https://github.com/eifinger"><code>@​eifinger</code></a> (<a
href="https://redirect.github.com/astral-sh/setup-uv/issues/962">#962</a>)</li>
</ul>
</blockquote>
</details>
<details>
<summary>Commits</summary>
<ul>
<li><a
href="https://github.com/astral-sh/setup-uv/commit/c771a70e6277c0a99b617c7a806ffedaca235ff9"><code>c771a70</code></a>
chore(deps): roll up Dependabot updates (<a
href="https://redirect.github.com/astral-sh/setup-uv/issues/970">#970</a>)</li>
<li><a
href="https://github.com/astral-sh/setup-uv/commit/2f537ca87c1ffa233ca2a1b84815388e3e42d845"><code>2f537ca</code></a>
chore: update known checksums for 0.11.30 (<a
href="https://redirect.github.com/astral-sh/setup-uv/issues/968">#968</a>)</li>
<li><a
href="https://github.com/astral-sh/setup-uv/commit/2269552d547df6f50e57442326930d30d943afe3"><code>2269552</code></a>
Speed up version client by partial response reads (<a
href="https://redirect.github.com/astral-sh/setup-uv/issues/807">#807</a>)</li>
<li><a
href="https://github.com/astral-sh/setup-uv/commit/47a7f4fb2e900d6c33a5b5f231fa21dbfaeba52f"><code>47a7f4f</code></a>
Change <code>prune-cache</code> default to <code>false</code> (<a
href="https://redirect.github.com/astral-sh/setup-uv/issues/967">#967</a>)</li>
<li><a
href="https://github.com/astral-sh/setup-uv/commit/71966eff34a27b0a62ed4b9f6f6e383e071b1bb5"><code>71966ef</code></a>
chore(deps): roll up Dependabot updates (<a
href="https://redirect.github.com/astral-sh/setup-uv/issues/962">#962</a>)</li>
<li><a
href="https://github.com/astral-sh/setup-uv/commit/f12b1f0a84bd6dc2331b36b2bbdbb1d1e617dbcc"><code>f12b1f0</code></a>
fix: fall back to distribution ID when os-release has no version field
(<a
href="https://redirect.github.com/astral-sh/setup-uv/issues/961">#961</a>)</li>
<li><a
href="https://github.com/astral-sh/setup-uv/commit/ecd24dd710f2fb0dca1693a67af11fc4a5c5ec84"><code>ecd24dd</code></a>
chore: update known checksums for 0.11.29 (<a
href="https://redirect.github.com/astral-sh/setup-uv/issues/960">#960</a>)</li>
<li><a
href="https://github.com/astral-sh/setup-uv/commit/6a191366842ac1502ba6c07e9b5acd5c2d9d8db3"><code>6a19136</code></a>
docs: update version references to v8.3.2 (<a
href="https://redirect.github.com/astral-sh/setup-uv/issues/949">#949</a>)</li>
<li>See full diff in <a
href="https://github.com/astral-sh/setup-uv/compare/11f9893b081a58869d3b5fccaea48c9e9e46f990...c771a70e6277c0a99b617c7a806ffedaca235ff9">compare
view</a></li>
</ul>
</details>
<br />


[![Dependabot compatibility
score](https://dependabot-badges.githubapp.com/badges/compatibility_score?dependency-name=astral-sh/setup-uv&package-manager=github_actions&previous-version=8.3.2&new-version=9.0.0)](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores)

Dependabot will resolve any conflicts with this PR as long as you don't
alter it yourself. You can also trigger a rebase manually by commenting
`@dependabot rebase`.

[//]: # (dependabot-automerge-start)
[//]: # (dependabot-automerge-end)

---

<details>
<summary>Dependabot commands and options</summary>
<br />

You can trigger Dependabot actions by commenting on this PR:
- `@dependabot rebase` will rebase this PR
- `@dependabot recreate` will recreate this PR, overwriting any edits
that have been made to it
- `@dependabot show <dependency name> ignore conditions` will show all
of the ignore conditions of the specified dependency
- `@dependabot ignore this major version` will close this PR and stop
Dependabot creating any more for this major version (unless you reopen
the PR or upgrade to it yourself)
- `@dependabot ignore this minor version` will close this PR and stop
Dependabot creating any more for this minor version (unless you reopen
the PR or upgrade to it yourself)
- `@dependabot ignore this dependency` will close this PR and stop
Dependabot creating any more for this dependency (unless you reopen the
PR or upgrade to it yourself)


</details>

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
…fy to be textbook like as possible (apache#23761)

## Which issue does this PR close?

N/A

## Rationale for this change

SortMergeJoin bitwise stream implementation is very complex and hard to
understand while on paper it should be pretty simple.

the reason for that is we have to store state between polls (we had
`boundary`) and handle the case where both can get `Poll::Pending` from
child and `Poll::Ready` from child which further complicate the code

## What changes are included in this PR?

1. Move to async generators
2. Rewrote main loop to be textbook like as possible

(this was entirely written by Claude Fable, sorry, I tried manually but
the code was too complex to hold in my head 😅 )

## Are these changes tested?
existing tests

## Are there any user-facing changes?
The join_time now includes the time to read from the async spill stream
between pending which is arguable more correct since this time is part
of the operator, although long waits between pending calls will be
counted in the op `join_time` while the alternative is not counting the
read from file and decoding...
…lator (apache#23946)

## Rationale for this change

Follow-up to post-merge review feedback from @neilconway on apache#23913.

## What changes are included in this PR?

- Use the fast foldhash `RandomState` for the distinct-value count map
instead of the standard library's default SipHash (the shared
`GenericDistinctBuffer` already does this; the merged fix regressed to
SipHash on this hot path).
- Use `estimate_memory_size` in `size()` instead of a hand-rolled
capacity calculation.
- Adopt the null-free fast path in `update_batch`/`retract_batch` (skip
per-element validity checks when the input has no nulls), mirroring
`GenericDistinctBuffer`.
- Add `ORDER BY` to the sliding-window regression test so its row order
is deterministic.

## Are these changes tested?

Yes — existing percentile unit tests and the full `aggregate.slt` pass.

## Are there any user-facing changes?

No.
…pstream now carries

DataFusion 55.1.0 adds the same start == end guard (with try_extend_nulls)
immediately above ours, so the Spice copy was unreachable and its
deprecated extend_nulls call warned. The regression tests stay.
- tests use StatisticsContext and PruningPredicateBuilder instead of the
  deprecated partition_statistics / PruningPredicate::try_new
- proto: expect(unused_variables) instead of allow (clippy::allow_attributes)
- rustdoc: fix unresolved and private intra-doc links
- rustfmt and taplo formatting
The aggregate-unprojection arm used an `if let` match guard, which is not
stable on the workspace's rust-version (1.94.0): `cargo +1.94.1 check` fails
with E0658. Move the check into the unqualified-column arm as a let-chain;
same behaviour, and the workspace checks on 1.94.1 again.
Copilot AI balanced review requested due to automatic review settings September 30, 2026 21:56
Comment thread dev/depcheck/Cargo.lock Fixed
Comment thread datafusion/ffi/Cargo.toml Fixed
Comment thread dev/depcheck/Cargo.lock Fixed
Comment thread Cargo.lock Fixed
Comment thread Cargo.lock Fixed
Comment thread Cargo.lock Fixed
Comment thread dev/depcheck/Cargo.lock Fixed

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

estimate_join_cardinality carried each input's column statistics through an inner,
left, right or full join unchanged, sum_value included. A join repeats a row once
per match and drops the rows that match nothing, so the input's sum says nothing
about the output's; kept exact, it let AggregateStatistics answer SUM over the join
with the input table's total (a wrong result with no error). Semi and anti joins
already drop it, and the cross join scales it.
Jeadie added 3 commits October 1, 2026 20:39
ListingTable materializes metadata columns (_last_modified, _size, _location)
from each object's ObjectMeta, but pruning only used partition columns. A
predicate such as `_last_modified > <watermark>` fell through to a row-level
FilterExec, so every file was opened and read before its rows were dropped.
For a compressed row format (e.g. jsonl.gz) this decompressed the whole object
on every scan.

Prune the listing by these predicates before any file is opened:

- `pruned_partition_list_with_metadata` filters the listed ObjectMeta set by
  evaluating the metadata-only predicate against each object. The existing
  `pruned_partition_list` delegates to it with empty metadata arguments, so its
  signature and behavior are unchanged.
- `filter_by_metadata` evaluates the predicate via create_physical_expr against
  a one-row batch built from MetadataColumn::to_scalar_value, giving correct
  >, >=, <, <=, =, BETWEEN, IN, LIKE, cast and NULL semantics. A NULL (unknown)
  result prunes the file, matching SQL WHERE semantics.
- `supports_filters_pushdown` reports a metadata-only conjunct as Exact, so the
  redundant FilterExec is dropped. A mixed metadata+data predicate, or one under
  OR/NOT, stays Inexact and keeps a residual filter (correct, not optimal).

Metadata pruning is orthogonal to partition pruning and applies to both
partitioned and unpartitioned tables, and to all file formats.

Reproduces spiceai/spiceai#14264.
A caller that obtains an ObjectMeta from a HEAD on a known key (rather than a
listing) can apply the identical metadata-column prune before opening the file,
and report the same predicates as Exact.
Compile the metadata predicate once per listing instead of per file
(MetadataPredicate), threading the caller's session ExecutionProps
through so filters using session variables still evaluate correctly.
Deduplicate metadata-column-name extraction between scan_with_args and
supports_filters_pushdown, and add test coverage: unit-level predicate
tests, a listing-helper integration test, and a ListingTable plan-level
test asserting the Exact pushdown marking and scan wiring actually
prune file groups and residual filters as expected. Also documents the
Exact-pushdown contract so a TableProvider wrapping ListingTable with
its own scan path doesn't inherit it without enforcing it.
@krinart

krinart commented Oct 1, 2026

Copy link
Copy Markdown
Author

CI note: this repo's GitHub workflows never start (dispatches stay queued with 0 jobs on spiceai-54 too), so only the license check runs. The CI commands were run locally on this head: cargo clippy --all-targets --workspace --features avro,integration-tests,extended_tests -- -D warnings, rustdoc with -D warnings, rustfmt, taplo, and cargo +1.94.1 check --workspace (MSRV) are clean. #239 adds the SUM-over-join statistics fix.

fix(physical-plan): drop input sums from inner and outer join statistics
Copilot AI balanced review requested due to automatic review settings October 1, 2026 17:14

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Three predicate benchmark datasets violate their stated independence assumptions, invalidating the measurements they are intended to produce.

Review effort: Balanced
Findings: 3 Medium severity

Open (3)

Comment thread benchmarks/sql_benchmarks/predicate_eval/load/corr.sql
Comment thread benchmarks/sql_benchmarks/predicate_eval/load/ints.sql
Comment thread benchmarks/sql_benchmarks/predicate_eval/load/markers.sql
fix(unparser): spell a Date32 literal's cast with the dialect's date type (refs spiceai/spiceai#14491)
Brings across the Spice patches on the 54 line that were missing here:
the unparser spelling a Date32 literal's cast with the dialect's date
type (#237), and pruning ListingTable file listings by metadata-column
predicates (`_last_modified`, `_size`, `_location`), which
`supports_filters_pushdown` reports `Exact`.

table.rs conflicted with DataFusion 55's declared output partitioning.
Both behaviors are kept. `scan_with_args` separates out the metadata
filters and passes them to `list_files_for_scan_with_metadata`, which
dispatches the way 55's `list_files_for_scan` does. The regular scan
prunes by metadata while listing, before the file limit. The declared
partitioning path assigns files to partitions before any filter runs,
then removes files that fail the metadata predicate from their assigned
group, the same way it handles partition filters. A file's partition
therefore never depends on the query. Limits are still not pushed to
listing when output partitioning is declared or residual filters remain.
`MetadataPredicate` now passes the `PhysicalPlanningContext` argument
that 55's `create_physical_expr` requires.

Adds a test for a metadata filter on a table with declared output
partitioning. It checks the filter is `Exact`, that no FilterExec is
planned, that the partition count is kept, that the excluded file is
pruned, and that the correct rows come back.
Merge spiceai-54 into spiceai-55-patches-2 (#237, metadata-column listing pruning)
Copilot AI balanced review requested due to automatic review settings October 1, 2026 20:13

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

The webpack development server upgrade requires Node 22.15 while repository workflows still provision Node 18 and 20.

Review effort: Balanced
Findings: 3 Medium severity

Open (3)

@grokspice

Copy link
Copy Markdown

This looks superseded — worth a decision before any more work goes into it.

spiceai-55 has moved well past 55.1.0 since this was opened:

27351ab0  2026-10-05T20:52:27Z  Merge pull request #248 from spiceai/spiceai-55-patches-4
47b9ec42  2026-10-05T19:08:28Z  Merge upstream DataFusion 55.2.0-rc1 into spiceai-55
500744d5  2026-10-04T20:49:06Z  [branch-55] chore: update version 55.2.0 and add changelog (#26037)

spiceai-55...spiceai-55-patches-2 now reads ahead=7 behind=13 status=diverged, and the two changes still unique to this branch are both already accounted for elsewhere:

So merging this as-is would re-litigate #240 through a side door, and its stated purpose — bringing 55.1.0 in — is already done by the 55.2.0-rc1 merge.

Recommend closing this in favour of #240, or retargeting it to just the piece that is genuinely still missing. Either way it is an author/maintainer call, not one to make from a babysitting pass — leaving it untouched.

@krinart krinart closed this Oct 6, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.