Skip to content

perf: Fix performance regression in unary math UDFs - #26062

Merged
neilconway merged 7 commits into
apache:mainfrom
neilconway:neilc/perf-sqrt
Oct 8, 2026
Merged

neilconway merged 7 commits into
apache:mainfrom
neilconway:neilc/perf-sqrt

Conversation

@neilconway

@neilconway neilconway commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

Which issue does this PR close?

  • N/A

Rationale for this change

#22308 changed sqrt to raise an error on negative floating point inputs. In the course of doing that, it changed all of the unary math UDFs to use try_unary instead of unary. Switching to try_unary regressed the performance of those UDFs. The effect was the most extreme for sqrt, because #22308 also added a per-value check that inhibited vectorization, but the other UDFs also suffered from try_unary's additional overhead.

This PR preserves the error-handling change in #22308 but implements it in a different way: we add support for an optional "input check" function to the unary math macro. If supplied, the check function is applied to every input value (including null slot) in a branch-free loop, which doesn't inhibit vectorization. If that loop detects an erroneous input, we do a second branching loop to find the problematic value to report the error. With this scheme, all the math functions can go back to using unary.

Using unary is not always a win: for (1) expensive math functions on (2) NULL-heavy data sets, the wasted work from invoking the function on null slots can exceed the overhead of checking the NULL bitmap, and expensive math functions typically prevent vectorization anyway. It would be possible to try to be smarter here about the threshold when try_unary becomes faster than unary, but for now this PR reverts to the pre-#22308 performance.

Benchmarks: (x86, AMD EPYC Milan):

  • sqrt/f64: 32.3µs → 9.7µs, −70%
  • sqrt/f32: 27.6µs → 3.0µs, −89%
  • sqrt/f64 with nulls: 34.0µs → 9.8µs, −71%
  • sqrt/f32 with nulls: 30.8µs → 3.0µs, −90%
  • degrees/f64: 1.90µs → 1.32µs, −30%
  • degrees/f32: 1.05µs → 0.76µs, −27%
  • degrees/f64 with nulls: 5.1µs → 1.33µs, −74%
  • degrees/f32 with nulls: 4.5µs → 0.77µs, −83%

What changes are included in this PR?

See above. Also added a benchmark for sqrt and degrees (degrees is a trivial math UDF that effectively measures the dispatch overhead). Other, more expensive math UDFs still got slower but their relative slowdown was much less.

What is the testing strategy for this PR?

Existing tests pass; added new unit test for new error-check facility.

Are there any user-facing changes?

No.

Benchmark `sqrt` and `degrees` on Float64 and Float32 arrays, with and
without nulls.
apache#22308 made `sqrt` return an error for negative inputs by switching
every unary math function from `unary` to `try_unary` with a per-value
validator. `try_unary` can return early on any value, so its loop is
not vectorized, and for nullable input it visits valid indices one at a
time. That made `sqrt`, which otherwise compiles to vector square-root
instructions, several times slower.

Add `unary_with_input_check`, which applies the function and checks
every value in one branch-free loop, then looks for an error to report
only if some value failed the check. `sqrt`'s validator now returns an
error message instead of a `Result`, so the check never builds an
error. The other unary math functions go back to using `unary`.

The error is now reported as an execution error, rather than wrapped in
an Arrow compute error.
@neilconway

Copy link
Copy Markdown
Contributor Author

FYI @xiedeyantu @Dandandan

@codecov-commenter

codecov-commenter commented Oct 5, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 82.74%. Comparing base (1fa378d) to head (92e1f62).
⚠️ Report is 15 commits behind head on main.

Additional details and impacted files
@@           Coverage Diff            @@
##             main   #26062    +/-   ##
========================================
  Coverage   82.73%   82.74%            
========================================
  Files        1147     1147            
  Lines      449353   449795   +442     
  Branches   449353   449795   +442     
========================================
+ Hits       371783   372184   +401     
- Misses      54916    54943    +27     
- Partials    22654    22668    +14     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@alamb

alamb commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

run benchmark math_expressions

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark running (GKE) | trigger
Instance: c4a-highmem-16 (12 vCPU / 65 GiB) | Linux bench-c6006040989-3076-glj89 6.12.94+ #1 SMP Fri Aug 21 08:00:16 UTC 2026 aarch64 GNU/Linux

CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected

Comparing neilc/perf-sqrt (c9a5496) to 8248a57 (merge-base) diff

Run configuration
run benchmark math_expressions

Results will be posted here when complete


File an issue against this benchmark runner

@github-actions github-actions Bot added the functions Changes to functions implementation label Oct 6, 2026

@alamb alamb left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good to me -- thanks @neilconway -- I reviewed the logic and the comments and kicked off a benchmark

I also took the liberty of pushing a commit to fix the CI error: https://github.com/apache/datafusion/actions/runs/37379821165/job/111998601947?pr=26062

Comment thread datafusion/functions/src/math/common.rs Outdated
/// Use this for functions that return an error for some argument values, such
/// as `sqrt`, which returns an error for negative numbers. `try_unary` can also
/// return errors, but it can return early on any value, which keeps the
/// compiler from vectorizing its loop. That makes cheap functions like `sqrt`

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I feel like the mention down here about vectorization is "burying the lead" so to speak -- it might be better if the comment started with a mention of "faster version of try_unary that gives the compiler the best chance to vectorize the check and the operation" or something

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks, I revised the comment to change the emphasis and describe the performance tradeoffs more clearly.

/// slots, in the same loop as `op`; only if some value fails is the array
/// searched again for an error to report. `input_error` should therefore be a
/// cheap check, such as a comparison.
pub(crate) fn unary_with_input_check<T: ArrowPrimitiveType>(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

this seems like it might be a good one to propose porting upstream to arrows rs (or as an example on try_unary 🤔

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yeah, I've been thinking about that. Adding something to Arrow definitely makes sense, although because making the right choice depends on a bunch of factors (e.g., how expensive the op is, whether it is vectorizable, null density, CPU architecture / version of SIMD), we'd want to make sure that users have a clear decision process for which primitive to use.

.values()
.iter()
.map(|&x| {
any_invalid |= input_error(x).is_some();

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

it is interesting this is mre vectorizable -- I was sot of expecting two loops

Comment thread datafusion/functions/src/math/common.rs Outdated
.collect();

// The check above also ran on null slots, which can hold any value, so the
// failure may be spurious. Re-check just the non-null values.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

👍

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

though in theory this can run and slow down arrays with nulls

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ah, right, this is a very good point! try_unary on NULL-heavy inputs can actually be faster. Measuring on my local machine:

  • For cheap functions like sqrt, degrees, radians, unary is faster for <= 80% NULLs.
  • For expensive functions like exp, sin, and ln, the cross-over point is something like 25% NULLs.

In an extreme case of 90% nulls (randomly distributed), exp is about 5 times faster using try_unary than with unary. Whereas with no nulls, unary is about 10% faster.

Intuitively, expensive functions (a) can't be vectorized anyway (b) waste more work on the values in null slots.

I think reverting the blanket switch to try_unary stlll makes sense, because it was probably unintended, and I'd suspect that most people invoking math functions on large data sets won't have NULL-heavy data. But we could certainly try to make this more intelligent, e.g., by applying a heuristic based on NULL density to switch to try_unary, or by having the more expensive functions always use try_unary. Either way I'd be inclined to leave it to a followup PR.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A related behavior is that null slots can hold arbitrary values; if those values happen to fail the error check, we'll need to take the slow path, even if all the valid slots are not erroneous. But we can at least do better here than the initial version of the PR -- in the slow path, we were skipping NULLs on the recheck with a filter(), but it's faster to use a bitmap to mask out null slots. I added that optimization to the PR.

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark completed (GKE) | trigger

Instance: c4a-highmem-16 (12 vCPU / 65 GiB)

Comparing neilc/perf-sqrt (c9a5496) to 8248a57 (merge-base) diff

Run configuration
run benchmark math_expressions
CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected
Details

group                                            HEAD                                   neilc_perf-sqrt
-----                                            ----                                   ---------------
atan2 f32 array: 1024                            1.00      9.8±0.01µs        ? ?/sec    1.00      9.8±0.01µs        ? ?/sec
atan2 f32 array: 4096                            1.00     38.7±0.09µs        ? ?/sec    1.00     38.7±0.03µs        ? ?/sec
atan2 f32 array: 8192                            1.00     78.4±0.04µs        ? ?/sec    1.01     78.8±0.16µs        ? ?/sec
atan2 f32 scalar                                 1.00    158.3±0.32ns        ? ?/sec    1.02    161.6±1.86ns        ? ?/sec
atan2 f64 array: 1024                            1.01      9.7±0.01µs        ? ?/sec    1.00      9.6±0.01µs        ? ?/sec
atan2 f64 array: 4096                            1.00     38.0±0.06µs        ? ?/sec    1.00     37.9±0.03µs        ? ?/sec
atan2 f64 array: 8192                            1.00     75.9±0.30µs        ? ?/sec    1.00     75.9±0.05µs        ? ?/sec
atan2 f64 scalar                                 1.00    146.8±0.96ns        ? ?/sec    1.04    152.2±3.04ns        ? ?/sec
cot f32 array: 1024                              1.00      7.1±0.01µs        ? ?/sec    1.00      7.1±0.00µs        ? ?/sec
cot f32 array: 4096                              1.00     27.0±0.02µs        ? ?/sec    1.00     27.1±0.01µs        ? ?/sec
cot f32 array: 8192                              1.00     57.9±0.11µs        ? ?/sec    1.00     57.8±0.11µs        ? ?/sec
cot f32 scalar                                   1.00    136.9±0.63ns        ? ?/sec    1.01    137.9±0.78ns        ? ?/sec
cot f64 array: 1024                              1.00      6.3±0.01µs        ? ?/sec    1.00      6.3±0.00µs        ? ?/sec
cot f64 array: 4096                              1.00     23.9±0.01µs        ? ?/sec    1.00     23.9±0.02µs        ? ?/sec
cot f64 array: 8192                              1.00     54.1±0.18µs        ? ?/sec    1.00     54.1±0.14µs        ? ?/sec
cot f64 scalar                                   1.00    125.6±0.93ns        ? ?/sec    1.01    126.4±0.98ns        ? ?/sec
degrees/f32                                                                             1.00    642.0±2.04ns        ? ?/sec
degrees/f32 with nulls                                                                  1.00    657.6±3.09ns        ? ?/sec
degrees/f64                                                                             1.00   1057.9±4.56ns        ? ?/sec
degrees/f64 with nulls                                                                  1.00   1072.2±9.65ns        ? ?/sec
factorial_array                                  1.00      3.0±0.10µs        ? ?/sec    1.00      3.0±0.10µs        ? ?/sec
factorial_scalar                                 1.00    161.7±2.30ns        ? ?/sec    1.03    166.9±2.82ns        ? ?/sec
floor_ceil size=1024/ceil_f64_array              1.00    374.0±0.80ns        ? ?/sec    1.01    377.2±1.26ns        ? ?/sec
floor_ceil size=1024/ceil_f64_scalar             1.00    126.7±0.09ns        ? ?/sec    1.00    126.9±0.20ns        ? ?/sec
floor_ceil size=1024/floor_f64_array             1.02    381.7±0.28ns        ? ?/sec    1.00    376.1±0.30ns        ? ?/sec
floor_ceil size=1024/floor_f64_scalar            1.00    127.4±0.11ns        ? ?/sec    1.05    134.2±0.14ns        ? ?/sec
floor_ceil size=4096/ceil_f64_array              1.00    668.8±0.19ns        ? ?/sec    1.01    677.3±0.38ns        ? ?/sec
floor_ceil size=4096/ceil_f64_scalar             1.00    126.6±0.06ns        ? ?/sec    1.00    126.6±0.09ns        ? ?/sec
floor_ceil size=4096/floor_f64_array             1.00    676.0±0.75ns        ? ?/sec    1.00    677.0±0.60ns        ? ?/sec
floor_ceil size=4096/floor_f64_scalar            1.01    127.2±0.04ns        ? ?/sec    1.00    125.9±0.07ns        ? ?/sec
floor_ceil size=8192/ceil_f64_array              1.00   1081.0±2.52ns        ? ?/sec    1.00   1081.1±1.48ns        ? ?/sec
floor_ceil size=8192/ceil_f64_scalar             1.00    126.7±0.15ns        ? ?/sec    1.00    126.5±0.10ns        ? ?/sec
floor_ceil size=8192/floor_f64_array             1.01   1094.3±1.96ns        ? ?/sec    1.00   1082.1±2.70ns        ? ?/sec
floor_ceil size=8192/floor_f64_scalar            1.00    126.8±0.03ns        ? ?/sec    1.00    126.5±0.19ns        ? ?/sec
gcd array and scalar                             1.00   1771.2±2.41µs        ? ?/sec    1.00   1776.8±5.17µs        ? ?/sec
gcd both array                                   1.00   1416.2±2.94µs        ? ?/sec    1.00   1419.5±2.81µs        ? ?/sec
gcd both scalar                                  1.00    195.9±3.74ns        ? ?/sec    1.01    198.8±4.55ns        ? ?/sec
isnan f32 array: 1024                            1.00    441.5±1.32ns        ? ?/sec    1.00    442.7±0.96ns        ? ?/sec
isnan f32 array: 4096                            1.00   1034.2±1.47ns        ? ?/sec    1.00   1038.5±4.24ns        ? ?/sec
isnan f32 array: 8192                            1.00   1834.7±2.30ns        ? ?/sec    1.01   1845.1±3.21ns        ? ?/sec
isnan f64 array: 1024                            1.00    403.7±1.32ns        ? ?/sec    1.00    403.1±1.75ns        ? ?/sec
isnan f64 array: 4096                            1.00    882.2±1.35ns        ? ?/sec    1.00    880.8±1.09ns        ? ?/sec
isnan f64 array: 8192                            1.00   1542.0±2.65ns        ? ?/sec    1.01   1560.7±2.58ns        ? ?/sec
iszero f32 array: 1024                           1.00    388.4±2.31ns        ? ?/sec    1.01    391.2±3.48ns        ? ?/sec
iszero f32 array: 4096                           1.00    900.2±2.27ns        ? ?/sec    1.01    906.7±1.47ns        ? ?/sec
iszero f32 array: 8192                           1.00   1584.6±2.18ns        ? ?/sec    1.01   1598.1±2.69ns        ? ?/sec
iszero f32 scalar                                1.02    107.8±0.50ns        ? ?/sec    1.00    105.2±0.53ns        ? ?/sec
iszero f64 array: 1024                           1.00    374.6±2.27ns        ? ?/sec    1.01    377.7±2.34ns        ? ?/sec
iszero f64 array: 4096                           1.00    854.3±2.27ns        ? ?/sec    1.00    857.6±2.92ns        ? ?/sec
iszero f64 array: 8192                           1.00   1511.0±2.33ns        ? ?/sec    1.00   1512.0±3.00ns        ? ?/sec
iszero f64 scalar                                1.00     98.6±0.55ns        ? ?/sec    1.00     98.5±0.73ns        ? ?/sec
lcm both array                                   1.00   1518.1±3.93µs        ? ?/sec    1.00   1512.1±1.97µs        ? ?/sec
nanvl/array_f32/1024                             1.00    408.3±1.95ns        ? ?/sec    1.01    412.1±1.52ns        ? ?/sec
nanvl/array_f32/4096                             1.00    649.9±2.41ns        ? ?/sec    1.04    675.0±3.88ns        ? ?/sec
nanvl/array_f32/8192                             1.00   1008.6±4.18ns        ? ?/sec    1.04   1054.0±2.36ns        ? ?/sec
nanvl/array_f64/1024                             1.02    496.6±2.50ns        ? ?/sec    1.00    485.3±1.86ns        ? ?/sec
nanvl/array_f64/4096                             1.02   1032.7±4.85ns        ? ?/sec    1.00   1014.5±3.06ns        ? ?/sec
nanvl/array_f64/8192                             1.00   1778.6±3.27ns        ? ?/sec    1.09   1935.5±3.71ns        ? ?/sec
nanvl/array_f64_both_nulls/1024                  1.00      2.9±0.15µs        ? ?/sec    1.00      2.9±0.15µs        ? ?/sec
nanvl/array_f64_both_nulls/4096                  1.00     10.3±0.68µs        ? ?/sec    1.00     10.3±0.72µs        ? ?/sec
nanvl/array_f64_both_nulls/8192                  1.00     20.2±1.33µs        ? ?/sec    1.00     20.2±1.51µs        ? ?/sec
nanvl/array_f64_x_nulls/1024                     1.00      2.9±0.18µs        ? ?/sec    1.00      2.9±0.18µs        ? ?/sec
nanvl/array_f64_x_nulls/4096                     1.00     10.2±0.78µs        ? ?/sec    1.01     10.3±0.77µs        ? ?/sec
nanvl/array_f64_x_nulls/8192                     1.00     20.0±1.59µs        ? ?/sec    1.02     20.4±1.42µs        ? ?/sec
nanvl/array_f64_y_nulls/1024                     1.01      3.0±0.12µs        ? ?/sec    1.00      3.0±0.14µs        ? ?/sec
nanvl/array_f64_y_nulls/4096                     1.00     10.5±0.52µs        ? ?/sec    1.06     11.2±0.25µs        ? ?/sec
nanvl/array_f64_y_nulls/8192                     1.00     20.3±1.30µs        ? ?/sec    1.08     22.0±0.54µs        ? ?/sec
nanvl/scalar_f32                                 1.00    131.3±1.40ns        ? ?/sec    1.01    132.0±2.76ns        ? ?/sec
nanvl/scalar_f64                                 1.00    119.1±1.09ns        ? ?/sec    1.01    120.3±3.15ns        ? ?/sec
power f64 array x f64 array, n=1024              1.00      8.8±0.02µs        ? ?/sec    1.01      8.9±0.02µs        ? ?/sec
power f64 array x f64 array, n=8192              1.00     65.8±0.16µs        ? ?/sec    1.00     65.9±0.18µs        ? ?/sec
power f64 array x f64 scalar, exp=0.5, n=1024    1.00      9.2±0.05µs        ? ?/sec    1.00      9.3±0.05µs        ? ?/sec
power f64 array x f64 scalar, exp=0.5, n=8192    1.00     71.6±0.38µs        ? ?/sec    1.00     71.7±0.39µs        ? ?/sec
power f64 array x f64 scalar, exp=2, n=1024      1.00      9.3±0.05µs        ? ?/sec    1.00      9.3±0.05µs        ? ?/sec
power f64 array x f64 scalar, exp=2, n=8192      1.00     71.6±0.39µs        ? ?/sec    1.00     71.7±0.39µs        ? ?/sec
random_1M_rows_batch_128                         1.00     13.9±0.54ms        ? ?/sec    1.00     14.0±0.55ms        ? ?/sec
random_1M_rows_batch_8192                        1.00     12.4±0.57ms        ? ?/sec    1.00     12.4±0.57ms        ? ?/sec
round size=1024/round_f32_array                  1.00    604.1±1.67ns        ? ?/sec    1.02    617.0±1.45ns        ? ?/sec
round size=1024/round_f32_scalar                 1.00    184.6±1.22ns        ? ?/sec    1.04    192.3±0.86ns        ? ?/sec
round size=1024/round_f64_array                  1.00   1214.8±5.24ns        ? ?/sec    1.01   1223.3±1.64ns        ? ?/sec
round size=1024/round_f64_scalar                 1.00    186.9±0.88ns        ? ?/sec    1.01    188.7±0.36ns        ? ?/sec
round size=4096/round_f32_array                  1.00   1377.5±3.70ns        ? ?/sec    1.01   1394.1±2.41ns        ? ?/sec
round size=4096/round_f32_scalar                 1.00    185.5±1.39ns        ? ?/sec    1.03    191.5±1.21ns        ? ?/sec
round size=4096/round_f64_array                  1.00      3.8±0.00µs        ? ?/sec    1.00      3.8±0.00µs        ? ?/sec
round size=4096/round_f64_scalar                 1.00    184.2±1.30ns        ? ?/sec    1.01    186.8±0.58ns        ? ?/sec
round size=8192/round_f32_array                  1.00      2.4±0.00µs        ? ?/sec    1.00      2.4±0.00µs        ? ?/sec
round size=8192/round_f32_scalar                 1.00    185.4±1.39ns        ? ?/sec    1.04    192.8±1.24ns        ? ?/sec
round size=8192/round_f64_array                  1.00      7.2±0.00µs        ? ?/sec    1.00      7.2±0.00µs        ? ?/sec
round size=8192/round_f64_scalar                 1.00    183.2±1.27ns        ? ?/sec    1.02    186.7±0.70ns        ? ?/sec
round_dense_f32/1024                             1.00    585.0±1.47ns        ? ?/sec    1.02    594.8±5.01ns        ? ?/sec
round_dense_f32/4096                             1.00   1367.7±3.82ns        ? ?/sec    1.00   1366.8±4.99ns        ? ?/sec
round_dense_f32/8192                             1.00      2.4±0.00µs        ? ?/sec    1.00      2.4±0.00µs        ? ?/sec
round_dense_f64/1024                             1.00   1192.8±2.37ns        ? ?/sec    1.01   1204.8±7.62ns        ? ?/sec
round_dense_f64/4096                             1.00      3.8±0.00µs        ? ?/sec    1.00      3.8±0.00µs        ? ?/sec
round_dense_f64/8192                             1.00      7.2±0.00µs        ? ?/sec    1.00      7.2±0.01µs        ? ?/sec
signum f32 array: 1024                           1.00    402.6±0.97ns        ? ?/sec    1.02    411.3±2.12ns        ? ?/sec
signum f32 array: 4096                           1.00    873.5±1.09ns        ? ?/sec    1.00    875.0±1.11ns        ? ?/sec
signum f32 array: 8192                           1.00   1501.0±1.21ns        ? ?/sec    1.00   1497.0±1.23ns        ? ?/sec
signum f32 scalar: 1024                          1.05    115.9±6.56ns        ? ?/sec    1.00    110.3±0.26ns        ? ?/sec
signum f32 scalar: 4096                          1.00    108.8±0.21ns        ? ?/sec    1.01    110.3±0.26ns        ? ?/sec
signum f32 scalar: 8192                          1.00    108.9±0.36ns        ? ?/sec    1.01    110.3±0.24ns        ? ?/sec
signum f64 array: 1024                           1.00   1017.8±1.55ns        ? ?/sec    1.01   1022.9±2.95ns        ? ?/sec
signum f64 array: 4096                           1.00      3.3±0.00µs        ? ?/sec    1.00      3.3±0.00µs        ? ?/sec
signum f64 array: 8192                           1.00      6.4±0.00µs        ? ?/sec    1.00      6.4±0.01µs        ? ?/sec
signum f64 scalar: 1024                          1.00    100.8±0.90ns        ? ?/sec    1.01    102.2±0.78ns        ? ?/sec
signum f64 scalar: 4096                          1.06    107.9±2.95ns        ? ?/sec    1.00    102.1±0.71ns        ? ?/sec
signum f64 scalar: 8192                          1.00    100.9±0.90ns        ? ?/sec    1.01    102.2±0.71ns        ? ?/sec
sqrt/f32                                                                                1.00      2.7±0.01µs        ? ?/sec
sqrt/f32 with nulls                                                                     1.00      2.7±0.00µs        ? ?/sec
sqrt/f64                                                                                1.00      9.9±0.01µs        ? ?/sec
sqrt/f64 with nulls                                                                     1.00      9.9±0.00µs        ? ?/sec
trunc f32 array: 1024                            1.00    365.4±1.27ns        ? ?/sec    1.00    366.9±1.11ns        ? ?/sec
trunc f32 array: 4096                            1.00    675.6±2.85ns        ? ?/sec    1.00    674.6±2.47ns        ? ?/sec
trunc f32 array: 8192                            1.00   1100.6±1.11ns        ? ?/sec    1.00   1105.7±1.80ns        ? ?/sec
trunc f32 precision array: 1024                  1.00    502.6±0.96ns        ? ?/sec    1.01    507.4±2.66ns        ? ?/sec
trunc f32 precision array: 4096                  1.00   1287.4±5.30ns        ? ?/sec    1.00   1288.0±3.09ns        ? ?/sec
trunc f32 precision array: 8192                  1.00      2.3±0.00µs        ? ?/sec    1.00      2.3±0.00µs        ? ?/sec
trunc f32 scalar                                 1.00     67.1±0.40ns        ? ?/sec    1.00     66.9±0.34ns        ? ?/sec
trunc f64 array: 1024                            1.00    423.4±2.06ns        ? ?/sec    1.01    428.8±1.93ns        ? ?/sec
trunc f64 array: 4096                            1.00    916.3±3.30ns        ? ?/sec    1.01    923.8±2.07ns        ? ?/sec
trunc f64 array: 8192                            1.00   1564.4±3.47ns        ? ?/sec    1.00   1558.1±3.93ns        ? ?/sec
trunc f64 precision array: 1024                  1.00   1089.1±1.73ns        ? ?/sec    1.01   1096.0±2.70ns        ? ?/sec
trunc f64 precision array: 4096                  1.00      3.6±0.00µs        ? ?/sec    1.00      3.6±0.00µs        ? ?/sec
trunc f64 precision array: 8192                  1.00      7.0±0.00µs        ? ?/sec    1.00      7.0±0.00µs        ? ?/sec
trunc f64 scalar                                 1.01     64.0±0.55ns        ? ?/sec    1.00     63.7±0.66ns        ? ?/sec

Resource Usage

math_expressions — base (merge-base)

Metric Value
Wall time 1390.3s
Peak memory 37.0 MiB
Avg memory 30.0 MiB
CPU user 1484.4s
CPU sys 1.3s
Peak spill 0 B

math_expressions — branch

Metric Value
Wall time 1475.3s
Peak memory 38.1 MiB
Avg memory 31.6 MiB
CPU user 1587.6s
CPU sys 1.4s
Peak spill 0 B

File an issue against this benchmark runner

@neilconway

Copy link
Copy Markdown
Contributor Author

@alamb FWIW the math_expressions benchmarks (prior to this PR) don't cover any of the UDFs that are implemented with make_math_unary_udf!, so I wouldn't expect to see wins. I can split the PRs in this benchmark out into a separate PR if you like.

Lead with what the helper is for (vectorizing the input check and `op`)
instead of burying it, and say when `try_unary` can still be faster:
arrays with many nulls, especially with expensive operations.
If a null slot holds a value that fails the input check, every batch
pays for the second pass even though there is no error. That happens
easily in practice: arithmetic kernels compute on null slots too, so
the null slots of `x - 10` usually hold -10.

Build the re-check as a bitmap with `collect_bool` and AND it with the
validity bitmap, instead of testing each value's null bit and
branching. On an array whose null slots fail the check, this takes the
re-check from about 5.5µs to 2.8µs per 8192 rows; arrays that pass the
first check are unaffected.

@jayzhan211 jayzhan211 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @neilconway , overall LGTM

Comment thread datafusion/functions/src/macros.rs
@neilconway

Copy link
Copy Markdown
Contributor Author

@jayzhan211 @alamb Thanks for the reviews!

@neilconway
neilconway enabled auto-merge October 8, 2026 13:48
@github-actions github-actions Bot added the sqllogictest SQL Logic Tests (.slt) label Oct 8, 2026
@neilconway
neilconway added this pull request to the merge queue Oct 8, 2026
@alamb

alamb commented Oct 8, 2026

Copy link
Copy Markdown
Contributor

Thank you

Merged via the queue into apache:main with commit 4978b30 Oct 8, 2026
42 checks passed
@neilconway
neilconway deleted the neilc/perf-sqrt branch October 8, 2026 15:50
@Dandandan

Copy link
Copy Markdown
Contributor

Nice!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

functions Changes to functions implementation sqllogictest SQL Logic Tests (.slt) v56.0.0

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants