You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
bench: Use variable-length inputs and scalar n in left/right benchmark
The Utf8 inputs were all exactly 32 bytes, while the Utf8View inputs
used a different length distribution, so the two array types were not
benchmarked on the same data. Fixed-length inputs also make per-row
work that depends on the input length, such as an ASCII check over the
whole string, perfectly predictable, so it looks nearly free. And every
case passed `n` as an array cycling through a short range, although `n`
is usually a literal.
Generate variable-length inputs for every case and build the Utf8 and
Utf8View arrays from the same strings. Pass `n` as a scalar, and add a
case with a different random `n` per row. Add cases for short results
from long inputs, for `n` exceeding the input length, and for negative
`n`.
Also derive the return field from the function instead of always using
Utf8View. Since #23330, `left` and `right` return Utf8 for Utf8 input,
and the mismatch fails the check in `ScalarUDF::invoke_with_args` when
debug assertions are enabled.
0 commit comments