bench: Improve benchmarks for left, right - #26026
Merged
neilconway merged 1 commit intoOct 5, 2026
Merged
Conversation
Contributor
Author
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #26026 +/- ##
========================================
Coverage 82.65% 82.66%
========================================
Files 1147 1147
Lines 446087 446357 +270
Branches 446087 446357 +270
========================================
+ Hits 368721 368971 +250
- Misses 54980 54996 +16
- Partials 22386 22390 +4 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
The Utf8 inputs were all exactly 32 bytes, while the Utf8View inputs used a different length distribution, so the two array types were not benchmarked on the same data. Fixed-length inputs also make per-row work that depends on the input length, such as an ASCII check over the whole string, perfectly predictable, so it looks nearly free. And every case passed `n` as an array cycling through a short range, although `n` is usually a literal. Generate variable-length inputs for every case and build the Utf8 and Utf8View arrays from the same strings. Pass `n` as a scalar, and add a case with a different random `n` per row. Add cases for short results from long inputs, for `n` exceeding the input length, and for negative `n`. Also derive the return field from the function instead of always using Utf8View. Since apache#23330, `left` and `right` return Utf8 for Utf8 input, and the mismatch fails the check in `ScalarUDF::invoke_with_args` when debug assertions are enabled.
neilconway
force-pushed
the
neilc/bench-left-right-variable
branch
from
October 4, 2026 17:26
a412ae8 to
b86d26b
Compare
Jefffrey
approved these changes
Oct 5, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Which issue does this PR close?
Rationale for this change
The
left/rightbenchmarks had some shortcomings:Utf8cases used fixed-length (32-byte) strings, whereasUtf8Viewused variable-length strings. Drawing the inputs from different distributions made it hard to compare results between the two data types; also, using fixed-length strings results in unrealistically good branch prediction, which can hide performance problems on typical real-world data.left(s, k)for a constantkis more common.Utf8View, so the benchmarks panicked in debug mode.What changes are included in this PR?
n, negativen, and fornexceeding the input string lengthWhat is the testing strategy for this PR?
Only benchmark changes.
Are there any user-facing changes?
No.