Skip to content

perf: Optimize left, right - #26039

Open
neilconway wants to merge 3 commits into
apache:mainfrom
neilconway:neilc/perf-left-right-ascii-prefix
Open

neilconway wants to merge 3 commits into
apache:mainfrom
neilconway:neilc/perf-left-right-ascii-prefix

Conversation

@neilconway

Copy link
Copy Markdown
Contributor

Which issue does this PR close?

  • N/A

Rationale for this change

#23762 added an ASCII fast-path for left and right. That improves performance in many scenarios, but the implementation calls is_ascii on every input string. That is expensive, particularly for the common case that left and right are used to fetch a small prefix/suffix from a much longer string. That check is also overly conservative: for example, we can take the ASCII fast-path for left(s, k) if the first k bytes in the string are ASCII, even if there are multibyte characters elsewhere in the string.

This PR implements two optimizations:

  1. Only call is_ascii on the bytes necessary to determine if we can take the fast-path, not the entire input string, as described above.
  2. Benchmarking identified that for Utf8View inputs, the first optimization regressed some benchmark cases (e.g., negative_n) on both ARM and x86. Claude's theory is that make_view is out-of-line and does an indirect jump on the length of the result string; it seems that after implementing the first optimization, this jump was not well-handled by the branch predictor. Instead, we add a helper sub_view that returns a view that is a substring of an existing view. This can be inlined and avoids the indirect jump incurred by make_view; it can also construct the new view from the old view with bitwise ops, rather than building the new view on the stack.

Benchmarks: (separate PR #26026)

x86 (AMD EPYC Milan)

  • left Utf8 long_result: 101.7 → 93.4 (−8%)
  • left Utf8 n_exceeds_len: 91.7 → 93.8 (+2%)
  • left Utf8 negative_n: 90.4 → 83.2 (−8%)
  • left Utf8 per_row_n: 93.9 → 89.5 (−5%)
  • left Utf8 short_result: 91.0 → 82.6 (−9%)
  • left Utf8 short_result_long_input: 310.4 → 89.4 (−71%)
  • left Utf8View long_result: 53.6 → 40.9 (−24%)
  • left Utf8View n_exceeds_len: 55.9 → 50.7 (−9%)
  • left Utf8View negative_n: 58.8 → 45.2 (−23%)
  • left Utf8View per_row_n: 89.0 → 45.2 (−49%)
  • left Utf8View short_result: 54.2 → 47.5 (−12%)
  • left Utf8View short_result_long_input: 249.1 → 52.7 (−79%)
  • right Utf8 long_result: 109.4 → 91.8 (−16%)
  • right Utf8 n_exceeds_len: 97.9 → 98.5 (+1%)
  • right Utf8 negative_n: 91.5 → 84.7 (−7%)
  • right Utf8 per_row_n: 100.7 → 88.1 (−13%)
  • right Utf8 short_result: 96.8 → 87.2 (−10%)
  • right Utf8 short_result_long_input: 261.2 → 95.5 (−63%)
  • right Utf8View long_result: 64.2 → 53.8 (−16%)
  • right Utf8View n_exceeds_len: 63.9 → 56.0 (−12%)
  • right Utf8View negative_n: 62.3 → 55.8 (−10%)
  • right Utf8View per_row_n: 101.0 → 54.9 (−46%)
  • right Utf8View short_result: 62.9 → 58.7 (−7%)
  • right Utf8View short_result_long_input: 202.2 → 67.7 (−67%)

ARM (Apple M4 Max)

  • left Utf8 long_result: 91.3 → 62.6 (−31%)
  • left Utf8 n_exceeds_len: 69.4 → 71.4 (+3%)
  • left Utf8 negative_n: 66.8 → 69.5 (+4%)
  • left Utf8 per_row_n: 66.6 → 60.6 (−9%)
  • left Utf8 short_result: 69.1 → 62.7 (−9%)
  • left Utf8 short_result_long_input: 91.1 → 59.6 (−35%)
  • left Utf8View long_result: 45.4 → 26.0 (−43%)
  • left Utf8View n_exceeds_len: 37.6 → 34.4 (−9%)
  • left Utf8View negative_n: 37.8 → 25.9 (−31%)
  • left Utf8View per_row_n: 68.1 → 29.4 (−57%)
  • left Utf8View short_result: 37.9 → 25.0 (−34%)
  • left Utf8View short_result_long_input: 61.1 → 28.9 (−53%)
  • right Utf8 long_result: 104.2 → 64.7 (−38%)
  • right Utf8 n_exceeds_len: 80.1 → 79.1 (−1%)
  • right Utf8 negative_n: 71.4 → 66.0 (−8%)
  • right Utf8 per_row_n: 74.0 → 64.3 (−13%)
  • right Utf8 short_result: 80.2 → 65.4 (−18%)
  • right Utf8 short_result_long_input: 101.1 → 63.4 (−37%)
  • right Utf8View long_result: 62.7 → 32.0 (−49%)
  • right Utf8View n_exceeds_len: 42.0 → 38.5 (−8%)
  • right Utf8View negative_n: 41.3 → 28.5 (−31%)
  • right Utf8View per_row_n: 72.0 → 33.5 (−53%)
  • right Utf8View short_result: 43.9 → 30.0 (−32%)
  • right Utf8View short_result_long_input: 63.7 → 33.3 (−48%)

What changes are included in this PR?

See above.

What is the testing strategy for this PR?

Existing tests pass; no functional changes.

Are there any user-facing changes?

No.

The Utf8 inputs were all exactly 32 bytes, while the Utf8View inputs
used a different length distribution, so the two array types were not
benchmarked on the same data. Fixed-length inputs also make per-row
work that depends on the input length, such as an ASCII check over the
whole string, perfectly predictable, so it looks nearly free. And every
case passed `n` as an array cycling through a short range, although `n`
is usually a literal.

Generate variable-length inputs for every case and build the Utf8 and
Utf8View arrays from the same strings. Pass `n` as a scalar, and add a
case with a different random `n` per row. Add cases for short results
from long inputs, for `n` exceeding the input length, and for negative
`n`.

Also derive the return field from the function instead of always using
Utf8View. Since apache#23330, `left` and `right` return Utf8 for Utf8 input,
and the mismatch fails the check in `ScalarUDF::invoke_with_args` when
debug assertions are enabled.
apache#23762 added an ASCII fast path to `left_right_byte_length` that calls
`is_ascii()` on the whole string for every row. For short results from
longer strings, such as Utf8View data read from Parquet, that scan costs
more than the per-character scan it replaced.

ASCII bytes are never part of a multi-byte UTF-8 sequence, so if the `n`
bytes at the relevant end of the string are ASCII, they are exactly the
`n` characters at that end. Check only those bytes, and fall back to the
per-character scan otherwise.
`make_view` selects per-length copy code with a jump on the result
length. When result lengths vary from row to row, as with a negative
`n`, the CPU often mispredicts that jump. Checking only the result
bytes for ASCII removed a length-dependent loop that had been making
the jump predictable, so those cases got slower.

Add `sub_view`, which builds the view for a substring of an existing
view with shifts and masks: from the view itself when the source is
inlined, and otherwise from a 12-byte window of the source that is
always in bounds. Use it for Utf8View input in `left` and `right`.
@github-actions github-actions Bot added the functions Changes to functions implementation label Oct 4, 2026
@neilconway

Copy link
Copy Markdown
Contributor Author

FYI @andygrove @comphead

Comment on lines +1239 to +1271
pub(crate) fn sub_view(view: u128, source: &[u8], range: Range<usize>) -> u128 {
debug_assert!(range.start <= range.end && range.end <= source.len());

// The substring's first bytes, in the low-order bits. Any bits past the
// end of the substring are masked off by `inline_view`.
let leading_bytes = if source.len() <= MAX_INLINE_LEN {
// `source` is stored in `view` itself, after its 4-byte length.
(view >> 32) >> (8 * range.start)
} else {
// `source` has more than 12 bytes, so read the 12 bytes starting at
// `range.start`, or the last 12 bytes if that would run past the end,
// and skip any that come before `range.start`.
let window_start = range.start.min(source.len() - MAX_INLINE_LEN);
let window = source[window_start..window_start + MAX_INLINE_LEN]
.try_into()
.unwrap();
read_12_bytes(window) >> (8 * (range.start - window_start))
};

let len = range.len();
if len <= MAX_INLINE_LEN {
inline_view(leading_bytes, len)
} else {
let original = ByteView::from(view);
ByteView {
length: len as u32,
prefix: leading_bytes as u32,
offset: original.offset + range.start as u32,
..original
}
.as_u128()
}
}

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This could potentially be used in other places (e.g., substr), but that will require more careful evaluation; I'll defer that for now.

@codecov-commenter

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 82.66%. Comparing base (9547b09) to head (b3e5014).
⚠️ Report is 2 commits behind head on main.

Additional details and impacted files
@@           Coverage Diff            @@
##             main   #26039    +/-   ##
========================================
  Coverage   82.65%   82.66%            
========================================
  Files        1147     1147            
  Lines      446087   446404   +317     
  Branches   446087   446404   +317     
========================================
+ Hits       368721   369010   +289     
- Misses      54980    55001    +21     
- Partials    22386    22393     +7     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

functions Changes to functions implementation

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants