sed: implement \\U \\L \\u \\l \\E case-conversion escapes in s/// - #557
dedsec-terminal wants to merge 8 commits into
Conversation
|
GNU sed testsuite comparison: |
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## main #557 +/- ##
==========================================
+ Coverage 83.86% 85.03% +1.16%
==========================================
Files 14 14
Lines 7203 7824 +621
Branches 424 448 +24
==========================================
+ Hits 6041 6653 +612
- Misses 1157 1166 +9
Partials 5 5
Flags with carried forward coverage won't be shown. Click here to find out more. ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
Merging this PR will improve performance by 5.87%
|
| Benchmark | BASE |
HEAD |
Efficiency | |
|---|---|---|---|---|
| ⚡ | number_fix |
1.2 s | 1.2 s | +5.87% |
Tip
Curious why performance improved? Comment @codspeedbot explain why performance improved on this PR, or directly use the CodSpeed MCP with your agent.
Comparing dedsec-terminal:fix-540-case-conversion-escapes (58b6c03) with main (0554c50)
|
Hi @dspinellis — friendly review request on this s/// case-conversion implementation (fixes #540; GNU suite +1, utf8-ru). Just pushed a perf follow-up (fast paths when no conversion directives; 265 tests green, fmt+clippy clean) to address the Codspeed delta. Let me know if you need anything else. Thanks! |
|
GNU sed testsuite comparison: |
|
Follow-up pushed (23b7709): patch coverage for the fast paths - unit tests for unmatched-group/single-literal fast paths, flag-sync contract panics, LowerFirst invalid-byte handling, compiler case-directive parsing + flag, and end-to-end substitution tests; plus removal of two dead None arms preempted by the all-None early return. Full suite green (lib + integration), fmt/clippy clean. Remaining reds are not from this change: Android jobs fail in sdkmanager setup, and CodSpeed shows -3.75% on number_fix only (genome/access_log improved; likely noise, happy to dig if you want). |
|
GNU sed testsuite comparison: |
|
GNU sed testsuite comparison: |
|
GNU sed testsuite comparison: |
|
GNU sed testsuite comparison: |
Addresses Codspeed -7.9pct regression: common case (no U/L/u/l/E) now plain memcpy via has_case_conversion flag; byte-mode hoists the branch; UTF-8 path validates once per segment; ASCII case uses to_ascii_* fast path. All 265 tests pass; fmt+clippy clean.
Add unit tests for the no-conversion fast paths (unmatched groups, single literals, flag-sync contracts), LowerFirst invalid-byte handling, compiler case-directive parsing and the has_case_conversion flag, plus end-to-end substitution coverage. Collapse two provably-dead None arms preempted by the all-None early return. Deduplicate GNU testsuite bot comments.
- revert unrelated GnuComment workflow modifications - update README to note \u/\U replacement behavior - consolidate case escape mapping in compiler - factor out take_case helper and shorten comments in command.rs
626ce4f to
7fefe2b
Compare
|
GNU sed testsuite comparison: |
7fefe2b to
ecd940c
Compare
|
GNU sed testsuite comparison: |
ecd940c to
75865ff
Compare
|
GNU sed testsuite comparison: |
- Encapsulate ReplacementTemplate.has_case_conversion with a private field and public getter - Replace panic arm in render_parts with unreachable!() and remove artificial panic tests - Drop backslash for unrecognized escapes in replacement strings to match GNU sed - Disable GNU case-conversion escapes in replacement strings under POSIX mode - Add unit and integration regression tests for unrecognized escapes and POSIX mode
75865ff to
58b6c03
Compare
|
GNU sed testsuite comparison: |


Fixes #540.
Implements GNU
s///replacement case-conversion escapes:\Uuppercase until\L/\E,\Llowercase until\U/\E\u/\lone-shot next-character conversion\Eends persistent conversion and clears pending one-shotDetails:
ReplacementPart::{Upper,Lower,UpperFirst,LowerFirst,End}incommand.rs;compile_replacementparses them beforeparse_char_escapeso\u/\Uin replacements are case-conversion directives rather than\uXXXX/\UXXXXXXXXUnicode escapes (matches GNU sed behavior; README updated accordingly). Escaped delimiter still wins (e.g.sU...U\UU).append_with_caseapplies persistent + one-shot to literals,&,\1..\9. Empty matches leave one-shot pending (GNUs/(b?)-/\u\1x/gcarry), any produced char consumes it. Fresh state per substituted occurrence sogdoes not propagate.take_case(single, persistent)helper.Verified against GNU sed 4.9 including the issue table (
ABC DEF,ABc def) plus\U&\E,\U\1-\2, empty-groupgcases,\u\l/\l\ulast-wins,\U\l/\L\ucombos.Tests:
cargo test --lib385 passed,cargo test --test tests268 passed,cargo clippy --libclean,cargo fmt --checkclean.