From 5c98e4d68fd5c2f976ccaebb3f31d790232bb5ae Mon Sep 17 00:00:00 2001 From: Sylvestre Ledru Date: Sat, 26 Sep 2026 19:21:54 +0200 Subject: [PATCH] docs: move extensions and incompatibilities to docs/src/extensions.md Start an mdBook under docs/, laid out as in coreutils, and move the list of extensions and incompatibilities out of the README into it. The README keeps a pointer. A Documentation workflow builds the book with mdbook 0.5.0, the version the uutils website uses, whenever docs/ changes, and uploads it as an artifact. mdbook 0.5 rejects the `multilingual` field coreutils still carries (the website strips it before building), so the book leaves it out. --- .github/workflows/documentation.yml | 41 ++++++++++++++++ README.md | 69 +-------------------------- docs/.gitignore | 1 + docs/book.toml | 7 +++ docs/src/SUMMARY.md | 5 ++ docs/src/extensions.md | 73 +++++++++++++++++++++++++++++ docs/src/index.md | 10 ++++ 7 files changed, 139 insertions(+), 67 deletions(-) create mode 100644 .github/workflows/documentation.yml create mode 100644 docs/.gitignore create mode 100644 docs/book.toml create mode 100644 docs/src/SUMMARY.md create mode 100644 docs/src/extensions.md create mode 100644 docs/src/index.md diff --git a/.github/workflows/documentation.yml b/.github/workflows/documentation.yml new file mode 100644 index 00000000..0b672e2d --- /dev/null +++ b/.github/workflows/documentation.yml @@ -0,0 +1,41 @@ +name: Documentation + +on: + pull_request: + paths: + - 'docs/**' + - '.github/workflows/documentation.yml' + push: + branches: + - main + paths: + - 'docs/**' + - '.github/workflows/documentation.yml' + +permissions: + contents: read + +jobs: + build-doc: + name: Build the mdBook documentation + runs-on: ubuntu-latest + + steps: + - name: Checkout repository + uses: actions/checkout@v7 + with: + persist-credentials: false + + - name: Install mdbook + uses: taiki-e/install-action@v2 + with: + tool: mdbook@0.5.0 + + - name: Build documentation + run: mdbook build docs + + - name: Upload documentation + uses: actions/upload-artifact@v7 + with: + name: documentation + path: docs/book diff --git a/README.md b/README.md index ce60d17e..40ca9b83 100644 --- a/README.md +++ b/README.md @@ -88,73 +88,8 @@ cargo test ``` ## Extensions and incompatibilities -### Supported GNU extensions -* Command-line arguments can be specified in long (`--`) form. -* Spaces can precede a regular expression modifier. -* `I` can be used in as a synonym for the `i` (case insensitive) substitution - flag. -* `M` and `m` substitution flags allow multi-line matching. -* In addition to `\n`, other escape sequences (octal, hex, C) are supported - in the strings of the `y` command. - Under POSIX these yield undefined behavior. -* The `a`, `c`, and `i` commands do not require an initial backslash, - allow text to appear on the same line, and support escape sequences - in the specified text. -* The `a`, `i`, `=`, `l`, `q` and `r` commands support address range as an extension to POSIX. -* The substitution command replacement group `\0` is a synonym for &. -* An `F` command outputs the name of the file currently being processed. -* A `Q` command (optionally followed by an exit code) quits immediately. -* The `q` command can be optionally followed by an exit code. -* A `W` command writes to a file the pattern's first line. -* An `R` command reads one line at a time from a file. -* The `l` command can be optionally followed by the output width. -* The `--follow-symlinks` option for in-place editing. -* The `--sandbox` option that limits potentially destructive commands. -* Address 0 can be used to specify an address range that is already - active on line 1 and can finish with the specified regular expression. -* Address steps can be specified in the form of start~step and start,~step - ranges. -* Address 0 can be used in the `r` command to prepend a file. - -### Supported BSD and GNU extensions -* The second address in a range can be specified as a relative address with +N. -* In-place editing of file with the `-i` flag. - -### New extensions -* Unicode characters can be specified in regular expression pattern, replacement - and transliteration sequences using `\uXXXX` or `\UXXXXXXXX` sequences. - -### Incompatible extensions -The `-U` or `--uutil-extensions` option enables useful extensions or bug fixes -that aren't compatible with GNU sed or POSIX. - -* The `l` command lists Unicode characters using the `\uXXXX` and `\UXXXXXXXX` - escapes rather than as octal UTF-8 byte sequences. - -### Incompatibilities -* Similarly to GNU _sed_, input is processed as raw bytes or as valid UTF-8 - (this includes 7-bit ASCII) based on the locale as specified by the - `LC_ALL`, `LC_CTYPE`, and `LANG` environment variables, - with the default being byte processing. - However, in contrast with GNU _sed_, other locales (e.g. ISO-8859-1) - are not supported. If the input is in another code page or encoding - and requires locale-specific processing (e.g. ignore/map case, - character classes), consider converting it through UTF-8 to ensure - the correct handling of locale-specific regular expressions. - This _sed_ program can also handle arbitrary byte sequences - if no part of the input requires treating it as a Rust String. -* Back-references aren't supported when input is processed as bytes - (`LC_ALL=C`). -* The command will report an error and fail if duplicate labels are found - in the script. - This matches the BSD behavior. The GNU version accepts duplicate labels. -* The last line (`$`) address is interpreted as the last non-empty line of - the last file. If files specified in subsequent arguments until the last - one are empty, then the last line condition will never be triggered. - This behavior is consistent with the - [original implementation](https://github.com/dspinellis/unix-history-repo/blob/Research-V7/usr/src/cmd/sed/sed1.c#L665). -* Labels are parsed for alphanumeric characters. The BSD version parses them - until the end of the line, preventing ; to be used as a separator. +The GNU, BSD and new extensions _sed_ supports, and where it differs from GNU +_sed_, are listed in [docs/src/extensions.md](docs/src/extensions.md). ## GNU test suite compatibility diff --git a/docs/.gitignore b/docs/.gitignore new file mode 100644 index 00000000..7585238e --- /dev/null +++ b/docs/.gitignore @@ -0,0 +1 @@ +book diff --git a/docs/book.toml b/docs/book.toml new file mode 100644 index 00000000..32dc524d --- /dev/null +++ b/docs/book.toml @@ -0,0 +1,7 @@ +[book] +language = "en" +src = "src" +title = "uutils sed Documentation" + +[output.html] +git-repository-url = "https://github.com/uutils/sed/tree/main/docs/src" diff --git a/docs/src/SUMMARY.md b/docs/src/SUMMARY.md new file mode 100644 index 00000000..21ccb4b9 --- /dev/null +++ b/docs/src/SUMMARY.md @@ -0,0 +1,5 @@ +# Summary + +[Introduction](index.md) + +* [Extensions and incompatibilities](extensions.md) diff --git a/docs/src/extensions.md b/docs/src/extensions.md new file mode 100644 index 00000000..c81b4b53 --- /dev/null +++ b/docs/src/extensions.md @@ -0,0 +1,73 @@ +# Extensions and incompatibilities + +The main goal of the project is compatibility with GNU _sed_, but _sed_ also +supports features that GNU _sed_ does not, and differs from it in a few places. +Below is a list of these extensions and incompatibilities. + +## Supported GNU extensions +* Command-line arguments can be specified in long (`--`) form. +* Spaces can precede a regular expression modifier. +* `I` can be used in as a synonym for the `i` (case insensitive) substitution + flag. +* `M` and `m` substitution flags allow multi-line matching. +* In addition to `\n`, other escape sequences (octal, hex, C) are supported + in the strings of the `y` command. + Under POSIX these yield undefined behavior. +* The `a`, `c`, and `i` commands do not require an initial backslash, + allow text to appear on the same line, and support escape sequences + in the specified text. +* The `a`, `i`, `=`, `l`, `q` and `r` commands support address range as an extension to POSIX. +* The substitution command replacement group `\0` is a synonym for &. +* An `F` command outputs the name of the file currently being processed. +* A `Q` command (optionally followed by an exit code) quits immediately. +* The `q` command can be optionally followed by an exit code. +* A `W` command writes to a file the pattern's first line. +* An `R` command reads one line at a time from a file. +* The `l` command can be optionally followed by the output width. +* The `--follow-symlinks` option for in-place editing. +* The `--sandbox` option that limits potentially destructive commands. +* Address 0 can be used to specify an address range that is already + active on line 1 and can finish with the specified regular expression. +* Address steps can be specified in the form of start~step and start,~step + ranges. +* Address 0 can be used in the `r` command to prepend a file. + +## Supported BSD and GNU extensions +* The second address in a range can be specified as a relative address with +N. +* In-place editing of file with the `-i` flag. + +## New extensions +* Unicode characters can be specified in regular expression pattern, replacement + and transliteration sequences using `\uXXXX` or `\UXXXXXXXX` sequences. + +## Incompatible extensions +The `-U` or `--uutil-extensions` option enables useful extensions or bug fixes +that aren't compatible with GNU sed or POSIX. + +* The `l` command lists Unicode characters using the `\uXXXX` and `\UXXXXXXXX` + escapes rather than as octal UTF-8 byte sequences. + +## Incompatibilities +* Similarly to GNU _sed_, input is processed as raw bytes or as valid UTF-8 + (this includes 7-bit ASCII) based on the locale as specified by the + `LC_ALL`, `LC_CTYPE`, and `LANG` environment variables, + with the default being byte processing. + However, in contrast with GNU _sed_, other locales (e.g. ISO-8859-1) + are not supported. If the input is in another code page or encoding + and requires locale-specific processing (e.g. ignore/map case, + character classes), consider converting it through UTF-8 to ensure + the correct handling of locale-specific regular expressions. + This _sed_ program can also handle arbitrary byte sequences + if no part of the input requires treating it as a Rust String. +* Back-references aren't supported when input is processed as bytes + (`LC_ALL=C`). +* The command will report an error and fail if duplicate labels are found + in the script. + This matches the BSD behavior. The GNU version accepts duplicate labels. +* The last line (`$`) address is interpreted as the last non-empty line of + the last file. If files specified in subsequent arguments until the last + one are empty, then the last line condition will never be triggered. + This behavior is consistent with the + [original implementation](https://github.com/dspinellis/unix-history-repo/blob/Research-V7/usr/src/cmd/sed/sed1.c#L665). +* Labels are parsed for alphanumeric characters. The BSD version parses them + until the end of the line, preventing ; to be used as a separator. diff --git a/docs/src/index.md b/docs/src/index.md new file mode 100644 index 00000000..12622f2c --- /dev/null +++ b/docs/src/index.md @@ -0,0 +1,10 @@ +# uutils sed + +_sed_ is a Rust reimplementation of the +[sed utility](https://pubs.opengroup.org/onlinepubs/9799919799/utilities/sed.html) +with some [GNU sed](https://www.gnu.org/software/sed/manual/sed.html), +[FreeBSD sed](https://man.freebsd.org/cgi/man.cgi?sed(1)), +and other extensions. + +It is part of the [uutils](https://uutils.github.io/) project. The source code +is on [GitHub](https://github.com/uutils/sed).