Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
41 changes: 41 additions & 0 deletions .github/workflows/documentation.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,41 @@
name: Documentation

on:
pull_request:
paths:
- 'docs/**'
- '.github/workflows/documentation.yml'
push:
branches:
- main
paths:
- 'docs/**'
- '.github/workflows/documentation.yml'

permissions:
contents: read

jobs:
build-doc:
name: Build the mdBook documentation
runs-on: ubuntu-latest

steps:
- name: Checkout repository
uses: actions/checkout@v7
with:
persist-credentials: false

- name: Install mdbook
uses: taiki-e/install-action@v2
with:
tool: mdbook@0.5.0

- name: Build documentation
run: mdbook build docs

- name: Upload documentation
uses: actions/upload-artifact@v7
with:
name: documentation
path: docs/book
69 changes: 2 additions & 67 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -88,73 +88,8 @@ cargo test
```

## Extensions and incompatibilities
### Supported GNU extensions
* Command-line arguments can be specified in long (`--`) form.
* Spaces can precede a regular expression modifier.
* `I` can be used in as a synonym for the `i` (case insensitive) substitution
flag.
* `M` and `m` substitution flags allow multi-line matching.
* In addition to `\n`, other escape sequences (octal, hex, C) are supported
in the strings of the `y` command.
Under POSIX these yield undefined behavior.
* The `a`, `c`, and `i` commands do not require an initial backslash,
allow text to appear on the same line, and support escape sequences
in the specified text.
* The `a`, `i`, `=`, `l`, `q` and `r` commands support address range as an extension to POSIX.
* The substitution command replacement group `\0` is a synonym for &.
* An `F` command outputs the name of the file currently being processed.
* A `Q` command (optionally followed by an exit code) quits immediately.
* The `q` command can be optionally followed by an exit code.
* A `W` command writes to a file the pattern's first line.
* An `R` command reads one line at a time from a file.
* The `l` command can be optionally followed by the output width.
* The `--follow-symlinks` option for in-place editing.
* The `--sandbox` option that limits potentially destructive commands.
* Address 0 can be used to specify an address range that is already
active on line 1 and can finish with the specified regular expression.
* Address steps can be specified in the form of start~step and start,~step
ranges.
* Address 0 can be used in the `r` command to prepend a file.

### Supported BSD and GNU extensions
* The second address in a range can be specified as a relative address with +N.
* In-place editing of file with the `-i` flag.

### New extensions
* Unicode characters can be specified in regular expression pattern, replacement
and transliteration sequences using `\uXXXX` or `\UXXXXXXXX` sequences.

### Incompatible extensions
The `-U` or `--uutil-extensions` option enables useful extensions or bug fixes
that aren't compatible with GNU sed or POSIX.

* The `l` command lists Unicode characters using the `\uXXXX` and `\UXXXXXXXX`
escapes rather than as octal UTF-8 byte sequences.

### Incompatibilities
* Similarly to GNU _sed_, input is processed as raw bytes or as valid UTF-8
(this includes 7-bit ASCII) based on the locale as specified by the
`LC_ALL`, `LC_CTYPE`, and `LANG` environment variables,
with the default being byte processing.
However, in contrast with GNU _sed_, other locales (e.g. ISO-8859-1)
are not supported. If the input is in another code page or encoding
and requires locale-specific processing (e.g. ignore/map case,
character classes), consider converting it through UTF-8 to ensure
the correct handling of locale-specific regular expressions.
This _sed_ program can also handle arbitrary byte sequences
if no part of the input requires treating it as a Rust String.
* Back-references aren't supported when input is processed as bytes
(`LC_ALL=C`).
* The command will report an error and fail if duplicate labels are found
in the script.
This matches the BSD behavior. The GNU version accepts duplicate labels.
* The last line (`$`) address is interpreted as the last non-empty line of
the last file. If files specified in subsequent arguments until the last
one are empty, then the last line condition will never be triggered.
This behavior is consistent with the
[original implementation](https://github.com/dspinellis/unix-history-repo/blob/Research-V7/usr/src/cmd/sed/sed1.c#L665).
* Labels are parsed for alphanumeric characters. The BSD version parses them
until the end of the line, preventing ; to be used as a separator.
The GNU, BSD and new extensions _sed_ supports, and where it differs from GNU
_sed_, are listed in [docs/src/extensions.md](docs/src/extensions.md).

## GNU test suite compatibility

Expand Down
1 change: 1 addition & 0 deletions docs/.gitignore
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
book
7 changes: 7 additions & 0 deletions docs/book.toml
Original file line number Diff line number Diff line change
@@ -0,0 +1,7 @@
[book]
language = "en"
src = "src"
title = "uutils sed Documentation"

[output.html]
git-repository-url = "https://github.com/uutils/sed/tree/main/docs/src"
5 changes: 5 additions & 0 deletions docs/src/SUMMARY.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
# Summary

[Introduction](index.md)

* [Extensions and incompatibilities](extensions.md)
73 changes: 73 additions & 0 deletions docs/src/extensions.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,73 @@
# Extensions and incompatibilities

The main goal of the project is compatibility with GNU _sed_, but _sed_ also
supports features that GNU _sed_ does not, and differs from it in a few places.
Below is a list of these extensions and incompatibilities.

## Supported GNU extensions
* Command-line arguments can be specified in long (`--`) form.
* Spaces can precede a regular expression modifier.
* `I` can be used in as a synonym for the `i` (case insensitive) substitution
flag.
* `M` and `m` substitution flags allow multi-line matching.
* In addition to `\n`, other escape sequences (octal, hex, C) are supported
in the strings of the `y` command.
Under POSIX these yield undefined behavior.
* The `a`, `c`, and `i` commands do not require an initial backslash,
allow text to appear on the same line, and support escape sequences
in the specified text.
* The `a`, `i`, `=`, `l`, `q` and `r` commands support address range as an extension to POSIX.
* The substitution command replacement group `\0` is a synonym for &.
* An `F` command outputs the name of the file currently being processed.
* A `Q` command (optionally followed by an exit code) quits immediately.
* The `q` command can be optionally followed by an exit code.
* A `W` command writes to a file the pattern's first line.
* An `R` command reads one line at a time from a file.
* The `l` command can be optionally followed by the output width.
* The `--follow-symlinks` option for in-place editing.
* The `--sandbox` option that limits potentially destructive commands.
* Address 0 can be used to specify an address range that is already
active on line 1 and can finish with the specified regular expression.
* Address steps can be specified in the form of start~step and start,~step
ranges.
* Address 0 can be used in the `r` command to prepend a file.

## Supported BSD and GNU extensions
* The second address in a range can be specified as a relative address with +N.
* In-place editing of file with the `-i` flag.

## New extensions
* Unicode characters can be specified in regular expression pattern, replacement
and transliteration sequences using `\uXXXX` or `\UXXXXXXXX` sequences.

## Incompatible extensions
The `-U` or `--uutil-extensions` option enables useful extensions or bug fixes
that aren't compatible with GNU sed or POSIX.

* The `l` command lists Unicode characters using the `\uXXXX` and `\UXXXXXXXX`
escapes rather than as octal UTF-8 byte sequences.

## Incompatibilities
* Similarly to GNU _sed_, input is processed as raw bytes or as valid UTF-8
(this includes 7-bit ASCII) based on the locale as specified by the
`LC_ALL`, `LC_CTYPE`, and `LANG` environment variables,
with the default being byte processing.
However, in contrast with GNU _sed_, other locales (e.g. ISO-8859-1)
are not supported. If the input is in another code page or encoding
and requires locale-specific processing (e.g. ignore/map case,
character classes), consider converting it through UTF-8 to ensure
the correct handling of locale-specific regular expressions.
This _sed_ program can also handle arbitrary byte sequences
if no part of the input requires treating it as a Rust String.
* Back-references aren't supported when input is processed as bytes
(`LC_ALL=C`).
* The command will report an error and fail if duplicate labels are found
in the script.
This matches the BSD behavior. The GNU version accepts duplicate labels.
* The last line (`$`) address is interpreted as the last non-empty line of
the last file. If files specified in subsequent arguments until the last
one are empty, then the last line condition will never be triggered.
This behavior is consistent with the
[original implementation](https://github.com/dspinellis/unix-history-repo/blob/Research-V7/usr/src/cmd/sed/sed1.c#L665).
* Labels are parsed for alphanumeric characters. The BSD version parses them
until the end of the line, preventing ; to be used as a separator.
10 changes: 10 additions & 0 deletions docs/src/index.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,10 @@
# uutils sed

_sed_ is a Rust reimplementation of the
[sed utility](https://pubs.opengroup.org/onlinepubs/9799919799/utilities/sed.html)
with some [GNU sed](https://www.gnu.org/software/sed/manual/sed.html),
[FreeBSD sed](https://man.freebsd.org/cgi/man.cgi?sed(1)),
and other extensions.

It is part of the [uutils](https://uutils.github.io/) project. The source code
is on [GitHub](https://github.com/uutils/sed).
Loading