Skip to content

shaik-abdul-thouhid/ezi-code

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

120 Commits
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

ezicode

A Unicode library for Zig. Three layers, stacked:

  • encoding — UTF-8, UTF-16, and UTF-32 codecs, plus BOM detection/stripping (encoding.bom).
  • transcoding — conversion between the three encoding forms, plus a chunked UTF-8 stream decoder.
  • unicode — character properties and text algorithms backed by the Unicode Character Database: normalization, casing, segmentation (grapheme / word / sentence / line), width, scripts, bidi, numeric, blocks, hangul, age.
  • collation — Default Unicode Collation Algorithm (UCA) using the DUCET, with sort key serialization for allocation-free comparison and database indexing.

It has no dependencies. The UCD tables are generated into Zig source and committed, so a normal build doesn't touch the network or the ucd/ inputs.

Status

Version 0.5.0-dev in main (latest release: v0.4.1). Pre-1.0 in the literal sense: the API is allowed to change. Tracks a recent Zig dev build (0.17.0-dev.657+2faf8debf minimum); it does not build against stable 0.16. If your toolchain isn't on a current master, this will not compile, and that is the intended trade-off until Zig 0.17 lands.

What works is well-tested. The unicode submodule includes exhaustive 0..=0x10FFFF sweeps and runs against the official UCD conformance vectors (GraphemeBreakTest.txt, WordBreakTest.txt, SentenceBreakTest.txt, LineBreakTest.txt, NormalizationTest.txt, CollationTest*.txt) under a build flag. The bidi algorithm went in with the rule-numbered adversarial test set you'd expect for UAX #9.

Installing

Via git ref (resolves the tag at fetch time):

zig fetch --save git+https://github.com/shaik-abdul-thouhid/ezi-code.git#v0.4.1

Or via plain HTTP tarball (pins the content hash in build.zig.zon):

zig fetch --save https://github.com/shaik-abdul-thouhid/ezi-code/archive/refs/tags/v0.4.1.tar.gz

Then in build.zig:

const ezi_code = b.dependency("ezi_code", .{
    .target = target,
    .optimize = optimize,
});
exe.root_module.addImport("ezi_code", ezi_code.module("ezi_code"));

Individual modules (encoding, transcoding, unicode, utils) are also exported, so you can depend on just the layer you need.

Tracking main (nightly / unreleased features)

To try unreleased work — features staged under ## [Unreleased] in CHANGELOG.md before a version is cut — fetch the main branch ref instead of a tag:

zig fetch --save "git+https://github.com/shaik-abdul-thouhid/ezi-code.git#main"

main is pre-release and may change without notice, so pin a tag for anything you actually depend on. Re-run the same command to advance to the newest main commit (delete the dependency's hash in build.zig.zon first so the fetch re-resolves).

Quick look

const ezi = @import("ezi_code");

// Decode and iterate UTF-8 scalars.
var scalar_count: usize = 0;
const view = try ezi.utf8.initUTF8View("héllo, мир", &scalar_count);
var it = view.iter();
while (it.next()) |cp| {
    // cp: u21
    _ = cp;
}

// Where, not just whether: byte offset of the first malformed sequence.
if (ezi.utf8.invalidIndex(input_bytes)) |at| return reportInvalidAt(at);

// Detect and strip a leading byte-order mark.
const body = ezi.bom.strip(file_bytes);

// Convert between encoding forms.
const utf16 = try ezi.transcoding.utf8ToUtf16(allocator, "Καλημέρα");
defer allocator.free(utf16);

// Normalize ([]const CodePoint in, []const CodePoint out).
const nfc = try ezi.unicode.nfc(allocator, &.{ 'c', 'a', 'f', 'e', 0x0301 });
defer allocator.free(nfc);   // { 'c', 'a', 'f', 0xE9 }

// Case-fold a UTF-8 string for caseless matching (ß → "ss").
const folded = try ezi.unicode.casing.foldFullUtf8Alloc(allocator, "Straße");
defer allocator.free(folded);   // "strasse"

// Caseless search without allocating: "STRASSE" found inside "…Straße…".
const at = try ezi.unicode.casing.indexOfFold(.full, "die Straße hier", "STRASSE");
_ = at; // 4

// Titlecase per the Unicode default algorithm (UAX #29 words).
const titled = try ezi.unicode.casing.titlecaseUtf8Alloc(allocator, "ΜΕΓΑΣ ΣΟΦΟΣ");
defer allocator.free(titled);   // "Μεγας Σοφος" (final sigma handled)

// Measure display width on a monospace grid (wide CJK counts as 2).
const cols = try ezi.unicode.width.stringWidth("コード");  // 6

// Bidi: resolve embedding levels and get the visual order of a line.
var paragraph = try ezi.unicode.bidi.resolveParagraph(allocator, code_points, .auto);
defer paragraph.deinit();
const visual_order = try paragraph.reorderVisual(allocator);
defer allocator.free(visual_order);

// Collate: compare two strings per the Unicode Collation Algorithm.
var collator = ezi.collation.Collator.init(.{});
const order = try collator.compareUtf8(allocator, "café", "cafe");
_ = order; // .gt

// Build a sort key once, then serialize to bytes for allocation-free comparison.
var key: ezi.collation.Key = .{};
defer key.deinit(allocator);
try collator.buildKey(allocator, code_points, &key);
const sort_bytes = try key.serializeAlloc(allocator, collator.options);
defer allocator.free(sort_bytes);
// std.mem.order(u8, sort_bytes_a, sort_bytes_b) == collator.compareKeys(key_a, key_b)

The per-module READMEs in src/encoding/, src/transcoding/, src/unicode/, and src/collation/ document the full surface. Read them before reaching for the top-level types — the interesting design is at that level, not in the facade.

What's in the unicode module

Submodule What it does
properties General Category, Bidi Class, CCC, Derived Core Properties, PropList
casing Simple / full / special casing (Turkic etc.), case folding
normalization NFC, NFD, NFKC, NFKD, Quick_Check, streaming Normalizer
segmentation Grapheme / word / sentence / line breaking (UAX #14, UAX #29) + iterators
emoji UTS #51 emoji properties (Emoji, Emoji_Presentation, Extended_Pictographic, …) + range tables
width East Asian Width
scripts Script and Script_Extensions
bidi UAX #9: mirroring, paired brackets, and the full reordering algorithm
numeric Numeric_Type and Numeric_Value
blocks Block membership and names
hangul Hangul_Syllable_Type plus algorithmic Hangul composition
age Derived_Age (Unicode version a code point was assigned in)

All lookups are deduplicated two-level page tables — two array indexes per query — so the table cost stays small even though every submodule covers the whole code space.

Unicode version and regenerating the tables

The committed tables track the UCD files in ucd/ (UnicodeData.txt, DerivedCoreProperties.txt, BidiMirroring.txt, BidiBrackets.txt, the break-property and -test files, etc.). To bump the Unicode version:

  1. Replace the relevant files under ucd/.
  2. Run zig build generate. This rebuilds the deduplicated tables under each submodule's generated/ directory.
  3. Re-run the conformance suite (below).

For day-to-day work, you don't need this — the generated files are checked in.

Building, testing, benchmarking

# Build the placeholder executable.
zig build

# Run all tests (Debug). Note: the unicode sweeps are slow in Debug.
zig build test

# Run a specific suite. Selectors: all, encoding, transcoding, unicode, collation, utils, conformance.
zig build test -Dinclude-test=unicode -Doptimize=ReleaseSafe

# Run the UCD conformance vectors.
zig build test -Dinclude-test=conformance -Doptimize=ReleaseSafe

# Regenerate Unicode tables from ucd/.
zig build generate

# Run benchmarks. Defaults to ReleaseFast for the library and the driver.
zig build bench
zig build bench -- --list                 # list registered modules
zig build bench -- encoding/utf8          # run one
zig build bench -- --size=524288 unicode  # custom corpus size

The bench driver reports mean of 7 runs ± stddev with throughput and tracked allocator memory, over three corpora (ASCII, multilingual, pathological). Module list is in bench/main.zig.

Generating documentation

zig build docs

This emits the Zig autodoc bundle into the repo-root docs/ directory: index.html, main.js, main.wasm, and sources.tar (the source the viewer loads on demand). Every public declaration carries a @stable-since: vX.Y.Z marker in its doc comment, so you can see when each API entered the stable surface.

Opening the docs

The viewer is a WebAssembly app that fetches sources.tar at runtime, so most browsers refuse to load it from a file:// path (CORS). Serve docs/ over HTTP instead:

zig build docs                       # (re)generate ./docs
cd docs && python3 -m http.server 8000
# then open http://localhost:8000 in a browser

Any static file server works — pick whatever you have:

npx http-server docs -p 8000         # Node
php -S localhost:8000 -t docs        # PHP
ruby -run -e httpd docs -p 8000      # Ruby

Some Chromium builds will still open docs/index.html directly from disk; if the page loads but shows no declarations, switch to the HTTP-server method above.

Layout

src/
  encoding/        UTF-8, UTF-16, UTF-32 codecs, BOM utilities + per-module README
  transcoding/     Cross-encoding converters and UTF8Stream + per-module README
  unicode/         All UCD-backed properties and algorithms + per-module README
    age/  bidi/  blocks/  casing/  emoji/  hangul/  normalization/
    numeric/  properties/  scripts/  segmentation/  width/
    tests/         UCD conformance test runners
  collation/       UCA/DUCET collation + sort key serialization + per-module README
    generated/      Generated DUCET tables
    tests/          CollationTest conformance + sort key serialization tests
  utils/           Internal helpers (search, slices). Not part of the public API.
bench/             Benchmark driver, framework, corpora, per-module suites
ucd/               Raw UCD inputs (only needed for `zig build generate`)
licences/          Upstream licenses for bundled third-party code and data

Design notes worth knowing before reading the code

  • Decode paths come in three flavours everywhere they exist: strict (full validation, fine-grained errors), unchecked (assume valid, skip checks), and lossy (replace malformed runs with U+FFFD, never error — there is no panic anywhere on a lossy path). Error sets are per-codec; "overlong", "surrogate", and "too large" are different failures because callers want to treat them differently.
  • Unchecked means one thing everywhere: the caller guarantees the documented preconditions; violations are asserted (safety-checked — they trap in Debug/ReleaseSafe and are undefined in ReleaseFast/ReleaseSmall); unchecked functions never return errors and never panic.
  • CodePoint (u21) is a contract: a value of this type is presumed to be a valid Unicode scalar (in range, not a surrogate). APIs that produce one uphold the contract; APIs that accept one — the []const CodePoint variants of casing, search, normalization, collation, and the encodeCodePoints* bulk encoders — rely on it and skip decoding and validation entirely. Pass already-decoded text through the CodePoint variants and you pay validation exactly once, at the boundary.
  • Validation answers where, not just whether: invalidIndex on every codec reports the offset of the first malformed sequence, and utf8.StreamingValidator does the same across arbitrarily-chunked input with no buffering.
  • The codec layer doesn't allocate. Where you need owned output, there is an explicit allocating variant or a buffer variant that writes into your memory.
  • Backward UTF-8 traversal is a first-class operation (codePointLenReverse, decodeCodePointReverseUnchecked, etc.) — necessary for cursor-style editing without scanning from the start.
  • The transcoders check the maximum possible expansion against usize overflow before any allocation. A hostile length is rejected before a single source unit is read.
  • The bidi algorithm follows the spec's rule numbering literally. If you're reading the code, keep UAX #9 open next to it.
  • utils/ is internal. It's exported by build.zig so submodules can share it, not because it's part of the API.

License

MIT for the source. See LICENSE.

Two third-party components ship with their own terms in licences/:

  • Björn Höhrmann's "Flexible and Economical UTF-8 Decoder" (MIT) — the DFA used by the UTF-8 codec.
  • Unicode Character Database data files (Unicode License V3) — the inputs in ucd/ and, transitively, the generated tables derived from them.

About

A Zig native Unicode (v17) engine: UTF-8/16/32, normalization, segmentation, unicode properties, bidi, collation (DUCET)

Topics

Resources

License

Stars

1 star

Watchers

0 watching

Forks

Contributors

Languages