library --workers N: the--sweepcandidate pass runs through N concurrent epubcheck workers (default 1: serial, byte-for-byte unchanged). epubcheck runs in a subprocess that releases the GIL, so threads parallelize the oracle honestly; this is the pass that measured ~4.4 s/book on the 2026-08-27 full-library walk (5,070 EPUBs ≈ 6 hours serial). Books are checked in windows of N consumed in input order, so the candidate set, the before-measurements, and every emitted line match the serial sweep, and--limitstays lazy within one window of overshoot. The repair phase stays serial on purpose: that is where the shared workdir and the atomic-replacement contract live. Without--sweep,--workersis a no-op (with a note); counts below 1 are a usage error.
The acquisition pathway's EPUB steps are now first-class CLI surface. Nothing here adds a repair class or changes an existing mode; the new verbs compose what already ships, in the phase-1 and phase-3 skills' documented order.
audit --json FILE: the audit subcommand writes machine-readable per-file verdicts in thelibrary --jsonshape (mode/root/summary plus one record per book), from all three modes: directory, library, and single-book. Records carry a status (clean/problem/error) and per-analyzer verdicts (problem/status/details); the always-on archive/spine verdicts appear OK when they were silent, emptytext is omitted from a record when the archive verdict owns the body-text story, and a scan error becomes its own record shape.--jsonwith--idaccepts exactly one book id, because each single-book run writes the file wholesale. This is the contract downstream consumers of the phase-1 EPUB slice read.run phase1 DIR [--json OUT] [--apply-lossy] [--backup DIR]: the pre-import vetting slice over a directory of loose files, in the phase-1 skill's documented order: the audit battery (corruption sweep, content battery, monolithic), then one fused repair sweep (epubcheck, watermark detection, repairability; the skill's two sweeps collapse into one epubcheck pass, with watermark hits read out of the per-book fix summary). Read-only until--apply-lossy, which IS the recorded lossy-strip consent;--backup DIRpasses through (keep backups outside the vetted directory). Exit codes 0/1/2 per the library contract; a consent question alone is not trouble.run phase3 --ids CSV [--json OUT]: the post-import scoped repair step, run from the library directory: exactlylibrary --id --sweep --only all --apply --all --install-to-calibreplus a pre/post epubcheck summary over the swept books. The phase-3 skill's step-10 scope warning is mechanized: an unscoped library-wide call is refused with exit 2, never run by accident. The verb drives the shipped library runner through the real parser, so its flags cannot drift from the subcommand it wraps.- Non-interactive contract: both run verbs take
--non-interactiveand never prompt (nothing in bindery prompts); open questions surface in the JSON asdecisions_neededentries: theapply_lossyconsent question in a read-only phase1,manual_repair/investigatefor books left partial or unreadable in phase3. - Skill sync: the phase-1 and phase-3 skills name the run verbs and the machine-readable contracts (the skill files live in the library directory, outside this repo).
- Dependencies: vir-tui 2.3.0 folded into the lock; the pyproject floor
stays
vir-tui>=2.2.0(additive bump at this release). - Docs and hygiene: CLAUDE.md's last surviving
calibredb add_formatline retired (it contradicted the v0.24.0 subprocess removal and carried a-applytypo); README's--sweepbullet and the pyproject description no longer advertise the validation daemon's speedup, which never engages because the daemon's classpath omits epubcheck'slib/dependencies (the fix is a recorded gate-behavior decision, not a drive-by);venv/and.venv_ci/are gitignored. Test count: 277 → 294.
The 2026-09-05 full-library sweep (testing_facility/top500candidates/REPORT.md)
catalogued the top recurring codes: RSC-005 (164,302), RSC-020 (8,431), PKG-010
(4,112), RSC-007 (3,219), RSC-012 (1,940), HTM-025 (800). Bulk RSC-005 repair stays
rejected (the spec non-goal stands); everything else gained a deterministic,
epubcheck-gated, opt-in repair. The structural-repair set is now twelve; --all
includes all three.
--prune-missing-resources(PKG-010/RSC-007) removes references to files the archive does not contain: dead<link>elements (dead_links_pruned), anchors' href to absent files with the text kept (missing_file_hrefs_stripped), absent<img>sources replaced by their escaped alt text when they carry one (missing_imgs_unwrapped) or dropped when they do not (missing_imgs_pruned), and orphaned non-spine OPF manifest items (manifest_items_pruned). Spine documents are never pruned: a missing spine document is a damaged fragment for the spine-integrity report, not a pruning candidate.--strip-broken-anchors(RSC-020/RSC-012, plus the HTM-025 scheme half) strips href attributes that cannot resolve while keeping the anchor text byte-for-byte: a#fragmentthe target document does not define (broken_fragment_hrefs_stripped; the anchor keeps its text, NCX navTargets fall back to the document target so chapter navigation survives, and the fragment is never re-pointed at a guessed sibling document), and non-resolvable URI schemes (nonfile_scheme_hrefs_stripped; fixed resolvable set: http, https, mailto). NCX counter:ncx_fragments_stripped. The id set for fragment checks is built from the documents as they will look when the pass runs (core transforms,--reserialize,--fix-id-colonsapplied), so an id rename can never make a valid reference look broken.--encode-url-spaces(the Death Masks shape: 67 RSC-020 "not a valid URL") percent-encodes raw spaces in src/href attribute values across the package — OPF manifest href, NCX content src, content documents (counterurl_spaces_encoded). Scope is fixed to the space character; archive entry names are untouched.- HTM-025 ampersand half closed by test:
escape_bare_ampalways covered attribute values (transforms run on raw text); the URL-attribute case is now pinned by an explicit test instead of an assumption.
Fixture validation (all /tmp copies, gate-accepted, zero regressions):
| Book | Before | After | Fixes |
|---|---|---|---|
| Swann's Way | 0f/31e/1w | 0f/2e/0w | dead_links_pruned:29, nonfile_scheme_hrefs_stripped:1 |
| Sodom and Gomorrah | 0f/53e/0w | 0f/42e/0w | ncx_fragments_stripped:9, broken_fragment_hrefs_stripped:2 |
| Death Masks | 0f/68e/34w | 0f/1e/34w | url_spaces_encoded:67 |
Every surviving finding is RSC-005 (the non-goal) or the Death Masks entry-name warnings (see below). The full label-vs-reality findings postscript lives in roadmap.md Phase 12; headline: the staged "PKG-010" book has no missing resources at all — its 34 PKG-010s are "file name contains spaces" warnings about the zip entry names, which no shipped flag touches. Renaming entries and rewriting every reference is a possible future flag.
scripts/FastSweep.javanow merges the audit (CSVfatals,errors,warnings,path, the formatbindery library --auditreads) and extract (path ||| CODE,...) prototypes behind--mode=audit/--mode=extract, keeping the stdin +parallelStream()shape.- The prototype's extract mode never worked:
CheckingReportbuffers every message and writes nothing to its writer untilgenerate()— andgenerate()NPEs without a precedinginitialize(). Every code list in the prototype'sraw_errors*.txtruns came back empty. The merged harness follows the CLI's own lifecycle (initialize→ validate →generate) and extracts codes from the JSON. The committed REPORT.md's per-book code attributions should be treated as unverified until regenerated with the fixed harness. scripts/fast_sweep.py(stdlib-only) is the wrapper: locates the epubcheck jar (--jar/EPUBCHECK_JAR/ parsed from the launcher), compiles the harness once with--release 25(cached on mtime), scans a directory or a path list, and--summaryaggregates extract output into the per-code REPORT-style table. Measured on the 14 staged fixtures: ~4.5 s across all cores vs ~42 s sequential.- The audit CSV feeds
bindery library --auditcleanly (smoke-tested for candidate selection both ways).
- epubcheck 5.3.0 surfaces disagree on counts: the human listing prints raw message
totals while the
--jsonchecker block prints deduplicated ones (Swann's Way: 31 vs 2 errors). bindery's gate reads the JSON surface consistently for both measurements, so gate decisions are unaffected. FastSweep's audit mode carries the raw CheckingReport counts; that CSV is used for candidate selection only. - Related, unchanged on purpose: validate.py's persistent daemon launches
java -cp ".:epubcheck.jar"without epubcheck'slib/dependencies, so it always fails its first check and silently falls back to the subprocess JSON path on this machine. Fixing the classpath would flip gate measurements from deduplicated to raw counts — a gate-behavior change that gets its own decision. - Tests 260 → 277: transform units (encoding, protection, idempotency, HTM-025), repair-level opt-in/counter/output tests for all three flags, the structural default-off case list, and the CLI wiring tuples.
- The element downgrade. The second half of the option-B verdict:
--downgrade-epub3-tagsrewrites EPUB3/HTML5 semantic elements to their EPUB2 equivalents —figureandsectionbecomediv,figcaptionbecomesp(counterepub3_tags_downgraded). Existing classes are kept and the semantic name is appended (<figure class="photo">becomes<div class="photo figure">), so class-selector stylesheets keep working. - CSS protection, generalized.
css_protected_tags/style_block_tagstake the tag set as a parameter now; a book whose stylesheet stylesfigure { ... }(element selector) has that name protected book-wide and the tags left untouched — preservation wins over compliance, and the book keeps its RSC-005 findings. The illegal-tags flow is behaviorally unchanged (same helpers, same defaults). - The structural-repair set is nine;
--allincludes the flag; the epubcheck gate applies as with every fix. - Tests 254 → 260: the three mappings, class preservation (both quote styles), the self-closing form, the protected-name path, and the end-to-end run with and without a styling stylesheet.
- The attribute scrub. Brandon picked option B on the NEW-AUDIT brief:
the low-risk attribute half ships first, the element downgrade separately
with its CSS-protection work.
--strip-epub3-attrsremoves the three EPUB3-only attributes epubcheck rejects on an EPUB2 package —page-progression-direction,epub:type,aria-label(counterepub3_attrs_stripped), in content documents and the OPF alike. - Fixed set, documented. The transform matches exactly those three names
and nothing else:
type="text"and the wider aria family survive, so a reader-legitimate attribute can never be swept up by accident. Extending the set requires a named epubcheck finding, per the roadmap note. - The structural-repair set is eight now;
--allincludes the flag; the epubcheck gate applies as with every fix. - Tests 250 → 254: the three attributes in both quote styles, the lookalike-survival guarantee, and the off-by-default / on / end-to-end run over an OPF and a content document.
--fix-page-mapnormalizes the two page-map defects epubcheck rejects on older HarperCollins / Anna's Archive conversions: the non-standardpage-map="page-map"attribute on the OPF<spine>(counterpage_map_stripped), and the classless NCX<pageList>that trips RSC-005 'missing required attribute "class"' (class="pages"injected, counterpagelist_class_added; a pageList that already carries a class is untouched).- Gated as a structural repair: opt-in, never part of the well-formedness
core, included by
--all, and epubcheck-gated like every other fix. The structural-repair set is seven now. - Tests 244 → 250: both quote styles for the spine attribute, a spine
without the attribute, injected/self-closing/already-classed pageList
forms, and the off-by-default / on / end-to-end zip render through
repair_epub.
bindery library --install-to-calibrewrites through cquarry now. Thecalibredb_replaceshell-out (acalibredb add_formatsubprocess) is replaced byinstall_format(): the repaired EPUB is placed atomically at the catalogued path — same directory, samedata.name, so Calibre's layout never changes — and thedatarow is re-registered throughcquarry.write.WritableCalibreDB(remove_format+add_formatin onebatch()transaction, becauseadd_formatrefuses case-insensitive duplicates). Two things improve beyond the crash fix: the storeduncompressed_sizebecomes the repaired file's true size (the CLI path updated it, but the row's truthfulness now holds for the fresh-format case too), and the book lands inmetadata_dirtied, so Calibre regenerates the sidecar.opfand re-pushes metadata to wireless readers — the raw subprocess never queued anything.- Fresh-format case. A book with no catalogued EPUB row gets the repaired
file placed under the repaired file's stem and a row registered fresh (the
resolver cannot see such books — its id map IS the format rows — so the
legacy
(id)guess plusCALIBRE_DBPATH, the old calibredb library contract, carries the id). - Failure posture. A database failure (locked, missing book) degrades to the atomic in-place save with a warning instead of losing the repair or crashing a multi-hour sweep; when the file was already placed, the warning says the row may be stale.
- Dependency.
uv.lockmoves to cquarry 1.9.0 (f22bbe7) in the same release, closing the lock-refresh note recorded at the 1.9 sync. - Tests 240 → 244: the routing battery was rewritten against the real write
module (id-wins-over-directory, row follow-through + queueing, guess via
CALIBRE_DBPATH, degrade paths for a wrong guess / missing library / database failure, the fresh-format placement).
- Fixed:
bindery library --install-to-calibrecrashed swapping repaired files back into the library. Thecalibredb_replacecall passed--replacetocalibredb add_format, but that flag does not exist — replacement is add_format's default behavior (--dont-replaceopts out) — so every install died withcalibredb: error: no such option: --replaceand aCalledProcessError(hit on a real repair pass, 2026-08-31). The subprocess call is now the plain four-argument form. The routing test that pinned the bad flag flips to reject it, and a dedicated regression test pins the exact command shape. spec.md and README no longer advertise the nonexistent flag. Phase 11 (roadmap) later retires the calibredb subprocess entirely in favor ofcquarry.write.WritableCalibreDB.add_format.
library --id <ids>: Added comma-separated book-id scoping for the sweep, mirroringaudit --id. Resolves EPUBs through cquarry'sget_format_path()viaCalibreIdResolver.audit --idcomma lists: The--idflag now supports a comma-separated list of IDs.- Spine-integrity reporting: Both audit and repair reports gain a field counting manifest/NCX references whose target files are absent from the archive, with a classification:
convention(consecutive chapter span despite absent files, e.g. bloated ToCs) orfragment(broken span).
bindery auditnames the corruption. The audit's single decompression pass now fully reads EVERY archive entry (CRC + real decompression, not just the central directory's word) before the text analysis runs, and a corrupt entry is reported as its own CORRUPT verdict —corrupt:N (first: <entry>)— in every mode, tagged like any other finding. A CRC-broken entry decompresses to nothing, and emptytext calling that EMPTY was the right alarm for the wrong disease: the re-source advice that follows from EMPTY ("content-less stub") mislabels a file whose problem is a damaged archive, not missing content. Emptytext now steps aside for corrupt books (classify returns CORRUPT; its buckets exclude them).library --sweep'sunreadablebucket splits by disease:not_a_zip/truncated/encrypted/corrupt_entry, so a CRC-damaged download is distinguishable from a DRM'd or truncated one without leaving the sweep (zipfile's message already names the broken entry; the summary prints per-reason counts).- The phase-1 skill's hand-rolled
zipfilecorruption sweep is retired — §2 now names the built-in verdict.test_bad_crc_entry_reads_emptyflipped to assert the corruption verdict instead of pinning the old empty-text reading.
bindery audit monolithicjoins the analyzer set (andall). One spine document at or above 300,000 characters flags — a book epubcheck and every other analyzer pass that still will not render on real readers (the motivating case: a ~30M-char dictionary EPUB, invisible to emptytext because it measures whole-book volume and to epubcheck because it never sees renderer memory limits).--max-doc-chars Nmoves the floor; the report carriesmax_doc_charsplus the offending document's href. Under-threshold books stay silent, like THIN. Reuses the shared visible-texts cache — no second decompression pass.- Wiring:
--id,--tag, directory and library modes all compose unchanged; flagged books tag like every other analyzer's. Phase 7's version-pin half shipped in v0.19.2. - Docs: spec's audit section, README, and the phase-1 skill's
"Chars per content document" step (no tool; run inline →
bindery audit monolithic).
CalibreIdResolver._loadoverformat_path_index(). The path→id map is now cquarry's canonical index (onedata ⋈ booksquery built exactly likeget_format_path()) instead of a per-bookget_format_path()loop — same construction, N queries collapsed to one. Keys are re-normalized the resolver's historical way (Path.resolve().lower()) so symlinked library directories and case-insensitive matching behave exactly as before. The(123)directory-name regex remains the documented no-catalog fallback. Theaudit.pyraw-join swap was waived at the Phase 9 design session (its three targeted joins are cheaper than whole-library hydration and bindery-cli wants per-book maps, not row dicts); do not "improve" it into a regression.- Dependency: requires cquarry ≥ 1.8 (
format_path_index). Theuv.lockstill pins the pre-1.8@maincommit until cquarry's 1.8.0 commits are pushed; after that, oneuv lock --upgrade-package cquarrylands the lock on current main. CI installs@mainfresh and self-corrects the same way.
- Fixed:
bindery audit --idwas unreachable from the CLI. v0.19.0 documented--idand shippedrun_single, but only the module's own argparse main (python -m bindery.audit) registered the flag; thebinderyconsole script's audit subparser never did, so the documented invocation died with "unrecognized arguments". The subparser now carries--id BOOK_IDandrun_audit_cmdroutes it torun_single(rejecting a directory argument with the module main's error text). Regression tests mirror the v0.18.1--tagwiring battery, which this bug duplicated: the earlier one shipped the write path unwired, this one the read path. - Removed a dead branch in audit's module main.
main()testedargs.dry_run, which no parser defines, so the "[DRY RUN]" header variant could never fire. - Version pin. New
tests/test_version.pyfails the suite whenpyproject.tomlandsrc/bindery/__init__.pydisagree (the v0.19.1 drift class); Phase 7's version-sync box. - Docs. pyproject's dependency comment and CLAUDE.md's hard constraint said
both git deps were hash-pinned; only
vir-tuiis (cquarry deliberately tracks@mainsince v0.19.0), and both texts now say so. The README's empty## Usageheading now introduces the three verbs.
- Dependency bump (deliberate pin move). The
vir-tuipin moved fromd13ad0etoe53f17e(vir-tui 2.2.0), per the repo policy of pinning exact commits and bumping deliberately. 2.2.0 is additive —progress_box(),interactive_session(),prompt_float/prompt_path/confirm,out_note,text_mode(), and results-pager search — and every API bindery-cli uses (ui.tqdm,ui.info,ui.print_header) is unchanged, so no behavior change is expected;uv.lockregenerated to match the new pin. - Version sync fix.
src/bindery/__init__.pystill said0.18.0whilepyproject.tomlsaid0.19.0; both now carry the release version (0.19.1) per the "pyproject must match VERSION below" rule.
- Accurate Calibre id resolution (no more guessing).
--install-to-calibrepreviously extracted the book id from the(123)directory-name fragment — a heuristic that replaces the wrong book whenever a directory was hand-renamed or its number no longer matches the catalog. Newlibrary.CalibreIdResolverbuilds a lazy, one-shot EPUB-path → id map through cquarry's own layout logic (CalibreDB.get_format_path()), so the id comes frommetadata.dband renamed directories cannot misroute a replacement. The(id)regex survives only as the no-catalog fallback; with neither source the repaired file is saved atomically in place (previous behaviour). - Single-entity audits:
bindery audit all --id 1234. New--idmode fetches exactly one row through cquarry'sCalibreDB.get_book()(no library-wide cache), resolves the EPUB viaget_format_path(), and reports the same verdicts as directory mode. Composes with--tag: the book is tagged only if the audit flags it. Mutually exclusive with the directory argument. - Dependency unpinned.
cquarrymoved from a frozen pre-1.1 commit to@main, matching Hermitage and CalibreQuarry — CI now always exercises current cquarry (1.6.x) instead of a year-old snapshot that lackedget_book/add_tag. - Roadmap: Phase 4 (Format Path Resolution, Safe Tag Application, Single-Entity Fetching) and the Phase 6 "cquarry Integration" checkbox are closed.
- Tests: 204 → 216 (12 new) — the resolver (catalogued hit, DB-truth-over-directory-name,
unknown file, missing catalog), the regex fallback, and the replace routing (resolver id
wins, regex fallback, atomic in-place fallback with
subprocessmocked);run_single(clean pass, unknown id, book without an EPUB) against a real minimalmetadata.db+ synthetic EPUB.
cquarry 1.2 adoption: audit tagging now actually reaches the write path, and tagged books regenerate their OPFs.
- Fixed:
audit --tagwas unreachable from the CLI. v0.18.0 documented--tag TAGand shipped_apply_audit_tag(), but theauditsubparser never registered the flag andrun_audit_cmdnever passed it — the write path existed only for Python callers. The flag now parses, defaults off (read-only audits stay read-only), and reachesrun_audit_library(tag=...); regression tests cover all three properties. - Tagged books are queued for OPF regeneration. The pinned cquarry bump (1.1 → 1.2) means every
WritableCalibreDB.add_tag()now records the book id in Calibre'smetadata_dirtiedqueue — the table upstream consumes to decide which sidecar.opfs to regenerate (and re-push to wireless readers) at next startup. Previously a bindery-cli-applied tag would appear in the GUI but never reach OPF/wireless sync. Verified end-to-end against a synthetic library: flagged book tagged,metadata_dirtiedgains exactly that id, clean books untouched. - Requires
cquarry@ e19c24c (v1.2.0); pin bumped accordingly.
cquarry 1.1 adoption: canonical format-path resolution and an opt-in tagging pass.
- Canonical EPUB resolution: library-mode audits (
bindery audit ...with no directory argument) now resolve each book's EPUB throughcquarry.db.CalibreDB.get_format_path()instead of hand-building<root>/<books.path>/<name>.epub. The storage-layout logic lives in exactly one place across the ecosystem, and a missing file surfaces as a normal scan error naming the expected path. --tag TAG(opt-in): after a library-mode audit, applyTAGto every flagged book via cquarry's separateWritableCalibreDBwrite module (trigger-safe: registers Calibre'stitle_sort/uuid4SQL functions, bumpslast_modified, cleans link tables before tags). The audit itself remains strictly read-only; nothing is written unless--tagis passed, and THIN emptytext advisories stay untagged. Books already carrying the tag are skipped (idempotent). Calibre must be closed so the write does not fight its lock.- Requires
cquarry>=1.1. - Tests: suite still green (201 tests) plus an end-to-end smoke run against a synthetic library:
foreign-language body text flagged, tagged
[Flagged],last_modifiedbumped, second run reports the book already tagged.
The default repair pass is well-formedness only again. v0.14.0 had promoted six structural
repairs into the always-on HTML_TRANSFORMS, so a plain bindery repair could delete
attributes, unwrap elements, and fabricate content on a book whose only defect was a stray
ampersand — exactly what spec.md forbids. Each now requires its own flag (on both repair and
library); --all still enables everything: --fix-empty-body, --fix-missing-title,
--fix-id-colons, --unwrap-block-in-inline, --strip-invalid-value,
--unwrap-illegal-tags.
The CSS precondition is real. --unwrap-illegal-tags no longer trusts callers to scan
stylesheets first. Every .css entry in the book plus each document's inline <style> blocks
are scanned for illegal-tag names used as element selectors (w { }, p st, x > w { },
pagebreak.new:after {}; class/id selectors like .st and #w do not protect), and those
names are protected for the whole book — styled formatting can never be silently destroyed.
Nested at-rules work (@media print { o > sentence {} }). This ports
scripts/find_css_illegal_tags.py into the library as transforms.css_protected_tags.
Housekeeping: git dependencies pinned to exact commits (vir-tui d13ad0ed = 2.0.0,
cquarry 4b771aa = 1.0.2) instead of floating branches, so builds are reproducible;
__init__.py VERSION resynced with pyproject.toml (drifted three releases behind at 0.15.0 vs
0.16.3); first tags cut (v0.16.3 backfilled on its release commit, v0.17.0 now);
spec.md / CLAUDE.md / README.md brought back in line with the code (audit subcommand, the pins,
the restored opt-in contract); roadmap Phase 6 boxes closed; run_epubcheck's docstring moved
above the daemon call where it is an actual docstring. Suite: 183 -> 201 tests.
- Fix: Resolved
test_audit.pyimport path failure. - Fix: Fixed
unwrap_block_in_inlinedata loss regex logic by capturing full text blocks. - Fix: Added
--replaceflag tocalibredb_replaceto prevent DB crash when file exists. - Fix: Corrected
analyze_brokentagsloop to properly analyze spines and docs instead of nonexistent iter method. - Fix: Removed hardcoded JVM release target in
EpubcheckDaemon. - Fix: Added regex word boundaries in
fix_id_colonsto prevent URL corruption. - Fix: Corrected
strip_invalid_valueregex to safely strip whitespace. - Fix: Fixed empty tag replacement regex in
fix_missing_title.
- Replaced local
ui.pymodule with standardizedvir-tuipackage.
- Build: Cleaned up orphaned scratch files to resolve strict ruff linting failures.
Centralized Calibre DB Access: bindery-cli has transitioned to the unified cquarry library for all read-only metadata.db accesses. The internal bindery.audit DB connection logic has been completely replaced with cquarry.db.CalibreDB. This inherits robust Calibre lock handling, database snapshotting, and ensures query logic stays perfectly synchronized with CalibreQuarry and Hermitage.
- Integrated EPUB Auditing (
bindery audit): Successfully absorbed the externalaudit_epubcodebase directly into bindery-cli as a native subcommand. This transforms bindery-cli from a pure repair tool into a complete EPUB lifecycle toolkit (Audit -> Repair -> Validate).content: Identifies non-English text by sniffing the text blocks, reporting the predominant script block (e.g., Cyrillic, CJK) to catch untranslated or mis-encoded EPUBs.pagenumbers: Detects books polluted with hardcoded print page numbers (using sliding window regex to find high-density sequential digits interrupting prose), seamlessly bridging intobindery repair --strip-paginationfor the fix.emptytext: Audits EPUBs to find effectively empty books or spine stubs (below configurable--min-charsand--thin-charsthresholds).ocr: Scans for systemic OCR damage by detecting disproportionately high densities of hyphenation, disjointed characters, or garbage sequences.- Library Integration: Directly operates on a Calibre library tree or loose directories, producing actionable console reports designed to be piped into
bindery library --audit.
- Supplementary Phase Structural Repairs: Integrated six highly targeted regex-based transforms to resolve the most common markup-level EPUB errors across the library, all safely guarded by the EPUBCheck gate:
- Duplicate NCX
playOrder: Safely rewrites duplicated playOrder attributes sequentially. - XML
idColon Violations: Translates illegal colons inid="X:Y"and their matching#X:Yfragment references to valid underscores. - Empty
<body>: Appends a non-breaking space to strictly empty body tags to satisfy parser requirements. - Missing
<title>: Injects a<title>Unknown</title>fallback in the<head>if missing, handling both empty<title/>self-closing tags and entirely absent tags. - Block-in-Inline Nesting: Safely unwraps
<span>tags that illegally contain a block-level element (e.g.<div>or<p>), leaving the block element intact. - Invalid
valueAttributes: Systematically strips invalidvalue="..."attributes from elements like<div>,<span>,<p>, etc.
- Duplicate NCX
- Illegal Tags Unwrapper: Added a sweeping transform that targets completely invalid or deprecated HTML tags that break EPUB3 validation (like
<st>,<font>,<sentence>,<o>,<w>, and<pagebreak>). Unwraps the tags without deleting their inner text. Automatically excludes tags when referenced by an EPUB CSS stylesheet to guarantee 100% format preservation.
- Epubcheck Daemon: Radically accelerated the
bindery libraryvalidation gate by implementing a transparent, persistent Java daemon. bindery-cli now automatically compilesFastDaemon.javain the background and pipes EPUB paths to it, eliminating the JVM startup penalty. - Validation time dropped from ~5 seconds per book down to under ~0.05 seconds per book, cutting library sweeps from 6 hours to 10 minutes.
- UI Upgrade: Integrated
tqdmprogress bars directly into the corebindery librarycommand, providing a real-time visual progress bar and ETA during the long-running sweep and candidate-processing phases. - Standardized documentation to reflect that
tqdmis a core required dependency, formally moving away from the "strictly stdlib-only" constraint.
- CLI Help Upgrade: Explicitly surfaced the
--install-to-calibreoption on the rootbindery --helpmenu by creating a dedicatedlibrary-specific integrationargparse group. - Documentation: Added a prominent code example for
--install-to-calibreto the Usage section of theREADME.md.
- CLI Help Formatting: Restructured the root
bindery --helpmenu to natively render the shared fix flags as a properly alignedargparsetable, matching the clean visual style of the subcommands and other portfolio tools.
- Oceanstrip Merged: Fully incorporated the
oceanstripstandalone tool directly into bindery-cli's primary pipeline. - Added
--strip-watermarksflag to systematically detect and remove producer/distributor watermarks (e.g. OceanofPDF stamps and zero-byte marker files, ABC Amber LIT Converter injections) from EPUBs. - This lossy transform is correctly governed by the
validate.no_worseEPUBCheck oracle, ensuring clean structural extraction without regressing file validity. - The
--strip-watermarksargument is automatically engaged when running bindery-cli with--all.
--strip-broken-tags: A new lossy transform to find and safely remove leaked HTML closing tags missing their open brackets (e.g.,</p>) that render as raw text in older uncorrected EPUBs. Guarded by thevalidate.no_worsegate.--install-to-calibre: A safer alternative to atomic replacement for Calibre users. Usescalibredb add_format --replaceto swap the corrected EPUB natively into the database, preserving all metadata, custom IDs, and read progress. Falls back to atomic replacement if the Calibre ID cannot be parsed from the filepath.--all: Added a meta-flag to automatically enable all opt-in non-fatal fixes and lossy strips (pagination, watermarks, tags, attributes, entities, image alt tags, etc.) simultaneously for a comprehensive repair.
- UI Upgrade: CLI scripts now feature rich terminal output with ANSI formatting,
tqdmprogress bars, and clear summary blocks. The project now relies ontqdmfor output formatting.
Repairing the same book twice now produces identical bytes. It did not before,
and had not for as long as the archive rewrite has existed. The mimetype entry was
written as zout.writestr("mimetype", ...) with a bare string arcname, and a string
arcname makes zipfile mint a fresh entry stamped with the current clock. Every other
entry in the archive is written from its source ZipInfo and keeps its own timestamp,
so this was the single entry that moved: two repairs seconds apart differed in exactly
those two bytes, with every entry's content identical.
Nothing malfunctioned because of it. No reader and no epubcheck run cares about that field. What it cost was the ability to checksum or diff two repairs to confirm they agree, which is exactly the check the v0.10.1 sweep leaned on, and it quietly made "deterministic repair" false at the byte level. spec.md now says determinism is meant at the byte level, so the claim is testable rather than aspirational.
The entry now carries the source entry's timestamp, or 1980-01-01 (the earliest a zip
can represent) when the entry is being added and there is nothing to inherit. It is
still built as a fresh ZipInfo rather than reusing the source one, because OCF
requires the mimetype entry to carry no extra field and reusing a broken book's entry
wholesale would propagate that violation; external_attr is set to what writestr
used to apply, so the timestamp is the only change to the output.
Measured, not assumed: three separate processes spaced over five seconds now produce one identical hash where they produced three. Over 99 real books, content, per-entry metadata, and fix reports are unchanged from v0.10.1, while timestamp preservation goes from 0/99 to 99/99. Three tests cover it, one of which documents that the naive "repair twice and compare" check passed by luck before the fix whenever both runs landed in the same clock second. Suite 122 to 125 tests.
A maintenance sweep, plus the roadmap decision recorded below.
Three bugs. Each fix was measured against the real library rather than argued from a fixture: the old and new code were run side by side over all 4,926 books (268,190 documents, content plus NCX). One of the three was actively firing; the other two are traps that were waiting for the right input.
--strip-bad-attrscould destroy a tag's self-closing slash. The unquoted-attribute-value pattern[^\s>]+swallowed the/in<img 31=x/>, so dropping the offending attribute left<img>: the one fix whose whole promise is "only the offending attribute is dropped" turned a well-formed self-closed tag into an unclosed one, introducing the very fatal it exists to remove. The gate would then reject the book, costing it every other repair it had earned. The pattern now stops before a/that ends the tag, while a/inside a value (a URL) is still consumed. No book in the library carries the pattern today, so this one was latent.strip_prolog_junkcounted a phantom fix for legal whitespace. Whitespace in the prolog is legal XML unless an XML declaration follows it, but any leading whitespace was stripped and counted. That marked the document changed, which forcesrepair_epub'sdecode("utf-8", "replace")round-trip on a file that had nothing wrong with it, and turned books that should reportnochangeinto books that spend two multi-second epubcheck runs to concludeequal. Whitespace before a declaration (where it really is fatal), a BOM, and any non-whitespace junk are still stripped. This is the one that was live: across the library the transform fired 119 times before and 106 after, so 13 phantom fixes in 7 books are gone, and every genuine case (a BOM ahead of an XML declaration, the real form in this library) still fires. Running the whole pipeline over those 7 books before and after confirms the change is surgical: every other fix count is untouched (fix_ncx_ids10 stays 10,fix_named_entities30 stays 30,stripped_pagination1 stays 1), only the phantom rewrites disappear.- Invalid-id renaming was not deterministic.
--fix-idsiterated the id set, so when two invalid ids wanted the same replacement (1:2and1_2both yieldid_1_2) which one received the extra_prefix depended on the hash seed: the same book repaired to different bytes from run to run, against a project that sells itself on deterministic repair. Both the OPF and NCX paths now plan renames through one sorted helper,_plan_renames.
The lint result no longer depends on which ruff is installed. There was no
[tool.ruff] section, so the rule set came from ruff's defaults: CI's pinned
0.15.20 passed clean while 0.16.2 reported eight findings on the same tree. The
selection is now pinned in pyproject.toml (ruff's default set plus import
sorting, pyupgrade, and bugbear), with a note on the two families deliberately left
out because they fight intentional code. CI moves to actions/checkout@v5, pins
ruff 0.16.2, and lints . rather than three named directories so a new top-level
script cannot escape.
Smaller: repair_epub opens the source archive once instead of twice; the dead
skip_empty parameter is gone from pagination._nearest; library --limit 20 on
a tree with 3 candidates counts [1/3] instead of [1/20]; and pyproject.toml
catches up to the VERSION bump in __init__.py. Suite 114 to 122 tests.
The Calibre hook question is settled, and it is FileTypePlugin with
on_import = True. Researched against the Calibre source rather than guessed:
run(path_to_ebook) is handed the file being imported and returns the path to a
modified copy, and Calibre imports that instead. So a repaired EPUB is what enters
the library, the user's original on disk is never touched, and nothing writes to
metadata.db.
That kills a roadmap item. The "optional metadata.db nudge so Calibre notices
the new file size" existed only because in-place replacement changes a file behind
Calibre's back. Under on_import the repair happens before Calibre reads the
file, so the size it records is already right. The nudge was solving a problem
that the correct hook does not create.
Two constraints recorded with the decision, because they shape what the plugin can
be: it runs inside Calibre's bundled Python, so it must vendor its own code and
cannot assume bindery-cli is installed (the minimal-dependency rule is what makes this
practical, and the optional html5lib path cannot come along); and epubcheck
cannot gate an import, being an external Java process of seconds per book, so the
plugin does the deterministic transforms and leaves gated work to the CLI where a
human is watching.
Locale hardening for the epubcheck wrapper (roadmap 5.2, the last open Phase 5
item). The wrapper parsed counts from epubcheck's human-readable summary line,
which the JVM localizes: on a non-English locale every book parsed as None and
was reported as error. Counts now come from epubcheck's locale-independent
--json - output first (the checker totals, verified against epubcheck
5.3.0), with the English summary-line regex kept as the fallback for
epubchecks too old for --json, and the JVM pinned to English via
JAVA_TOOL_OPTIONS (appended, so existing JVM flags survive) so that fallback
stays meaningful. New tests/test_validate.py (9 tests) mocks
subprocess.run itself to pin the JSON path, both fallbacks, the env pin, and
the failure modes; suite grows 105 to 114.
Two new fixes, born from a Dan Brown import batch whose books opened fine but carried 89 and 225 epubcheck errors.
--fix-idsnow covers the NCX. Old conversions stamp navPoint ids from UUIDs (digit-led) or colon-bearing strings; epubcheck rejects every one as RSC-005. Invalid NCX ids are renamed with the sameid_scheme as manifest ids. NCX ids are internal to the NCX (nothing in the OPF or content documents references them), so the rename needs no cross-file bookkeeping. Counted asfix_ncx_idsin reports.- New:
--add-img-alt(opt-in). Addsalt=""to<img>elements missing the required attribute. Rendering is unchanged (an empty alt draws nothing), but this is the one transform that adds markup the author never wrote, andalt=""asserts "decorative" to a screen reader where a missing alt did not; hence opt-in, never a core transform. Quote-aware, idempotent, CDATA and comments never rewritten; counted asimg_alt_added.
Real-world validation: the two motivating books went 89 errors to 0 and 225 to 190 (the remainder are dead NCX fragment identifiers, a fix candidate we deliberately passed on), both accepted by the normal gate.
New: --escape-unknown-entities (opt-in). The last fix candidate from the
v0.5.0 audit. An entity name that is neither XML-predefined nor in the HTML5 table
stays a fatal "entity not declared" (the core fix_named_entities deliberately
leaves it); with this flag such references are escaped (&foo; -> &foo;),
which renders exactly as browsers already render an unknown entity: the literal
text.
- Conditionally semantics-preserving, hence opt-in: rendering is identical except
against a document whose DOCTYPE internal subset declares the entity, so any
document carrying an internal subset (
<!DOCTYPE ... [) is skipped wholesale. - CDATA sections and comments are never rewritten (the standing transform
invariant), the normal epubcheck gate applies, and the fix is idempotent (the
&it emits is predefined and stays put on a re-run). - Available on both
repairandlibrary; counted asescape_unknown_entitiesin reports.
Phase 2 closes out: the audit workflow is now self-contained, and the mimetype fix joins the core repair set.
--sweep: re-audit integration.library --only fatals --sweepruns a live epubcheck sweep for candidate selection, replacing the separate CSV step (and with it the whole audit-path-mismatch bug class). Each sweep result is reused as that book's before-measurement, so no book is epubchecked twice. Mutually exclusive with--auditand--no-validate;--limitkeeps the sweep lazy.--json FILE: machine-readable run report. Per-book path, status, before/after counts, fix summary, and applied flag, plus the summary totals, for scripting and cross-run comparison.--manual-list FILE: the manual follow-up export. One path per line for every book the run did not (or could not) auto-repair: nochange, equal, partial, reject, error, unreadable.- A missing
mimetypeentry is added, and wrong or whitespace-padded content is normalized to the OCF constantapplication/epub+zip. The content is spec-constant, so this is deterministic and semantics-preserving; it is counted (mimetype_added/mimetype_normalized) and gate-checked like any other fix. - spec.md documents void end-tag swallowing (the 5.6 gap):
self_close_voidalso removes orphaned end tags for void elements (</br>,</col>), which are always invalid and cannot change what renders. Behavior unchanged since v0.4.2; the spec now says so.
The Phase 5 audit sweep: three confirmed safety bugs fixed, packaging honesty, and a round of CLI hardening and UX. Every fix ships with a unittest regression test.
Safety and correctness:
--strip-paginationcan no longer auto-apply a still-fatal book. Theno_worseacceptance bar for the lossy strip used to overwrite the gate's verdict outright, so a book going 3 fatals -> 1 fatal was classifiedacceptandlibrary --applyatomically replaced a book that still does not open.no_worsenow relaxes only the improvement demand: a result with remaining fatals is demoted topartialand never applied. This closes a hole in the hard rule that still-fatal books are manual work.- One corrupt
.epubno longer aborts an entirelibraryrun. A non-zip, truncated, or encrypted book raised out of the sweep and killed a multi-hour run with a traceback. Each book is now guarded individually; unreadable books are reported, counted in a newunreadable:summary line, and skipped.repairprints a clean error for the same case instead of a traceback. - The pagination strip no longer deletes
<p id=...>navigation targets. An id on the removed paragraph itself (the common<p id="page7">7</p>page-anchor shape) vanished with the block, breaking NCX page-lists and internal links; only inner<a id=...>anchors were rescued. Both removal paths now preserve it: delete-only keeps an emptied<p id=...></p>shell, and a merge hoists the id as an empty anchor. Single-quoted ids are recognized too.
Packaging:
html5libis now the optional extra the docs always promised. It moved fromdependenciesto[project.optional-dependencies], so a plain install is genuinely minimal;bindery[reserialize]pulls it in for--reserialize.
CLI hardening:
repairlabels partial output honestly. A book whose fatals were reduced but not cleared was written with arepaired:line that read as fixed; it now printsPARTIAL (still has fatals; needs manual work):.repairrefuses to overwrite an existing output file unless--forceis given.- Single-quoted attributes are visible to the OPF/NCX regexes. The NCX-001 sync, OPF location, and unique-identifier lookup all required double quotes, so a single-quoting toolchain made them silently no-op. All accept either quote now.
Book.EPUBis found. The library scan matches the.epubsuffix case-insensitively (Calibre emits lowercase, but a hand-added file should not be invisible).
UX:
- Progress output for long runs. A
[123/4051] Author/Title.epubline per book goes to stderr, so a mostly-clean library no longer shows hours of silence; stdout stays a clean report.--quietsuppresses it. - A warning fires when the audit CSV matches zero scanned books (the silent path-mismatch trap that read as "library is clean"). Paths are resolved on both sides first, so relative-vs-absolute mismatches no longer occur at all.
- Backup flags warn when inert.
--backup/--backup-inplacewithout--applyprint a note;--apply --strip-paginationwithout any backup flag prints a loud recommendation (the one lossy mode deserves a backup). libraryexits 2 when any book was rejected, unreadable, or failed epubcheck, so scripts and cron can detect trouble; 0 is a clean sweep, 1 a usage error.--limitlimits the scan, not just the work. Candidates are consumed lazily, so--only ncx --limit 20stops opening archives after the 20th candidate instead of probing every book in the tree.
New: --strip-pagination (opt-in, lossy). Removes print page numbers and running headers that a PDF/OCR conversion baked into the body text as literal paragraphs, which reflow into the middle of sentences ("where the hay cart 16 was taking him"). This is the first mode that removes visible content, so it is a deliberate, fenced-off exception to bindery-cli's semantics-preserving rule: off by default, and accepted by a new no_worse bar (no net-new fatals or errors) instead of the improvement-demanding gate, because a baked page number is valid markup that epubcheck cannot see.
- Removes only injected furniture, never the author's prose. Where a number split a sentence it rejoins the two paragraphs (closing up a word split like
compli-/mentary); page-listidanchors are hoisted into the merged paragraph so navigation still resolves. - A book is treated as paginated only when it has both a dense run of standalone arabic numbers (>= 20) and several confident mid-sentence interrupts (>= 3), so a merely chapter-numbered book is never touched. Roman chapter/front-matter numerals and year-range values are preserved.
- Three independent safety nets guard every edit, any failure leaving the document unchanged: character conservation (no prose character lost or fabricated),
<p>/<a>tag balance, and the epubcheck no-regression check. - Validated on /tmp copies of the real library: Fingersmith 372 numbers removed (zero left), Animal Farm 54 removed with all ten roman chapter numbers intact, zero prose characters changed in either, epubcheck no worse.
- EPUB 3 namespace prefixes are preserved.
--strip-bad-attrsno longer drops perfectly valid EPUB 3 prefixed attributes. It now correctly parsesepub:prefixandprefixdeclarations (e.g.epub:prefix="math: ...") instead of strictly requiring anxmlns:declaration to bind a prefix. - Nested orphaned void end tags are swallowed globally.
self_close_voidnow completely strips all explicit end tags for void elements globally (like</br>and</img>) after self-closing the start tag, preventing fatal XML parse errors when tools generate deeply nested orphaned void end tags (e.g.,<br><br></br></br>).
Bugfix and cleanup sweep. No new fixes or flags; several of these close real holes in the safety contract.
-
CDATA sections and comments are never rewritten. The transforms used to escape
&, convert entities, and self-close<br>inside<![CDATA[...]]>and<!-- -->, where that content is literal and already legal XML; escaping a&in CDATA-wrapped CSS/JS changes what renders. All body-text transforms (including--strip-bad-attrs) now skip these spans. This is now a spec invariant. -
Hyphenated custom elements are no longer mangled.
-,:, and.are valid XML name characters but not word characters, so the v0.2.0�boundary still let<colmatch inside<col-group>and self-close it. The matcher now requires whitespace,/, or>after the element name. -
The OPF is located via
META-INF/container.xmlinstead of "first.opfin archive order", so a stray duplicate OPF can no longer win and sync the wrong uid into the NCX. Falls back to the old behavior when container.xml is absent. -
--fix-idsupdates all references, not just the spine.fallback=,media-overlay=, and the EPUB 2<meta name="cover" content="...">also point at manifest ids; leaving them stale orphaned fallback chains and broke Calibre's cover detection when the cover item's id was renamed. -
Duplicate entry names survive the rewrite.
ZipFile.read(name)returns the first entry's bytes for every same-named duplicate (seen in broken EPUBs); entries are now read individually. -
Multi-codepoint entities are converted (
≂̸and friends become one numeric reference per codepoint) instead of being skipped. -
atomic_replacecleans up after itself and syncs the directory. A failure mid-replace no longer leaves a.bindery.tmpin the library, and the rename is fsynced so a crash right after a replace cannot lose it. -
--limit 0now means "process nothing" instead of being ignored; Ctrl-C during a long run exits cleanly (130) instead of dumping a traceback. -
repairnow writes the exact bytes the gate accepted. It used to produce the final output with a second repair pass that dropped--fix-ids,--reserialize, and--strip-bad-attrs, so the written file could be missing the very repairs epubcheck had just validated. The gated temp file is now copied to the output. -
An epubcheck failure no longer bypasses the gate. When epubcheck crashed, timed out, or produced unparsable output mid-run, the book was classified
unvalidatedandlibrary --applyreplaced it as if it had passed. Such books are now a distincterroroutcome: reported, counted, and never applied or written. Only an explicit--no-validateskips the gate. -
Unchanged archive entries are copied byte-for-byte. Eligible entries were decoded with
utf-8/replaceand re-encoded even when no transform fired, which would silently swap non-UTF-8 bytes for U+FFFD in otherwise untouched files. -
library --only fatalswithout--auditis an error. It used to silently treat every book in the library as a candidate; the README always saidfatalsneeds the audit CSV, and now the CLI enforces it. -
Audit CSVs without a header row no longer lose their first book; blank rows are skipped instead of crashing the load.
-
NCX
dtb:uidreplacement inserts the uid literally (a uid containing\1was previously parsed as a regex replacement template), and per-file change accounting no longer leaks across multiple.ncxentries in one archive. -
Cleanup: shared CLI flags are defined once for both subcommands, the transform pipeline is properly typed, and the spec/README document the new
erroroutcome and the byte-preservation guarantee.
- New
--strip-bad-attrs. Drops attributes that are invalid XML and so make a document unparseable: a name starting with a digit (e.g. a mangled31="") or a namespaced name whose prefix is never declared (e.g. Office VMLv:shapeswith noxmlns:v). It is surgical (only the offending attribute is removed) and a no-op on well-formed files, since those cannot contain such attributes. Off by default. - This cleared the last 2 markup-fatal library books that survived
--reserialize: The Selfish Gene (v:shapes) and The Rustonomicon (broken SVG31=""). Both now validate with zero fatals, open in Calibre, and preserve their full text. With this, the entire 38-book fatal set from the original audit is resolved.
- New
--reserialize(structural repair). Rebuilds content documents that are still not well-formed by re-parsing them with html5lib (lenient HTML5 recovery, like a browser) and re-emitting XHTML. This closes unclosed non-void elements (<p>,<div>,<span>,<blockquote>,<body>) that the regex transforms cannot, and even recovers some corrupted tag names. It runs only on documents that are not already well-formed, so good files are left byte-for-byte unchanged, and only when opted in. - New dependency: html5lib (for
--reserializeonly). Imported lazily; every other mode runs with no third-party dependency. This is the one approved exception to the minimal-dependency design. - Verified on the 12 markup-fatal library books:
--reserialize --fix-idsclears 10 of 12 to zero fatals (content preserved; the 2 holdouts are Office-VML and broken-SVG foreign content). All gate-accepted.
- Hardened
self_close_void. The matcher now requires a word boundary after the element name and is quote-aware, fixing a bug where<colmatched inside<colgroup>(self-closing it and orphaning the end-tag) and where a>inside an attribute value ended the tag early. This introduced fatals on 19 books during the library run; the gate rejected them, and they are now repaired cleanly. Already-self-closed tags are left untouched and not counted. - New
--fix-ids(RSC-005). Optionally rewrite manifest item ids that are not valid XML names (start with a digit, contain a colon) and update their spine references. Off by default, since it touches the OPF; the dc: metadata is never altered. On a real book this cleared 36 bad ids (794 to 723 errors), gate-accepted.
First release. A focused EPUB repair tool, sibling to oceanstrip, born from auditing a 3713-book Calibre library where 38 books carried fatal parse errors.
- Deterministic, semantics-preserving transforms: self-close void elements, named
entity to numeric reference, escape bare
&, strip pre-prolog junk, collapse a duplicated rootxmlns. - NCX-001 fix: sync
toc.ncxdtb:uidto the OPF unique identifier. - mimetype ordering/compression repair on rewrite.
- Two-mode epubcheck gate that understands fatal unmasking: when a book had fatals, success is fewer fatals, and the error count rising as hidden errors surface is not treated as a regression.
repair(single file) andlibrary(batch) CLI modes. Library mode is a dry run by default;--applyreplaces accepted books in place, atomically, with optional backups. Only the.epubis touched, so Calibre's Quality Check sync can reconcile the database.unittestsuite; minimal dependencies.
Validated on the real library: 24 of ~40 fatal books fully de-fataled (they now open), 6 partially improved and flagged for manual finish, the rest left untouched, and zero epubcheck regressions.