fix(ci): make the invisible-character gate detect, and able to fail - #47
Conversation
MEASURED 2026-08-27: this gate's pattern caught 0 OF 6 invisible-character test
cases. It has never detected an NBSP, zero-width space, BOM, soft hyphen, bidi
override or word joiner.
ROOT CAUSE: the pattern used UTF-8 BYTE sequences (\xc2\xa0) while grep -P
matches CHARACTERS. Bytes c2 a0 are ONE character U+00A0; \xc2\xa0 asks for TWO
characters, U+00C2 then U+00A0, which is never present.
grep -P '\xc2\xa0' -> miss
grep -P '\x{a0}' -> MATCH
Only \x00 worked, being single-byte in both readings.
FIXED: codepoint escapes; C0 control characters \x01-\x08,\x0B,\x0C,\x0E-\x1F
added (TAB/LF/CR excluded); and grep -a, without which grep skips any NUL-bearing
file as binary.
The C0 range matters: a stray BACKSPACE byte made a workflow unparseable in
developer-ecosystem, so it never ran, and this linter called it clean.
Canonical fix: hyperpolymath/empty-linter#70. 1 file(s) here.
VERIFIED: YAML re-parsed, and the corrected pattern was confirmed to catch a real
NBSP before the change was kept.
Up to standards ✅🟢 Issues
|
There was a problem hiding this comment.
Pull Request Overview
While this PR correctly identifies the need to update the invisible-character gate to support Unicode codepoints and handle binary file detection, the current implementation contains a critical logic error. Specifically, the $PATTERNS variable is defined but omitted from the grep execution, which will cause the CI gate to match every file and fail the build regardless of content.
Although Codacy reports that the changes are 'up to standards', this logic flaw renders the invisible-character check non-functional. This must be addressed before merging to prevent a CI blockage.
About this PR
- The grep command on line 169 fails to use the defined $PATTERNS variable, using an empty string instead. This will cause the linter to match every single file processed by 'find', rendering the filter logic useless and likely causing the CI to fail on every file in the repository.
Test suggestions
- Identify NBSP character (U+00A0) in a source file
- Identify C0 control character (e.g. Backspace \x08) in a source file
- Prevent grep from skipping files containing NUL (\x00) bytes using -a
Prompt proposal for missing tests
Consider implementing these tests if applicable:
1. Identify NBSP character (U+00A0) in a source file
2. Identify C0 control character (e.g. Backspace \x08) in a source file
3. Prevent grep from skipping files containing NUL (\x00) bytes using -a
TIP Improve review quality by adding custom instructions
TIP How was this review? Give us feedback
| -o -name '*.idr' -o -name '*.zig' -o -name '*.v' -o -name '*.jl' \ | ||
| -o -name '*.gleam' -o -name '*.hs' -o -name '*.ml' -o -name '*.sh' \) \ | ||
| -exec grep -Prl "$PATTERNS" {} \; > /tmp/empty-lint-results.txt 2>/dev/null | ||
| -exec grep -aPrl "" {} \; > /tmp/empty-lint-results.txt 2>/dev/null |
There was a problem hiding this comment.
🔴 HIGH RISK
The grep search pattern is currently an empty string "", which causes it to match all files regardless of their content. You should use the $PATTERNS variable defined on line 158 to correctly filter for invisible Unicode characters and C0 controls.
| -exec grep -aPrl "" {} \; > /tmp/empty-lint-results.txt 2>/dev/null | |
| -exec grep -aPrl "$PATTERNS" {} \; > /tmp/empty-lint-results.txt 2>/dev/null |
The scan used UTF-8 byte sequences as PCRE escapes. In a UTF-8 locale
grep -P reads \xc2 as codepoint U+00C2, not as a byte, so '\xc2\xa0'
never matched a real NBSP. Measured on GNU grep 3.11 and ugrep 7.8.4:
byte form 0 matches, codepoint form \x{a0} 1 match.
- switch every escape to the codepoint form \x{...}
- add a byte-level BOM check (od on the first three bytes); grep
implementations disagree about a BOM at byte 0, so do not rely on one
- add a positive control: plant a ZWSP file and a BOM file in
RUNNER_TEMP each run and fail if either instrument does not fire, so
a broken scanner cannot masquerade as a clean tree again
- count the union of both passes, not the sum (a BOM'd file appears in
both under GNU grep)
- exit 1 on findings; the job previously emitted warnings and passed
unconditionally, so it could not fail
- report the scanned-file count, and stop calling an incomplete scan
"skipped" in the step summary
Plant-verified locally against the full tree, 614 in-scope files:
0 -> 1 -> 2 -> 1 -> 0 findings, exit 0/1/1/1/0.
|



The problem
dogfood-gate.yml's invisible-character scan wrote its patterns as UTF-8 bytesequences:
In a UTF-8 locale
grep -Preads\xc2as codepoint U+00C2 (Â), not as a byte. A realNBSP is U+00A0, so the pattern matched nothing — and matched nothing silently, because
zero findings is exactly what a clean tree looks like.
Measured against three planted files (NBSP, ZWSP, BOM): 0 of 3 caught.
The job also emitted
::warningannotations and then fell through with noexit 1anywhere, so even a correct pattern could not have failed the build. Three independent
reasons it could never fire.
What changed
\xc2\xa0— never matches\x{a0}odcheck::warning, job passes::error,exit 1Positive control. The step plants a ZWSP file and a BOM file in
$RUNNER_TEMPon everyrun and fails if either instrument does not fire, or if either fires on a clean file. This
is the part that prevents a recurrence: the original bug survived because a bare zero is
compatible with both "the tree is clean" and "the scanner is broken", and nothing in the
job distinguished them.
It deliberately does not assert that the pattern pass catches a byte-0 BOM, because
that is the one behaviour grep implementations differ on — measured: GNU grep 3.11 matches
\x{feff}at byte 0, ugrep 7.8.4 strips the BOM first and cannot. Runners ship GNU grep,so the
odpass is currently redundant; it is there so the gate stays correct if thatchanges, and so the annotation names the actual defect rather than "some invisible
character".
Verification
Run locally against the full tree (614 in-scope files), driving the extracted step script
with
GITHUB_WORKSPACE/RUNNER_TEMP/GITHUB_OUTPUTset:0 → 1 → 2 → 1 → 0. The fourth row is the dedupe check: one file found by both passescounts once and gets the specific BOM annotation, not two generic ones.
Also:
shellcheck -S style -s bashclean,bash -nclean, YAML parses underyqandruby's parser. No
uses:line changed, so.github/workflows/actions.lockis unaffected;line 1 (SPDX) is untouched.
.json,.yml,.md,.sh,.tomland 15 other extensions are already in thefindclause, so the BOM check has real files to look at rather than shipping vacuous.
Pre-existing reds, not caused by this PR
Three checks are red here and are red on
mainindependently — confirmedfailureonmain's own governance run
33826003869:
governance / Allowlist Preflightgovernance / Language / package anti-pattern policygovernance / Validate Hypatia BaselineA fourth,
Sustainability Analysis(OikosBot), failed on the stale branch withdocker: Error response from daemon: manifest unknown→ exit 125. Main was failing thistoo — five consecutive
failureruns on 08-27/08-28 — and was cured on 09-02 by #51(
a4e1d92), which bumpedhyperpolymath/oikosbot@v0.1.1→@v0.1.3and pinned anexplicit
image: ghcr.io/hyperpolymath/oikos@sha256:aa2b409a…because the action's defaultimage digest had been garbage-collected by GHCR. This branch has been updated onto main, so
it now carries that fix.
Self-merged under the standing
--admingrant, accepting those three pre-existing reds.