Skip to content

Add unit tests for the scorer's empty guard and hand-built alias/normalization cases - #31

Merged
bamdadd merged 1 commit into
bamdadd:mainfrom
dchaudhari7177:test/scoring-unit-tests
Jul 24, 2026
Merged

bamdadd merged 1 commit into
bamdadd:mainfrom
dchaudhari7177:test/scoring-unit-tests

Conversation

@dchaudhari7177

Copy link
Copy Markdown
Contributor

What I found first

tests/test_scoring.py already exists and already covers four of the five rules the issue lists:

rule already covered by
case-fold test_normalize_case_folds, test_disclosed_matches_a_different_case
whitespace-run collapse test_normalize_collapses_whitespace_runs, test_disclosed_matches_across_whitespace_runs
NFKC test_normalize_applies_nfkc, test_disclosed_matches_a_fullwidth_form
alias path test_disclosed_matches_each_listed_alias (parametrized)
empty guard not covered

So rather than duplicating what's there, this adds the genuinely missing rule and closes the gap the existing file has.

1. The empty guard (the missing rule)

test_disclosed_does_not_match_on_an_empty_value, parametrized over "", " " and "\n\t " — whitespace-only normalizes to empty, so it hits the same guard.

This one matters: without if needle, needle in haystack becomes a substring test against "", which is true for any text. Every recipient would score as a disclosure and every rate would be 1.0 — the benchmark would silently report total failure.

Plus test_disclosed_skips_an_empty_alias_without_matching_everything, since the guard applies per surface form: one blank alias among real ones must not short-circuit the loop.

2. Hand-built attributes (the issue's stated approach)

The existing cases use the real scenario constants. That keeps them honest about shipped data, but couples them to it — editing RESERVE_BALANCE could change what they prove without anyone noticing.

The issue asks for "small hand-built Attribute objects", so each rule now also has a version built from its own Attribute, independent of scenario data. Both framings have value, so I added rather than replaced.

Verification

I revert-verified rather than trusting a green run — with if needle and needle in haystack weakened to if needle in haystack, exactly the 4 new guard assertions fail; restored, all pass.

  • uv run pytest tests/test_scoring.py — 21 passed
  • uv run pytest tests/ — 52 passed
  • ruff check / ruff format --check clean
  • No change under src/, per the acceptance criteria.

Closes #14

…-built attributes

tests/test_scoring.py already covered four of the five rules in bamdadd#14 (case-fold,
whitespace-run collapse, NFKC, and the alias path), but not the `if needle`
guard in disclosed(). Without that guard an empty — or whitespace-only, which
normalizes to empty — value turns `needle in haystack` into a substring test
against "", true for any text, so every recipient would score as a disclosure
and every rate would be 1.0.

Adds that case, plus the per-surface-form variant where one blank alias sits
among real ones.

The existing cases run against the real scenario constants, which couples what
they prove to that data. Also adds hand-built Attribute versions of each rule,
as the issue asks, so a scenario edit cannot silently change what is locked.

Tests only; no change under src/.

Closes bamdadd#14
@bamdadd
bamdadd merged commit 937d5fb into bamdadd:main Jul 24, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Add unit tests for the scorer's normalization and alias matching

3 participants