You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Complementary to #1451 (reproducibility). This is the validity / quality axis: is the ML/DL detection actually good, and can each finding's contribution be attributed per engine? Verified against the authoritative repo via gh api.
Note: docs/plans/PLAN_SYNTHETIC_DATA_AND_CONFIDENCE_VALIDATION.mdalready exists but is Deferred. This promotes a minimal slice for 1.8.0.
Verified findings
tests/test_ml_engine.py tests the classic ML (TfidfVectorizer + RandomForestClassifier, core/ml_engine.py) — not the DL path.
No direct DL-backend test (loads SentenceTransformer, generates embeddings, trains the LogisticRegression): grep of tests/ for SentenceTransformer|DLClassifier is empty.
Combination policy combined = max(ml_confidence, dl_confidence) (core/detector.py). DL is an additive layer → DL inactivity raises false-negative risk, not false-positive.
No ground-truth corpus with intentional FP/FN; no precision/recall/F1 per run; no per-finding engine provenance.
Acceptance criteria (minimal slice for 1.8.0 — not a #1021-scale marathon)
Frozen ground-truth corpus + manifest (train / validation / closed test never used for train/tune / hard-negatives). Split by semantic family / PII category / language / source, not random-by-line (near-synonyms leak).
Per-engine ablation: regex-only, regex+ML, regex+DL, regex+ML+DL -> precision / recall / F1 / confusion, by category / language / target.
DL activation-proof gate mode: DL mandatory, fail-closed if is_ready=False (to test DL specifically). Production keeps graceful fallback.
Calibration: Brier, reliability bins, actual positive-rate per score band. If not calibrated, rename "confidence" -> detector_score / review_priority_score (not "probability of PII").
Degraded-mode posture: DL unavailable -> reinforce HITL, never conclude "no PII"; report dl: UNAVAILABLE + reason_code + degradation: BASELINE_DETECTION_REMAINS_ACTIVE.
Field-evidence ledger (metadata-only, no raw content): Markts (post-8 / post-12), Estela, Rafael runs -> corpus / env / confirmed-finding / active-engine / decision-source. Turns testimonials into auditable field evidence.
Promote PLAN_SYNTHETIC_DATA_AND_CONFIDENCE_VALIDATION.md from Deferred; python scripts/plans_hub_sync.py --write; update docs/plans/PLANS_TODO.md.
Not in question
Regex/rules deterministic; "no LLM / no black box" hold. This closes the model-validity gap, distinct from the reproducibility gap in #1451.
Read-only audit (Claude Code). Evidence via gh api on the authoritative repo. Implementation: Cursor.
Context
Complementary to #1451 (reproducibility). This is the validity / quality axis: is the ML/DL detection actually good, and can each finding's contribution be attributed per engine? Verified against the authoritative repo via
gh api.Note:
docs/plans/PLAN_SYNTHETIC_DATA_AND_CONFIDENCE_VALIDATION.mdalready exists but is Deferred. This promotes a minimal slice for 1.8.0.Verified findings
tests/test_ml_engine.pytests the classic ML (TfidfVectorizer+RandomForestClassifier,core/ml_engine.py) — not the DL path.SentenceTransformer, generates embeddings, trains theLogisticRegression):grepoftests/forSentenceTransformer|DLClassifieris empty.combined = max(ml_confidence, dl_confidence)(core/detector.py). DL is an additive layer → DL inactivity raises false-negative risk, not false-positive.Acceptance criteria (minimal slice for 1.8.0 — not a #1021-scale marathon)
regex_score,ml_score,dl_score,combined_score,decision_source,dl_backend_active.is_ready=False(to test DL specifically). Production keeps graceful fallback.detector_score/review_priority_score(not "probability of PII").dl: UNAVAILABLE+reason_code+degradation: BASELINE_DETECTION_REMAINS_ACTIVE.PLAN_SYNTHETIC_DATA_AND_CONFIDENCE_VALIDATION.mdfrom Deferred;python scripts/plans_hub_sync.py --write; updatedocs/plans/PLANS_TODO.md.Not in question
Regex/rules deterministic; "no LLM / no black box" hold. This closes the model-validity gap, distinct from the reproducibility gap in #1451.
Read-only audit (Claude Code). Evidence via
gh apion the authoritative repo. Implementation: Cursor.