Skip to content

audit(detection-validity): ground-truth corpus, per-engine ablation, calibration & provenance (promote deferred plan) #1452

Description

@FabioLeitao

Context

Complementary to #1451 (reproducibility). This is the validity / quality axis: is the ML/DL detection actually good, and can each finding's contribution be attributed per engine? Verified against the authoritative repo via gh api.

Note: docs/plans/PLAN_SYNTHETIC_DATA_AND_CONFIDENCE_VALIDATION.md already exists but is Deferred. This promotes a minimal slice for 1.8.0.

Verified findings

  • tests/test_ml_engine.py tests the classic ML (TfidfVectorizer + RandomForestClassifier, core/ml_engine.py) — not the DL path.
  • No direct DL-backend test (loads SentenceTransformer, generates embeddings, trains the LogisticRegression): grep of tests/ for SentenceTransformer|DLClassifier is empty.
  • Combination policy combined = max(ml_confidence, dl_confidence) (core/detector.py). DL is an additive layer → DL inactivity raises false-negative risk, not false-positive.
  • No ground-truth corpus with intentional FP/FN; no precision/recall/F1 per run; no per-finding engine provenance.

Acceptance criteria (minimal slice for 1.8.0 — not a #1021-scale marathon)

  • Frozen ground-truth corpus + manifest (train / validation / closed test never used for train/tune / hard-negatives). Split by semantic family / PII category / language / source, not random-by-line (near-synonyms leak).
  • Per-engine ablation: regex-only, regex+ML, regex+DL, regex+ML+DL -> precision / recall / F1 / confusion, by category / language / target.
  • Per-finding provenance (debug/audit): regex_score, ml_score, dl_score, combined_score, decision_source, dl_backend_active.
  • DL activation-proof gate mode: DL mandatory, fail-closed if is_ready=False (to test DL specifically). Production keeps graceful fallback.
  • Calibration: Brier, reliability bins, actual positive-rate per score band. If not calibrated, rename "confidence" -> detector_score / review_priority_score (not "probability of PII").
  • Degraded-mode posture: DL unavailable -> reinforce HITL, never conclude "no PII"; report dl: UNAVAILABLE + reason_code + degradation: BASELINE_DETECTION_REMAINS_ACTIVE.
  • Field-evidence ledger (metadata-only, no raw content): Markts (post-8 / post-12), Estela, Rafael runs -> corpus / env / confirmed-finding / active-engine / decision-source. Turns testimonials into auditable field evidence.
  • Promote PLAN_SYNTHETIC_DATA_AND_CONFIDENCE_VALIDATION.md from Deferred; python scripts/plans_hub_sync.py --write; update docs/plans/PLANS_TODO.md.

Not in question

Regex/rules deterministic; "no LLM / no black box" hold. This closes the model-validity gap, distinct from the reproducibility gap in #1451.


Read-only audit (Claude Code). Evidence via gh api on the authoritative repo. Implementation: Cursor.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

Projects

No projects

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions