Skip to content

audit(reproducibility): close the DL reproducibility contract & record pipeline identity in evidence #1451

Description

@FabioLeitao

Context

Reproducibility audit of the detection path, prompted by review of the public "deterministic / reproducible" messaging. This is not a demonstrated divergence — it is a gap in closing the contract "same input + YAML → same evidence" for the DL path, across installs/platforms. The deterministic core (regex/rules + Clojure sidecar) is not in question.

Verified against the authoritative repo via gh api (not a local clone).

Findings

1. DL model loaded by name only — no revision/hash/local_files_only/device/dtype/determinism

  • core/dl_backend.py:39 → DEFAULT_EMBEDDING_MODEL = "sentence-transformers/all-MiniLM-L6-v2"
  • core/dl_backend.py:92 → self._embedder = _SentenceTransformer(embedding_model)
  • Absent: revision=, local_files_only, weight sha256 verification, fixed device/dtype, torch.use_deterministic_algorithms, seeds (Python/NumPy/Torch), PYTHONHASHSEED, deterministic cuDNN.
  • Surfaces: (a) upstream model-repo mutation; (b) CPU/CUDA float variation; (c) tokenization/embedding drift across versions.

2. Numeric/DL deps unpinned (>=, not ==)

  • pyproject.toml: numpy>=2.4.6, pandas>=3.0.3, scipy>=1.17.0, scikit-learn>=1.9.0, dl = ["sentence-transformers>=5.5.1"].
  • A future reinstall (pipx) may resolve different versions → different tokenization/embeddings/math.

3. Silent fallback changes the pipeline with no signal

  • core/dl_backend.py:85,91-100 → init wrapped in try/except Exception sets self._ready = False (does not raise).
  • core/detector.py:1225-1227 → constructs DLClassifier; drops it when not is_ready.
  • core/detector.py:1438,1516 → DL used only if self._dl_classifier.is_ready.
  • Result: host A (model loads) runs regex+ml+DL; host B (load fails) runs regex+ml. Both complete; results may differ — unrecorded. Legitimate graceful degradation, but it must appear in the evidence contract.

4. Evidence does not record pipeline identity (repo-wide)

  • Absent across the repo: backends_active, regex_backend, ml_backend, embedding_revision, embedding_sha256, degradation_mode, dependency_lock (only reason_code exists, in core/extras_runtime.py, for extras).
  • An operator comparing two runs cannot tell from the evidence that they compared different pipelines.

5. Current gate neither exercises DL nor validates cross-host output parity

  • Gate scripts are supply-chain/security (grype, docker-scout, trailer-attest, tripwire) — not output parity across hosts.
  • The dl extra does not appear in scripts/ or .github/ → the gate does not install/exercise DL. Cross-host DL reproducibility is unvalidated.

Impact on the public claim

We cannot certify "same input + YAML → same bytes" for the DL path across machines/reinstalls/dates. The deterministic core (regex/rules) holds. Public messaging ("byte-for-byte", "deterministic AI" spanning ML/DL) overreaches until the DL contract closes — the marketing site is being adjusted in parallel to "reproducible under pinned config; no LLM decides findings on the critical path".

The test bench already exists: Maestro + LabOp

The required "subprocess × reinstall × platform, byte-for-byte" test is a Maestro job over the min-spec gate hosts (alpine musl/no-AVX, T14 glibc/AVX, pi3b aarch64, Void, Zorin). alpine-emachines is the canary: torch/sentence-transformers most likely does not load there → silent fallback → a different pipeline than T14. Closing this contract = make the gate read the canary (the min-spec's long-intended 2nd axis, now named: reproducibility contract).

Acceptance criteria

  • Pin the model by immutable revision=<commit-sha>; add local_files_only for governed/air-gapped runs; verify weight sha256 before load.
  • Fix device (cpu default) and dtype (float32); torch.use_deterministic_algorithms(True); set Python/NumPy/Torch seeds + PYTHONHASHSEED; deterministic cuDNN when CUDA.
  • Exact-pin the dl extra + numeric deps (or a lockfile) + dependency_lock_sha256.
  • Record full pipeline identity in every evidence artifact: backends_active; regex_backend/ml_backend/dl_backend: active|unavailable; embedding_model/embedding_revision/embedding_sha256; device; dtype; config_sha256; dependency_lock_sha256. On fallback: dl_backend: unavailable, degradation_mode: regex+ml, reason_code: MODEL_LOAD_FAILED (no sensitive content/stack).
  • Test: subprocess × reinstall × platform comparing canonical output byte-for-byte (or documented tolerance) — as a Maestro job over the gate hosts.
  • Concurrency / input-ordering test (connector/worker ordering can change final output even when each classifier is deterministic).
  • Verify/record current gate behavior: does smoke run the dl extra? does evidence now record backends per host?
  • Widen automated proof of ML (TF-IDF + RF) determinism.
  • Until closed: scope the claim (docs + site) to the deterministic core; DL = "reproducible under pinned config".
  • Create docs/plans/PLAN_DL_REPRODUCIBILITY.md with <!-- plans-hub-summary: ... -->; run python scripts/plans_hub_sync.py --write; add a docs/plans/PLANS_TODO.md entry.

Not in question (still correct)

  • "No LLM / no generative / no hallucination on the critical path"
  • "No proprietary black box"
  • Regex/rules: deterministic by nature

Read-only audit (Claude Code). Evidence via gh api on the authoritative repo. Implementation: Cursor.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

Projects

No projects

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions