Context
Reproducibility audit of the detection path, prompted by review of the public "deterministic / reproducible" messaging. This is not a demonstrated divergence — it is a gap in closing the contract "same input + YAML → same evidence" for the DL path, across installs/platforms. The deterministic core (regex/rules + Clojure sidecar) is not in question.
Verified against the authoritative repo via gh api (not a local clone).
Findings
1. DL model loaded by name only — no revision/hash/local_files_only/device/dtype/determinism
core/dl_backend.py:39 → DEFAULT_EMBEDDING_MODEL = "sentence-transformers/all-MiniLM-L6-v2"
core/dl_backend.py:92 → self._embedder = _SentenceTransformer(embedding_model)
- Absent:
revision=, local_files_only, weight sha256 verification, fixed device/dtype, torch.use_deterministic_algorithms, seeds (Python/NumPy/Torch), PYTHONHASHSEED, deterministic cuDNN.
- Surfaces: (a) upstream model-repo mutation; (b) CPU/CUDA float variation; (c) tokenization/embedding drift across versions.
2. Numeric/DL deps unpinned (>=, not ==)
pyproject.toml: numpy>=2.4.6, pandas>=3.0.3, scipy>=1.17.0, scikit-learn>=1.9.0, dl = ["sentence-transformers>=5.5.1"].
- A future reinstall (pipx) may resolve different versions → different tokenization/embeddings/math.
3. Silent fallback changes the pipeline with no signal
core/dl_backend.py:85,91-100 → init wrapped in try/except Exception sets self._ready = False (does not raise).
core/detector.py:1225-1227 → constructs DLClassifier; drops it when not is_ready.
core/detector.py:1438,1516 → DL used only if self._dl_classifier.is_ready.
- Result: host A (model loads) runs regex+ml+DL; host B (load fails) runs regex+ml. Both complete; results may differ — unrecorded. Legitimate graceful degradation, but it must appear in the evidence contract.
4. Evidence does not record pipeline identity (repo-wide)
- Absent across the repo:
backends_active, regex_backend, ml_backend, embedding_revision, embedding_sha256, degradation_mode, dependency_lock (only reason_code exists, in core/extras_runtime.py, for extras).
- An operator comparing two runs cannot tell from the evidence that they compared different pipelines.
5. Current gate neither exercises DL nor validates cross-host output parity
- Gate scripts are supply-chain/security (grype, docker-scout, trailer-attest, tripwire) — not output parity across hosts.
- The
dl extra does not appear in scripts/ or .github/ → the gate does not install/exercise DL. Cross-host DL reproducibility is unvalidated.
Impact on the public claim
We cannot certify "same input + YAML → same bytes" for the DL path across machines/reinstalls/dates. The deterministic core (regex/rules) holds. Public messaging ("byte-for-byte", "deterministic AI" spanning ML/DL) overreaches until the DL contract closes — the marketing site is being adjusted in parallel to "reproducible under pinned config; no LLM decides findings on the critical path".
The test bench already exists: Maestro + LabOp
The required "subprocess × reinstall × platform, byte-for-byte" test is a Maestro job over the min-spec gate hosts (alpine musl/no-AVX, T14 glibc/AVX, pi3b aarch64, Void, Zorin). alpine-emachines is the canary: torch/sentence-transformers most likely does not load there → silent fallback → a different pipeline than T14. Closing this contract = make the gate read the canary (the min-spec's long-intended 2nd axis, now named: reproducibility contract).
Acceptance criteria
Not in question (still correct)
- "No LLM / no generative / no hallucination on the critical path"
- "No proprietary black box"
- Regex/rules: deterministic by nature
Read-only audit (Claude Code). Evidence via gh api on the authoritative repo. Implementation: Cursor.
Context
Reproducibility audit of the detection path, prompted by review of the public "deterministic / reproducible" messaging. This is not a demonstrated divergence — it is a gap in closing the contract "same input + YAML → same evidence" for the DL path, across installs/platforms. The deterministic core (regex/rules + Clojure sidecar) is not in question.
Verified against the authoritative repo via
gh api(not a local clone).Findings
1. DL model loaded by name only — no revision/hash/local_files_only/device/dtype/determinism
core/dl_backend.py:39→DEFAULT_EMBEDDING_MODEL = "sentence-transformers/all-MiniLM-L6-v2"core/dl_backend.py:92→self._embedder = _SentenceTransformer(embedding_model)revision=,local_files_only, weightsha256verification, fixeddevice/dtype,torch.use_deterministic_algorithms, seeds (Python/NumPy/Torch),PYTHONHASHSEED, deterministic cuDNN.2. Numeric/DL deps unpinned (
>=, not==)pyproject.toml:numpy>=2.4.6,pandas>=3.0.3,scipy>=1.17.0,scikit-learn>=1.9.0,dl = ["sentence-transformers>=5.5.1"].3. Silent fallback changes the pipeline with no signal
core/dl_backend.py:85,91-100→ init wrapped intry/except Exceptionsetsself._ready = False(does not raise).core/detector.py:1225-1227→ constructsDLClassifier; drops it whennot is_ready.core/detector.py:1438,1516→ DL used onlyif self._dl_classifier.is_ready.4. Evidence does not record pipeline identity (repo-wide)
backends_active,regex_backend,ml_backend,embedding_revision,embedding_sha256,degradation_mode,dependency_lock(onlyreason_codeexists, incore/extras_runtime.py, for extras).5. Current gate neither exercises DL nor validates cross-host output parity
dlextra does not appear inscripts/or.github/→ the gate does not install/exercise DL. Cross-host DL reproducibility is unvalidated.Impact on the public claim
We cannot certify "same input + YAML → same bytes" for the DL path across machines/reinstalls/dates. The deterministic core (regex/rules) holds. Public messaging ("byte-for-byte", "deterministic AI" spanning ML/DL) overreaches until the DL contract closes — the marketing site is being adjusted in parallel to "reproducible under pinned config; no LLM decides findings on the critical path".
The test bench already exists: Maestro + LabOp
The required "subprocess × reinstall × platform, byte-for-byte" test is a Maestro job over the min-spec gate hosts (alpine musl/no-AVX, T14 glibc/AVX, pi3b aarch64, Void, Zorin). alpine-emachines is the canary: torch/sentence-transformers most likely does not load there → silent fallback → a different pipeline than T14. Closing this contract = make the gate read the canary (the min-spec's long-intended 2nd axis, now named: reproducibility contract).
Acceptance criteria
revision=<commit-sha>; addlocal_files_onlyfor governed/air-gapped runs; verify weight sha256 before load.torch.use_deterministic_algorithms(True); set Python/NumPy/Torch seeds +PYTHONHASHSEED; deterministic cuDNN when CUDA.dlextra + numeric deps (or a lockfile) +dependency_lock_sha256.backends_active;regex_backend/ml_backend/dl_backend: active|unavailable;embedding_model/embedding_revision/embedding_sha256;device;dtype;config_sha256;dependency_lock_sha256. On fallback:dl_backend: unavailable,degradation_mode: regex+ml,reason_code: MODEL_LOAD_FAILED(no sensitive content/stack).dlextra? does evidence now record backends per host?docs/plans/PLAN_DL_REPRODUCIBILITY.mdwith<!-- plans-hub-summary: ... -->; runpython scripts/plans_hub_sync.py --write; add adocs/plans/PLANS_TODO.mdentry.Not in question (still correct)
Read-only audit (Claude Code). Evidence via
gh apion the authoritative repo. Implementation: Cursor.