Public repository: https://github.com/rdinizcal/Diagnosis-replication-package
This package supports three distinct activities:
- Computational reproduction regenerates paper tables and figures from the
selected archived outputs in
raw_results/without invoking a checker. RQ1 is recomputed from per-seed trees and deterministic Cartesian grids; RQ2 is recomputed from the archived run reports and scheduler records. - Full rerunning executes the 480 materialized experiment/seed configurations against the bundled fixed inputs.
- Independent replication uses the registries, data/provenance records, semantics, tool snapshot, environment lock, and runbook to reimplement or vary the method without relying on undocumented source behavior.
The authoritative selection policy is the completed reconciliation. RQ1/RQ2
uses campaign 20260803171759, with declared per-experiment overrides to
20260812103600. RQ3 uses the 140 selected historical sensitivity runs. The
obsolete outputs, online_results, 20260728/20260729 campaigns, and nested
historical exp1..exp34 results are not imported.
python3 scripts/preflight.py # reports host readiness; no changes
./smoke # schema/parser/mutation/checker/cache/tree
./reproduce-analysis # writes derived_results/
./verify # validate evidence and analysis oracles
./reproduce-full # dry-run inventory
./reproduce-full --profile at1 # short single-token profile (exp1/seed0)
./reproduce-full --profile multi-token # exp2/seed0
./reproduce-full --profile discrete-operator # exp30/seed0
./reproduce-full --execute # expensive: authorized primary runsSee RUNBOOK.md for container and selective-run commands, docs/RQ1_RQ2.md
for the exact effectiveness/efficiency methodology, CLAIMS.md for the
claim-to-artifact map, experiments/registry.csv for canonical exp1–exp34,
and experiments/auxiliary_registry.csv for RQ3 and all auxiliary studies.
ARTIFACT_EVALUATION.md records the clean-room validation environment,
commands, and outcomes.
The selected RQ1/RQ2 runs captured effective JSON configurations but not their
tool commit or image digest. The package therefore records UNRECORDED and
separately identifies 05d0750395972662e9be0a1083b09a6dce7e3e65 as the
rerun source base. The bundled tool applies the documented bounded-quantifier
guard correction in tool/PATCHES.md. RQ3 did not capture
effective configs either; its 140 materialized configs are evidence-based
reconstructions and are not represented as byte-identical reruns of the archive.
They preserve the historical position namespace and provide translated current
positions for execution. Archived AT53 runs used legacy binder-only quantifier
flips; the current tool instead repairs range guards.
Optional auxiliary material is inventoried separately from the core RQ1--RQ3 reproduction: checker-swap timing sources, external-baseline case metadata, and MS-sensitivity configurations. Their recorded availability is explicit and is not used to assert unsupported derived results.
Public interfaces use only canonical paper IDs. Values such as AT1/exp1 and
AT1/exp3 survive solely in historical_internal_alias for auditability.
data/: immutable requirements, 17 benchmark traces (10 AT rows, 6 CC rows, and RR), running-example inputs, and provenance.evaluation_inputs/: historical launch templates and requirement sources referenced by the archived RQ1/RQ2 configurations, plus the RQ3 sensitivity templates retained to audit the reconstructed configurations.raw_results/rq1_rq2/: the reconciliation-selected evidence for all 340 RQ1/RQ2 run slots, including configurations, generations, ARFFs, raw J48, caches, stopping records, logs, and scheduler metadata.derived_results/: disposable outputs regenerated by scripts.expected/: immutable machine-readable comparison oracles.tool/: archive-safe executable subset of the declared rerun source base, the documented quantifier correction, and package launch/semantics files; it is not an assertion about the unknown run commit.Diagnosis/: Git submodule pinned to the complete declared rerun source at05d0750395972662e9be0a1083b09a6dce7e3e65. It is provenance material; the self-contained release archive usestool/so it does not depend on a hosting service or submodule expansion.references/expert_trees/: 34 machine-readable expert trees, original evaluator scripts, trace witnesses where recoverable, and vector renderings.experiments/operator_coverage/: executable OP1--OP15/OP4a fixtures, full source/AST bundles, expected verdicts, and implementation/coverage matrix.experiments/implementation_fix/: exact faulty/fixed controllers, generated patch, eight pre/post traces, detailed verdicts, and regression tests.
The normalized executable configurations are under
experiments/rq1_effectiveness/configs/ (340 RQ1/RQ2 seed configurations) and
experiments/rq3_sensitivity/configs/ (140 reconstructed RQ3 seed
configurations). RQ2 analyzes the same runs as RQ1 and therefore has no separate
launch-configuration set.
MANIFEST.csv is the release inventory. Every immutable file has a SHA-256;
generated/rerun targets are listed with their generation command and status.
Clone with git clone --recurse-submodules to inspect the full upstream
Diagnosis snapshot. Ordinary analysis and smoke reproduction use the bundled
tool/ subset and also work from the materialized Zenodo archive.