Repository navigation
docs(audit): close out Milestone 2 against what the repository does - #48
Merged
Merged
Conversation
The live-R guard asked a different question from the record it stood in front of: it probed an empty package list, its consumer probed the four in R_PACKAGES, so on any checkout where R answers but the packages are absent -- every bare checkout -- the guard reported ready and the consumer threw outside any try, erroring the testset before its first @test (#43). The @info branch could not fire for the case it was written for. provenance.jl now answers the question once. RProbeStatus carries a status (:ok, :packages_missing, :r_unreachable, :no_lockfile) beside the record, and r_probe_status performs the single probe whose record the suite then asserts on, so guard and consumer cannot diverge. The failure's component names travel as a ProbeFailure field rather than as prose, so classification is a field read, and probe_r collects every absent package instead of stopping at the first: four missing packages report four. The skip is loud and named: the testset is titled with the status and the reason -- "r - not probed (packages_missing): ... dada2, Biostrings, ShortRead, vegan" -- and recorded with @test_skip, so it lands in the summary's own Broken column rather than in an @info that scrolls past. On CI, where ci.yml installs R and restores all 79 packages from renv.lock, the same block is a hard assertion: a skip there is a provisioning regression, not a bare checkout. The negative control is pure and runs everywhere the live probe skips: :r_unreachable, :packages_missing and :no_lockfile are asserted to be distinguishable in both symbol and text. Verified against the real module on a box with R 4.5.0 and none of the four packages: the :packages_missing path is exercised live, names all four, and records Broken rather than passing or erroring. Closes #43
Issue #15 asks for the Milestone 2 record to be audited. Re-ran the claims rather than reading them back, and two of them no longer described the repository: - the ">10% regression gate fails CI" claim, for the frontend script and for each bench/*.jl. The gate was built, then deliberately replaced by an informational comparison: ci.yml records two consecutive runs of identical benchmark code producing -16%..+52% deltas, because the workloads import no application code and so measure the runner. bench/table_loading/benchmark.jl prints NOTE and warns instead of failing. Corrected, with the reasoning kept. - the test-suite figures: 586 frontend tests across 16 files are now 604 (599 pass, 5 todo, 0 fail, 3368 expects) and 27 Julia unit files / 6830 lines are now 30 / 8310. Corrected as dated audit notes beside the original snapshot rather than overwritten, so the record shows both what was measured then and what holds now. Verified and left as recorded: coverage 40.23% funcs / 47.50% lines (claim 40/47), all five benchmark categories plus the comprehensive runner and analysis_config, the seven frontend workloads one-for-one, Codecov residue gone and gitar absent, CI runtime 14-29 min (claim 15-20). The project board is recorded as unverifiable with the credentials this audit had, not assumed. docs/audit/milestone-2-close-out.md carries the claim-by-claim evidence and the reproducible commands behind each verdict. Closes #15
|
|
Warning Review limit reachedNext included review available in 56 minutes. View limit detailsLimit details: You’ve used the included review currently available. You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository. Review configuration: ⚙️ Run configurationConfiguration used: Organization UI Review profile: ASSERTIVE Plan: Advanced Run ID: 📒 Files selected for processing (4)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.



Closes #15.
Audited the Milestone 2 record by re-running its claims instead of reading them back.
Verdict: all four acceptance criteria hold. Two claims in the milestone document
no longer described the repository, and both are corrected here:
ci.ymlrecords why, with the measurement: two consecutive runs of identicalbenchmark code produced per-workload deltas between -16% and +52%, because the
harness workloads import no application code — a delta measures the runner, not the
commit.
bench/table_loading/benchmark.jlagrees (NOTE+@warn, "Never gates inCI"). A gate below the noise floor blocks at random, and a gate that fails for
reasons the commit cannot influence trains people to ignore it. What still holds
for real: checksum hard-fails, the workload freeze policy, and deltas + machine
factor shipped as artifacts.
5 todo, 0 fail, 3368 expects, 461 ms); Julia 30 unit files / 8310 lines (from
27 / 6830). Corrected as dated notes beside the original snapshot rather than
overwritten, so both the then-measurement and today's hold together.
Verified and left as recorded: coverage 40.23% funcs / 47.50% lines (the
issue's 40/47), all five
bench/categories plus the comprehensive runner andanalysis_config, the seven frontend workloads matching the document one-for-one,Codecov residue gone and
gitarabsent, and CI runtime 14–29 min against the claimed15–20 — with the renv restore (79 R packages from source) as the variance driver. The
cladistics category is precise already:
tree-rendering-clade-cumulusandCladeCumulus.tsxexist,test_clade_cumulus.jlremains future work as stated.Recorded as unverified rather than assumed: the project board
(users/hyperpolymath/projects/45) is a user-level Projects v2 board, which needs
read:project— not granted to the token this audit used. Everything else on theaudit's table was re-derived from the repository or the API.
docs/audit/milestone-2-close-out.mdcarries the commands and output behind everyverdict. No code change is implied by the audit: the tests, benchmarks and CI wiring
are as the milestone describes them.