Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
39 commits
Select commit Hold shift + click to select a range
c3f2340
Add abstaining finite-sample design stress test
FlorianPfaff Aug 25, 2026
794cd04
Bind stress budgets to the frozen signal-to-noise target
FlorianPfaff Aug 25, 2026
810c12a
Add checksum-bound design stress artifact runner
FlorianPfaff Aug 25, 2026
9c0701d
Reject unused certified stress allocations
FlorianPfaff Aug 25, 2026
e6d624a
Test calibrated abstention stress and artifact provenance
FlorianPfaff Aug 25, 2026
d12f49d
Document the bounded post-freeze design sensitivity
FlorianPfaff Aug 25, 2026
6b1ff21
Expose the design stress artifact command
FlorianPfaff Aug 25, 2026
9569d56
Fix slot-safe stress gate test construction
FlorianPfaff Aug 25, 2026
36b5b75
Add a bounded out-of-span nonlinear stress probe
FlorianPfaff Aug 25, 2026
9bbf34d
Bind the nonlinear probe table into stress artifacts
FlorianPfaff Aug 25, 2026
ac0c123
Resolve stress quality diagnostics
FlorianPfaff Aug 25, 2026
6dd9d6f
Cover the nonlinear stress probe and artifact table
FlorianPfaff Aug 25, 2026
2ce8bb0
Scope the fixed nonlinear misspecification probe
FlorianPfaff Aug 25, 2026
5ec894b
Harden frozen stress configuration validation
FlorianPfaff Aug 25, 2026
b00eae9
Test rejection of ill-defined stress settings
FlorianPfaff Aug 25, 2026
c0a8474
Make the locked equal-N60 benchmark the primary stress schedule
FlorianPfaff Aug 25, 2026
ba68a3f
Normalize fixed-budget stress metadata
FlorianPfaff Aug 25, 2026
b06bc6f
Bind the chronologically locked N60 allocation
FlorianPfaff Aug 25, 2026
9816310
Test checksum-bound locked N60 allocations
FlorianPfaff Aug 25, 2026
d4704b3
Document the chronologically locked N60 primary stress
FlorianPfaff Aug 25, 2026
36c34fe
Normalize fixed-budget stress test configuration
FlorianPfaff Aug 25, 2026
f12d396
Record primary stress schedule chronology
FlorianPfaff Aug 25, 2026
a044478
Normalize locked-allocation test imports
FlorianPfaff Aug 25, 2026
513a3ef
Wrap fixed-budget validation
FlorianPfaff Aug 25, 2026
29426f9
Sort locked stress imports
FlorianPfaff Aug 25, 2026
f3c3a50
Wrap locked allocation provenance validation
FlorianPfaff Aug 25, 2026
0a0a5fb
Sort design stress imports
FlorianPfaff Aug 25, 2026
50f1d7f
Annotate locked stress allocations
FlorianPfaff Aug 25, 2026
b6755d1
Validate stress caps by design
FlorianPfaff Aug 25, 2026
3ff6b86
Verify frozen stress constructors
FlorianPfaff Aug 25, 2026
5676b9f
Test frozen stress construction contract
FlorianPfaff Aug 25, 2026
3438d5f
Document design-specific stress caps
FlorianPfaff Aug 25, 2026
b41f123
Annotate reconstructed stress allocations
FlorianPfaff Aug 25, 2026
97f07ff
Bind frozen stress allocation seed
FlorianPfaff Aug 25, 2026
b0c2097
Test explicit stress seed provenance
FlorianPfaff Aug 25, 2026
c71695f
Document explicit frozen allocation seed
FlorianPfaff Aug 25, 2026
d1251dd
Freeze locked N60 open-set stress boundary
FlorianPfaff Aug 25, 2026
88aa601
Add post-freeze stress geometry audit
FlorianPfaff Aug 25, 2026
183233e
Freeze post-stress geometry diagnostic
FlorianPfaff Aug 25, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions .gitattributes
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
results/design-open-set-stress-n60/*.csv -text -diff
31 changes: 31 additions & 0 deletions docs/design_open_set_stress_n60_result.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,31 @@
# Frozen locked-N60 open-set stress result

The immutable package in `results/design-open-set-stress-n60/` was produced from clean
commit `c71695fda83ae93407599a909097962ee3fa9e0e`. It byte-verifies and reconstructs the
chronologically locked equal-N60 allocation (SHA-256
`a823be49faf6c6cbebf60b11d4b5ca895cf7734d6e9c577ee98f97a5907b69b2`) from source
commit `1b2028929ac6ebc1cce0882f0c22af9918044342` and explicit allocation seed 7.

The frozen settings use 100 calibration, 100 calibration-audit, and 200 evaluation
replicates; threshold, audit, and evaluation seeds are 104729, 130363, and 155921. The
held-out fraction is 0.35. No threshold or evaluation was changed after inspecting results.

| Allocation | Weakest pure rate (Wilson lower; raw closed-set) | Worst 50/50 mixture false-pure rate (upper) | Null false-pure rate (upper) | Nonlinear probe false-pure rate (upper) |
|---|---:|---:|---:|---:|
| Coupled novelty | 0.030 (0.0138; 0.425) | 0.460 (0.5292) | 0.010 (0.0357) | 0.035 (0.0705) |
| Uniform factorial | 0.425 (0.3585; 0.720) | 0.545 (0.6125) | 0.015 (0.0432) | 0.005 (0.0278) |
| Locked heuristic maximin | 0.750 (0.6857; 0.820) | 0.800 (0.8495) | 0.020 (0.0503) | 0.515 (0.5833) |

For the locked heuristic maximin allocation, the pure-over-null, winner-over-runner, and
flexible-over-pure thresholds are 1.4684378147, 0.7589681024, and 2.0743104877. Its
worst mixture is gain plus update, while its weakest pure generator is surprise.

This is failure-boundary evidence, not open-set robustness. Null control is bounded, and
the locked maximin schedule retains reasonable matched-pure performance, but the current
adequacy rule frequently labels mixtures and the single declared nonlinear probe as pure.
The artifact does not justify claims for arbitrary mixtures, nonlinear alternatives,
serial dependence, sensor dynamics, subject hierarchy, or physical protocol feasibility.

The package checksum-table SHA-256 is
`44a5188c43bda52e6fc9dc7007cf2de44a9671e9c5477ac88c3173c06cfdbd80`; the manifest
SHA-256 is `d840a2ec34f5a386109c7f985033b53144fdcc1b1a0e0592b3274c0e38902b64`.
124 changes: 124 additions & 0 deletions docs/design_stress.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,124 @@
# Post-freeze design abstention stress

This module is a **versioned sensitivity analysis**, not part of the immutable
five-seed matched-generator evidence. It asks whether a deliberately conservative
classifier can avoid confident pure-candidate calls when the generator is null or
is an unmodelled 50/50 combination of two candidates. It does not turn synthetic
recovery into biological evidence.

## Frozen schedule

The primary run evaluates the chronologically earlier, equal-budget 60-draw
allocation frozen for the five-seed paper benchmark. The CLI verifies the exact
allocation-file SHA-256, source-code commit, and explicit seed metadata, loads all
three designs from that file, and reconstructs all three deterministic constructors
before accepting the counts. The accepted CSV does not itself contain a seed
column, so the seed is a required CLI input recorded in provenance rather than an
invented file field. The accepted summary did not serialize the optimizer cap, so this contract
also records the source-code setting: the maximin constructor uses
`max_point_fraction=0.15` (a cap of 9 at N=60), while the cap does not apply to
the coupled-novelty or uniform-factorial comparators. In the frozen file the
observed maxima are 8, 12, and 1 respectively. Applying the optimizer cap to the
novelty comparator would therefore change the accepted benchmark rather than
validate it. Hash, seed, constructor, or unused-override mismatches are rejected.
This prevents a later certificate or stress result from silently changing the
primary schedule.

The artifact still reports each design's population observation-equivalent index,
recomputed from its 60-draw geometry using

```math
G(R)=\tfrac12\log\!\left(1+a^2R/\sigma^2\right).
```

With standardized generating signals, `a=1`, `sigma=1`, and target held-out
log-score gap 5, those indices are 45 for the heuristic maximin design, 93 for
uniform factorial, and 1,113 for coupled novelty. The common primary budget of
60 is not asserted to equal any one of those indices. Both are effectively
independent Gaussian-observation counts, not physical trials, fluorescence
samples, sessions, or animals.

Optional `budget_factors` can generate separately labeled 0.5/1/2-type
sensitivity schedules, and a checksum-bound certified integer allocation can
replace an exactly matching optional maximin budget. Neither is part of the
primary frozen run. In particular, a certified-N45 diagnostic completed before
this primary freeze remains outside the primary artifact and was not used to
tune its thresholds. A count vector over independently instantiated grid cells
is still an allocation target, not an executable ordered history: no reset,
washout, carry-over, or history-realization protocol is provided here.

## Train-only scoring and abstention

Every replicate is randomly divided into training and held-out observations.
Each pure candidate is fitted with an intercept, slope, and training-residual
variance. The flexible adequacy model contains an intercept and all six
candidates. Its ridge penalty is selected only inside the training set by
three-fold cross-validation; its residual variance is also fitted on training
data. All comparisons are held-out profiled Gaussian log-score differences.

Three inequalities are required for a pure call:

1. the best pure candidate must beat the intercept-only model by more than the
upper null-calibration quantile. Because the statistic already maximizes over
all six pure candidates, this threshold is familywise for that candidate set;
2. the best pure candidate must beat the runner-up by more than a separately
null-calibrated ambiguity threshold;
3. the all-six flexible model may not beat the best pure model by more than the
worst-candidate pure-generator adequacy threshold.

The third rule is one-sided. The flexible model nests the pure model, so the two
population scores tie under a correctly specified pure generator; requiring the
pure model to beat the flexible model would be invalid. The adequacy threshold is
the largest conformal upper threshold across the six matched pure generators,
which protects the worst calibrated pure candidate at the declared finite
calibration resolution.

Threshold calibration, calibration audit, and final evaluation use three
disjoint deterministic seeds. The artifact reports the independent calibration
audit rather than treating threshold-training performance as validation.

## Scenarios and interpretation

Matched pure generators are evaluated with correct-call, wrong-call, abstention,
and raw closed-set winner rates. The null reports any non-null pure call as a
false call. All 15 unordered 50/50 candidate pairs are evaluated after equal
coefficients and unit-standard-deviation scaling on the full feasible grid. Any
pure call for a mixture—including a call naming one of its constituents—is
counted as false. Pointwise Wilson intervals accompany all reported rates; they
are not simultaneous confidence bounds over the many scenario cells.

The mixture family is open-set relative to the six pure labels but remains inside
their linear span. One additional, deliberately narrow out-of-span probe takes
`tanh(standardized surprise)`, removes its full-grid OLS projection on an
intercept and all six standardized candidates, and scales the residual to unit
standard deviation. Orthogonality is checked numerically on the full grid; a pure
call is false. This is a diagnostic of one fixed saturation-shaped residual, not
coverage of a biologically defined misspecification class.

The bounded artifact therefore does **not** certify robustness to arbitrary
out-of-span misspecification, other nonlinear/saturating combinations, serial
dependence, subject hierarchy, indicator dynamics, nuisance mismatch, or invalid
sequential histories. Matched-field simulation also cannot diagnose
misspecification of the candidate signals themselves.

## Reproducible artifact

From a clean checkout of the exact stress commit:

```bash
bayesian-ach-design-stress \
--repo-root . \
--code-sha <40-character-stress-commit> \
--output /absolute/path/design-stress-n60 \
--fixed-budgets 60 \
--locked-allocation /absolute/path/optimal_design_allocation_seed7.csv \
--locked-allocation-sha256 <frozen-64-character-sha256> \
--locked-design-code-sha <40-character-design-commit> \
--locked-allocation-seed 7
```

The command refuses a dirty or mismatched checkout. It writes the configuration,
thresholds, independent calibration audit, pure/null/mixture and fixed nonlinear-probe evaluations,
allocations, a provenance manifest, and `SHA256SUMS.csv`. The manifest binds the
producer commit, canonical configuration digest, every payload file, and every
supplied certificate package.
21 changes: 21 additions & 0 deletions docs/design_stress_geometry_diagnostic.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,21 @@
# Post-freeze stress geometry diagnostic

This explanatory diagnostic is downstream of the immutable locked-N60 stress package. It
does not alter its thresholds, seeds, allocations, simulations, or endpoints.

`scripts/analyze_design_stress_geometry.py` stratifies the frozen result by allocation and
replays the 15 maximin mixture evaluations to count each decision gate independently. For
each 50/50 pair it also computes, on the locked maximin support, the weighted affine
residual against the best constituent and best pure candidate, the oracle two-component
residual, candidate correlation, covariance condition number, and the population profiled
Gaussian gap for the frozen 21-sample held-out size.

`scripts/verify_design_stress_geometry.py` checks the immutable source-package digest,
producer/script provenance, every payload checksum, pair count, gate/false-call identities,
and the zero-residual two-component oracle. It then reruns all stratification, geometry, and
15 x 200 frozen evaluations and requires byte-identical JSON, CSV, and manifest outputs.

The diagnostic explains whether poor mixture rejection reflects population aliasing on the
locked support or finite-sample decision behavior. It is not a new endpoint, a robustness
claim, a threshold-tuning analysis, or permission to select a replacement gate after seeing
the frozen result.
1 change: 1 addition & 0 deletions pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -70,6 +70,7 @@ plot = ["matplotlib>=3.8"]
[project.scripts]
bayesian-ach = "bayesian_ach.cli_ext:main"
bayesian-ach-design = "bayesian_ach.design_cli:main"
bayesian-ach-design-stress = "bayesian_ach.design_stress_cli:main"
bayesian-ach-replay = "bayesian_ach.replay_cli:main"

[tool.setuptools.packages.find]
Expand Down
4 changes: 4 additions & 0 deletions results/design-open-set-stress-n60-geometry/SHA256SUMS.csv
Original file line number Diff line number Diff line change
@@ -0,0 +1,4 @@
file,bytes,sha256
stress_diagnostic.json,20761,854b97eae7c873bdfc2de3b8145003028b570c65cca3774cedb32827f8fc4a2a
maximin_mixture_geometry.csv,6752,8d179c5c5ae2a1623c955e96d00bb86ea510be0a91d2326870de6ca367e144f7
artifact_manifest.json,721,e248c92fae66931fa0fbf7a108d8bb2fe82572949c47f89e85365bed28286372
20 changes: 20 additions & 0 deletions results/design-open-set-stress-n60-geometry/artifact_manifest.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,20 @@
{
"artifact": "post_freeze_stress_geometry_diagnostic",
"diagnostic_script_sha256": "31ff6397b5a4bd1dbdf0b83f578c0b24e8e0e36f34058758d92aa4094ec89bb7",
"files": [
{
"bytes": 20761,
"path": "stress_diagnostic.json",
"sha256": "854b97eae7c873bdfc2de3b8145003028b570c65cca3774cedb32827f8fc4a2a"
},
{
"bytes": 6752,
"path": "maximin_mixture_geometry.csv",
"sha256": "8d179c5c5ae2a1623c955e96d00bb86ea510be0a91d2326870de6ca367e144f7"
}
],
"producer_commit": "88aa601dcf65a1cecf20d388ecdb2e304f6c4116",
"producer_git_dirty": false,
"schema_version": 1,
"source_artifact_sha256sums_sha256": "44a5188c43bda52e6fc9dc7007cf2de44a9671e9c5477ac88c3173c06cfdbd80"
}
Original file line number Diff line number Diff line change
@@ -0,0 +1,16 @@
first_candidate,second_candidate,false_pure_call_rate,wilson_upper,sequential_reason_counts,independent_gate_pass_counts,best_constituent,best_constituent_residual,best_any_pure,best_any_pure_residual,oracle_two_component_residual,best_pure_profiled_gap_n_test_21,oracle_two_component_profiled_gap_n_test_21,weighted_candidate_correlation,weighted_pair_covariance_condition_number
innovation_l2,surprise,0.745,0.8003950482796865,"{""flexible_model_better"":3,""null_not_rejected"":3,""pure_ambiguity"":45,""pure_call"":149}","{""all_three_pass"":149,""flexible_adequacy_pass"":197,""pure_over_null_pass"":197,""winner_over_runner_pass"":153}",surprise,0.06476225526808417,surprise,0.06476225526808417,1.7834111285025804e-31,0.6588911673833898,1.8725816849277094e-30,0.8467944254803421,12.054612191382652
innovation_l2,gain,0.555,0.6221939899704665,"{""flexible_model_better"":31,""null_not_rejected"":6,""pure_ambiguity"":52,""pure_call"":111}","{""all_three_pass"":111,""flexible_adequacy_pass"":151,""pure_over_null_pass"":194,""winner_over_runner_pass"":143}",gain,0.44165477400362824,update_l2,0.23773950726644638,7.233792871118458e-32,2.239510748574044,7.595482514674381e-31,-0.02131984334093419,1.7613618463469893
innovation_l2,update_l2,0.62,0.6844099415132637,"{""flexible_model_better"":25,""null_not_rejected"":1,""pure_ambiguity"":50,""pure_call"":124}","{""all_three_pass"":124,""flexible_adequacy_pass"":160,""pure_over_null_pass"":199,""winner_over_runner_pass"":149}",update_l2,0.2576905807957704,update_l2,0.2576905807957704,5.256248004221645e-31,2.4074102515385984,5.519060404432727e-30,0.4213845233433928,2.823121945397039
innovation_l2,information_gain,0.58,0.6462640083838327,"{""flexible_model_better"":16,""null_not_rejected"":3,""pure_ambiguity"":65,""pure_call"":116}","{""all_three_pass"":116,""flexible_adequacy_pass"":170,""pure_over_null_pass"":197,""winner_over_runner_pass"":133}",information_gain,0.2286673248345743,information_gain,0.2286673248345743,2.511402930060989e-32,2.162266115687133,2.636973076564038e-31,0.4586862807960559,2.8800745735556
innovation_l2,change_probability,0.615,0.6796669027307678,"{""flexible_model_better"":26,""null_not_rejected"":3,""pure_ambiguity"":48,""pure_call"":123}","{""all_three_pass"":123,""flexible_adequacy_pass"":157,""pure_over_null_pass"":197,""winner_over_runner_pass"":150}",change_probability,0.2715733105685233,change_probability,0.2715733105685233,3.425330603757512e-31,2.522677090253475,3.596597133945387e-30,0.3569282136837277,2.4570981315526876
surprise,gain,0.485,0.5538915080725428,"{""flexible_model_better"":22,""null_not_rejected"":2,""pure_ambiguity"":79,""pure_call"":97}","{""all_three_pass"":97,""flexible_adequacy_pass"":151,""pure_over_null_pass"":198,""winner_over_runner_pass"":120}",gain,0.44477271898790127,update_l2,0.3334070092601476,4.191979116953257e-31,3.02124194263797,4.40157807280092e-30,-0.043754286693241884,1.755875929808455
surprise,update_l2,0.575,0.6414638745648614,"{""flexible_model_better"":13,""null_not_rejected"":1,""pure_ambiguity"":71,""pure_call"":115}","{""all_three_pass"":115,""flexible_adequacy_pass"":174,""pure_over_null_pass"":199,""winner_over_runner_pass"":128}",update_l2,0.28827138069547065,update_l2,0.28827138069547065,3.0044507132440885e-32,2.6596637002520334,3.1546732489062928e-31,0.32590264303136707,2.3238272821220116
surprise,information_gain,0.72,0.7776311022605097,"{""flexible_model_better"":11,""null_not_rejected"":2,""pure_ambiguity"":43,""pure_call"":144}","{""all_three_pass"":144,""flexible_adequacy_pass"":181,""pure_over_null_pass"":198,""winner_over_runner_pass"":157}",information_gain,0.1751311847924566,information_gain,0.1751311847924566,2.2232935278006252e-31,1.6944877739577782,2.3344582041906565e-30,0.6044310151366707,4.259069867164642
surprise,change_probability,0.67,0.7314257412939605,"{""flexible_model_better"":39,""null_not_rejected"":5,""pure_ambiguity"":22,""pure_call"":134}","{""all_three_pass"":134,""flexible_adequacy_pass"":158,""pure_over_null_pass"":195,""winner_over_runner_pass"":174}",change_probability,0.24059459272699615,change_probability,0.24059459272699615,1.4359733665351229e-31,2.263703136999608,1.507772034861879e-30,0.4498740075388434,2.9744407398573025
gain,update_l2,0.8,0.8495479907390189,"{""flexible_model_better"":5,""null_not_rejected"":1,""pure_ambiguity"":34,""pure_call"":160}","{""all_three_pass"":160,""flexible_adequacy_pass"":192,""pure_over_null_pass"":199,""winner_over_runner_pass"":165}",gain,0.1515439180125079,gain,0.1515439180125079,7.5579912038097065e-31,1.4815875834877898,7.935890764000191e-30,0.791530832607952,8.604108451760718
gain,information_gain,0.61,0.6749165686253038,"{""flexible_model_better"":20,""pure_ambiguity"":58,""pure_call"":122}","{""all_three_pass"":122,""flexible_adequacy_pass"":166,""pure_over_null_pass"":200,""winner_over_runner_pass"":142}",gain,0.3094561867702409,update_l2,0.2454971557070175,2.429239636520434e-32,2.3051151066504176,2.5507016183464558e-31,0.4917349601612237,3.0003291825845855
gain,change_probability,0.27,0.33543435198668425,"{""flexible_model_better"":75,""null_not_rejected"":5,""pure_ambiguity"":66,""pure_call"":54}","{""all_three_pass"":54,""flexible_adequacy_pass"":82,""pure_over_null_pass"":195,""winner_over_runner_pass"":129}",gain,0.7194446631198409,gain,0.7194446631198409,2.8000132496473827e-31,5.691014368329873,2.940013912129752e-30,-0.0444760730140199,1.1234288396957601
update_l2,information_gain,0.745,0.8003950482796865,"{""flexible_model_better"":5,""null_not_rejected"":3,""pure_ambiguity"":43,""pure_call"":149}","{""all_three_pass"":149,""flexible_adequacy_pass"":191,""pure_over_null_pass"":197,""winner_over_runner_pass"":154}",update_l2,0.13734921104922593,update_l2,0.13734921104922593,7.415215471879735e-31,1.3513531640796759,7.785976245473722e-30,0.7713790932333591,7.802793892925121
update_l2,change_probability,0.325,0.39268014883970037,"{""flexible_model_better"":68,""null_not_rejected"":2,""pure_ambiguity"":65,""pure_call"":65}","{""all_three_pass"":65,""flexible_adequacy_pass"":100,""pure_over_null_pass"":198,""winner_over_runner_pass"":134}",update_l2,0.6095014120398539,update_l2,0.6095014120398539,6.510927165325115e-32,4.997206715257962,6.836473523591371e-31,0.09440005532043154,1.2090358929226095
information_gain,change_probability,0.535,0.6028143584299599,"{""flexible_model_better"":53,""null_not_rejected"":5,""pure_ambiguity"":35,""pure_call"":107}","{""all_three_pass"":107,""flexible_adequacy_pass"":131,""pure_over_null_pass"":195,""winner_over_runner_pass"":161}",change_probability,0.4585249642762335,change_probability,0.4585249642762335,8.693262582975106e-31,3.962969079130324,9.127925712123861e-30,0.22817185530184655,1.6223418188593566
Loading