Skip to content

Select the first bounded computational-materials benchmark - #3

Merged
chefmrfrizzle merged 2 commits into
mainfrom
agent/select-first-benchmark
Aug 10, 2026
Merged

Select the first bounded computational-materials benchmark#3
chefmrfrizzle merged 2 commits into
mainfrom
agent/select-first-benchmark

Conversation

@chefmrfrizzle

Copy link
Copy Markdown
Owner

What changed

  • compares four safe public or synthetic computational-materials benchmark candidates using current primary sources
  • adds a transparent weighted scorecard with sourced facts, local measurements, assumptions, estimates, and rights stop-conditions
  • recommends exactly one narrow workload: pinned spglib 2.7.0 wurtzite symmetry classification
  • defines deterministic logical inputs and discrete outputs, challenge cases, pass/fail rules, a reproducibility budget, failure cases, and required domain expertise
  • proposes ADR-0010 without adopting it
  • records D-021/D-022 as blocked with accountable owners and required evidence
  • records LearningEpisode 0004, including the C/Python lattice-convention counterexample

Why

Prompt 2 requires a legally safe, low-cost workload that can demonstrate deterministic execution, verification, replay, counterexamples, and second-machine feasibility before any implementation begins.

The selected candidate uses a synthetic four-site input and exposes a realistic failure: copying the upstream C lattice directly into the Python row-vector API returns a plausible but incorrect classification.

Scope and governance

This PR contains public research/planning documentation only. It does not implement schemas, adapters, workers, signing, payment, settlement, production infrastructure, or scientific acceptance policy.

Disclosure: @chefmrfrizzle and requested reviewer @11BUSD share one human operator. Any @11BUSD approval is separate-account workflow evidence under ADR-0009, not independent-person or domain-expert assurance. ADR-0010 remains proposed; D-018, D-019, D-021, and D-022 remain blocked.

Validation

  • ./tests/test_repository_baseline.sh
  • BASE_SHA=origin/main HEAD_SHA=HEAD ./scripts/check_repository_baseline.sh
  • git diff --check
  • all 26 primary-source links checked; 25 returned HTTP 200 and one publisher DOI endpoint returned HTTP 403 to the automated client while remaining a valid DOI
  • local discovery run: macOS arm64, CPython 3.9.6, spglib 2.7.0, NumPy 2.0.2; 1,000 repeated classifications produced one selected result

Review focus

  • scientific defensibility of the claim boundary and perturbation cases
  • whether sourced facts, measurements, assumptions, and estimates remain clearly separated
  • whether dataset/code licensing conclusions are appropriately conservative
  • whether D-021/D-022 block implementation until domain and second-machine evidence exist

@chefmrfrizzle
chefmrfrizzle requested a review from 11BUSD August 10, 2026 16:49
@chefmrfrizzle chefmrfrizzle self-assigned this Aug 10, 2026
@chefmrfrizzle
chefmrfrizzle marked this pull request as ready for review August 10, 2026 16:49
@chefmrfrizzle
chefmrfrizzle merged commit d379b76 into main Aug 10, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants