From dc6c258fa4ff62c3bba94413aa26a55902e505df Mon Sep 17 00:00:00 2001 From: hyperpolymath <6759885+hyperpolymath@users.noreply.github.com> Date: Fri, 25 Sep 2026 23:51:12 +0000 Subject: [PATCH] feat(analysis): wire calibration in (#48); fix both red CI checks MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Two long-standing red checks on `main`, plus the blocking per-file debt behind one of them. CI --- * Hypatia's single high/critical finding was research_extensions RE008: `dependabot-automerge.yml` gated auto-merge on the dispatch actor rather than the PR author. The PR-author gate is kept, the actor conjunct is gone. The explanatory comment describes the pattern instead of quoting it, because RE008 is a raw-text scan with no comment filtering and quoting it would re-arm the finding. * `Publish Image` failed with `cargo build --release --locked -p oikosbot-cli` exit 101: tree-sitter 0.27.0 declares rust-version 1.90 / edition 2024 while the builder was rust:1.88-slim. CI on ubuntu-latest passed because the runner ships a newer rustc. Builder is now the rolling rust:1-slim. Calibration (#48) ----------------- `Analyzer::estimate_resources()` maps a detected pattern onto an `OperationKind` via the new `calibration::operation_for_pattern` and prices it with `estimate_operation()`, whose `ResourceRange` carries the min/typical/max band and the confidence that row has earned — `Calibrated` for the measured kinds, `Estimated` for the host-dependent ones. Units with no recognised pattern keep the naive complexity path, labelled `Estimated` with no band claimed. `ResourceRange` moves to `oikosbot-metrics` and rides on `AnalysisResult` (serde-default, skipped when absent), and SARIF emits it as `properties.resource_range`. `calibrated_estimate()` and its duplicated multiplier table are deleted; `standard_objectives()` is deleted as dead code. Tests ----- * calibration: per-kind confidence ladder, min<=typical<=max on every axis, the pattern→kind mapping (including the deliberate `redundant-allocation` hole). * analyzer: a recognised pattern earns Calibrated + a band; unrecognised code stays Estimated with no band; the calibrated path changes the numbers. * `crates/oikosbot-cli/tests/check_gate.rs`: end-to-end proof that `compare --check` can block — an undocumented calibrated regression exits 1, the same regression without `--check` exits 0, a documented `Pareto-Trade-off:` is accepted, and a heuristic regression refuses to block *and says so*. Docs ---- `docs/usage.adoc` and `docs/ci-runbook.adoc` (both indexed), per-crate guides, and `analyzers/code-haskell/README.adoc`. The "`--check` cannot block a merge" caveats in README/QUICKSTART/STATUS/action.yml/COMPARISON now describe what is true, and name the eco-score saturation that follows from calibrated microjoule-scale estimates. Also: delete `policy-engine/deepproblog/eco_problog.pl` (dead by the 2026-07-28 ruling) and record `instant-sync.yml`'s settings-level disabled state in its header. NOT verified in CI by this change: there is no Rust toolchain in the authoring environment, so `cargo fmt/test` results are unknown until the PR runs. Before/after figures in the docs are derived by reasoning and labelled as such. Co-authored-by: arena-agent <297053741+arena-agent@users.noreply.github.com> --- .github/workflows/dependabot-automerge.yml | 9 +- .github/workflows/instant-sync.yml | 11 + .machine_readable/descriptiles/STATE.a2ml | 10 +- CHANGELOG.adoc | 62 +++ Containerfile | 7 +- DEBT.adoc | 118 ++++-- EXPLAINME.adoc | 15 +- QUICKSTART.adoc | 24 +- README.adoc | 12 +- action.yml | 10 +- analyzers/code-haskell/README.adoc | 66 ++++ crates/oikosbot-analysis/README.adoc | 74 ++++ crates/oikosbot-analysis/src/analyzer.rs | 169 +++++++- crates/oikosbot-analysis/src/calibration.rs | 193 ++++++--- crates/oikosbot-analysis/src/dependencies.rs | 1 + crates/oikosbot-analysis/src/security.rs | 1 + crates/oikosbot-capability/README.adoc | 34 ++ crates/oikosbot-cli/README.adoc | 64 +++ crates/oikosbot-cli/src/main.rs | 2 +- crates/oikosbot-cli/tests/check_gate.rs | 171 ++++++++ crates/oikosbot-dea/README.adoc | 34 ++ crates/oikosbot-eclexia/README.adoc | 42 ++ crates/oikosbot-eclexia/src/lib.rs | 2 + crates/oikosbot-fleet/src/lib.rs | 1 + crates/oikosbot-metrics/README.adoc | 42 ++ crates/oikosbot-metrics/src/lib.rs | 26 ++ crates/oikosbot-pareto/README.adoc | 38 ++ crates/oikosbot-pareto/src/lib.rs | 14 - crates/oikosbot-sarif/README.adoc | 42 ++ crates/oikosbot-sarif/src/lib.rs | 28 ++ crates/oikosbot-telemetry/README.adoc | 41 ++ docs/COMPARISON-climate-warrior.adoc | 11 +- docs/README.adoc | 40 +- docs/STATUS.adoc | 34 +- docs/ci-runbook.adoc | 227 +++++++++++ docs/usage.adoc | 394 +++++++++++++++++++ policy-engine/deepproblog/eco_problog.pl | 208 ---------- 37 files changed, 1901 insertions(+), 376 deletions(-) create mode 100644 analyzers/code-haskell/README.adoc create mode 100644 crates/oikosbot-analysis/README.adoc create mode 100644 crates/oikosbot-capability/README.adoc create mode 100644 crates/oikosbot-cli/README.adoc create mode 100644 crates/oikosbot-cli/tests/check_gate.rs create mode 100644 crates/oikosbot-dea/README.adoc create mode 100644 crates/oikosbot-eclexia/README.adoc create mode 100644 crates/oikosbot-metrics/README.adoc create mode 100644 crates/oikosbot-pareto/README.adoc create mode 100644 crates/oikosbot-sarif/README.adoc create mode 100644 crates/oikosbot-telemetry/README.adoc create mode 100644 docs/ci-runbook.adoc create mode 100644 docs/usage.adoc delete mode 100644 policy-engine/deepproblog/eco_problog.pl diff --git a/.github/workflows/dependabot-automerge.yml b/.github/workflows/dependabot-automerge.yml index 9a65cbf..131f097 100644 --- a/.github/workflows/dependabot-automerge.yml +++ b/.github/workflows/dependabot-automerge.yml @@ -49,7 +49,14 @@ jobs: contents: write # needed to enable auto-merge pull-requests: write # needed to approve # Only run for PRs actually authored by Dependabot. - if: github.actor == 'dependabot[bot]' && github.event.pull_request.user.login == 'dependabot[bot]' + # NB: gate on the PR author only. The dispatch-actor conjunct that used to + # sit here has been removed deliberately: that context is the user who + # *triggered the run*, which a fork, branch or rerun event can set without + # the PR being Dependabot's, and Hypatia's research_extensions RE008 flags + # exactly that shape of identity check as spoofable (it scans workflow text + # verbatim, which is also why this comment describes the pattern instead of + # quoting it). The PR-author check below is the real authorisation test. + if: github.event.pull_request.user.login == 'dependabot[bot]' runs-on: ubuntu-latest timeout-minutes: 15 steps: diff --git a/.github/workflows/instant-sync.yml b/.github/workflows/instant-sync.yml index f1a94fb..c2a2244 100644 --- a/.github/workflows/instant-sync.yml +++ b/.github/workflows/instant-sync.yml @@ -1,6 +1,17 @@ # SPDX-License-Identifier: MPL-2.0 # This workflow is managed by gh actions-lock. # Instant Forge Sync - Triggers propagation to all forges on push/release +# +# STATE: DISABLED at the repository settings level (Settings → Actions → +# Workflows), so it never runs on pushes or releases even though the triggers +# below are still declared. Repository-level workflow disabling is not visible +# from the repository contents, so it is recorded here. +# +# It is kept in the tree because the propagation mechanism is still the +# intended distribution path for the estate (see mirror.yml, which publishes to +# one mirror). Re-enable only with a `FARM_DISPATCH_TOKEN` secret present on the +# destination farm; without the secret the credential check below no-ops with a +# notice, so re-enabling before the secret exists is safe but useless. name: Instant Sync on: push: diff --git a/.machine_readable/descriptiles/STATE.a2ml b/.machine_readable/descriptiles/STATE.a2ml index 3901710..24d69e0 100644 --- a/.machine_readable/descriptiles/STATE.a2ml +++ b/.machine_readable/descriptiles/STATE.a2ml @@ -5,9 +5,9 @@ [metadata] project = "oikosbot" version = "0.1.0-dev" -last-updated = "2026-08-07" +last-updated = "2026-09-26" status = "active" -session = "2026-08-07 documentation + debt audit: DEBT.adoc written (licence/docs/code/proof/CI-CD, every item evidenced); removed an UNEARNED OpenSSF Best Practices badge (API returns empty — never registered); deleted ARCHITECTURE.md and GOVERNANCE.md (generic boilerplate describing separate src/ and tests/ directories this repo does not have, shadowing the real .adoc versions); deleted orphaned LICENSES/AGPL-3.0-or-later.txt (no file declares it); fixed docs/README.adoc's DUPLICATED SPDX header (MPL-2.0 on line 1 shadowing CC-BY-SA-4.0 on line 2 — a doc licensed as code, and a class of defect that passes the presence check while asserting the wrong licence); ARCHITECTURE.adoc banner-flagged as TARGET design with each unbuilt component named; tech-debt-2026-05-26.adoc marked superseded; wiki built out from a one-line stub; repo description and topics set. Prior session 2026-08-03/04 estate economics round one: #60 merged (oikosbot-telemetry/-capability/-dea + estate CLI), oikosbot-estate snapshot repo created, #65 in review (path-keyed gate detection). Prior: 2026-07-28 generation-1 go-live: Pareto engine made executable (#42 — crates/oikosbot-pareto: ε-tolerant dominance, normalized weighted frontier, base-vs-head verdicts with confidence gating, oikosbot compare, SARIF pareto_* properties, EconScore composition per ARCHITECTURE); .oikos.yml --config + auto-discovery (estate configs + governance flag now real); push-email-notify removed (dual-use ruling); publish-image root-caused to GHCR package access (permission_denied: write_package — owner grant, Containerfile verified sound via podman); composite action.yml + docs/COMPARISON-climate-warrior.adoc. Rulings: Action-mode first then App; NO interim listener (upstream AffineScript Http::Server instead); advisor + machine-checked trade-offs; Scallop replaces DeepProbLog. Prior session: 2026-06-21 close-out of the post-extraction work: fleet bridge → BotId::Oikosbot + ReScript-era containers removed (#5); finding taxonomy in NEUROSYM.a2ml [finding-taxonomy] + policies/finding_taxonomy.ecl (#9); robot-repo-automaton build fix + Rust build/test/clippy CI gate (gitbot-fleet); stale-identity sweep across SECURITY/CLAUDE/META (#11); standards reusable-workflow pin refresh (#13); LICENSE dual-SPDX MPL-2.0 + CC-BY-SA-4.0, SECURITY.md finalized (reporting → j.d.a.jewell@open.ac.uk), redundant trufflehog job dropped (#14); added docs/README.adoc documentation map. Open follow-ups: #12 (taxonomy vocab reconciliation), #16 (developer+maintainer docs + README split), #17 (end-user docs), #18 (taxonomy tags through Finding types)." +session = "2026-08-07 documentation + debt audit: DEBT.adoc written (licence/docs/code/proof/CI-CD, every item evidenced); removed an UNEARNED OpenSSF Best Practices badge (API returns empty — never registered); deleted ARCHITECTURE.md and GOVERNANCE.md (generic boilerplate describing separate src/ and tests/ directories this repo does not have, shadowing the real .adoc versions); deleted orphaned LICENSES/AGPL-3.0-or-later.txt (no file declares it); fixed docs/README.adoc's DUPLICATED SPDX header (MPL-2.0 on line 1 shadowing CC-BY-SA-4.0 on line 2 — a doc licensed as code, and a class of defect that passes the presence check while asserting the wrong licence); ARCHITECTURE.adoc banner-flagged as TARGET design with each unbuilt component named; tech-debt-2026-05-26.adoc marked superseded; wiki built out from a one-line stub; repo description and topics set. Prior session 2026-08-03/04 estate economics round one: #60 merged (oikosbot-telemetry/-capability/-dea + estate CLI), oikosbot-estate snapshot repo created, #65 in review (path-keyed gate detection). Prior: 2026-07-28 generation-1 go-live: Pareto engine made executable (#42 — crates/oikosbot-pareto: ε-tolerant dominance, normalized weighted frontier, base-vs-head verdicts with confidence gating, oikosbot compare, SARIF pareto_* properties, EconScore composition per ARCHITECTURE); .oikos.yml --config + auto-discovery (estate configs + governance flag now real); push-email-notify removed (dual-use ruling); publish-image root-caused to GHCR package access (permission_denied: write_package — owner grant, Containerfile verified sound via podman); composite action.yml + docs/COMPARISON-climate-warrior.adoc. Rulings: Action-mode first then App; NO interim listener (upstream AffineScript Http::Server instead); advisor + machine-checked trade-offs; Scallop replaces DeepProbLog. Prior session: 2026-06-21 close-out of the post-extraction work: fleet bridge → BotId::Oikosbot + ReScript-era containers removed (#5); finding taxonomy in NEUROSYM.a2ml [finding-taxonomy] + policies/finding_taxonomy.ecl (#9); robot-repo-automaton build fix + Rust build/test/clippy CI gate (gitbot-fleet); stale-identity sweep across SECURITY/CLAUDE/META (#11); standards reusable-workflow pin refresh (#13); LICENSE dual-SPDX MPL-2.0 + CC-BY-SA-4.0, SECURITY.md finalized (reporting → j.d.a.jewell@open.ac.uk), redundant trufflehog job dropped (#14); added docs/README.adoc documentation map. Open follow-ups: #12 (taxonomy vocab reconciliation), #16 (developer+maintainer docs + README split), #17 (end-user docs), #18 (taxonomy tags through Finding types). Session 2026-09-26 (branch arena/01a0dad6-oikosbot): two long-standing red checks root-caused and fixed — Hypatia's single high/critical was research_extensions RE008 on dependabot-automerge.yml (github.actor identity check; PR-author gate retained), and Publish Image's exit 101 was an MSRV drift (tree-sitter 0.27.0 needs Rust 1.90/edition 2024 vs rust:1.88-slim; builder now rust:1-slim). #48 calibration wiring landed (see calibration-wiring and enforcement-inert). Debt register updated for the resolved items; docs delivered: docs/usage.adoc, docs/ci-runbook.adoc, per-crate READMEs, analyzers/code-haskell/README.adoc. NOT verified in CI by the author (no Rust/Haskell/Elixir toolchain in the sandbox): verification means the PR's own CI run plus, for the image, a successful Publish Image on main." [project-context] name = "OikosBot" @@ -36,8 +36,8 @@ milestones = [ issues = [ { id = "github-actions-budget", severity = "operational", summary = "Actions spending limit exhausted intermittently across the estate. Signature: job conclusion 'failure' with ZERO steps and the annotation 'job was not started because recent account payments have failed'. Hit 2 of 15 swept consumer repos (chronicles-of-slavia, canonical-ums — the latter private). NOT a workflow defect; do not debug the workflow when steps==0." }, { id = "actions-lockfile-enforcement", severity = "resolved", summary = "RESOLVED for this repo. GitHub's Actions workflow-lockfile enforcement killed every workflow estate-wide at startup (0s startup_failure, error visible only on the run's HTML page). Cured by shipping .github/workflows/actions.lock in #61; reusable-caller permissions fixed in #63. CI now runs green (Hypatia, Secret Scanner, Governance, Language Policy, CodeQL). Consumers still need their own lockfiles — see consumer-fleet-dark." }, - { id = "enforcement-inert", severity = "design", summary = "PER-FILE path only: `--check` cannot block a merge, since only Measured/Calibrated inputs may fail a run and the analyzer emits only Estimated (calibration.rs exists but has ZERO callers; estimate_resources() is still naive complexity*0.1 J). Refusal is LOUD (::warning::) as of #47, never silent. Real fix tracked in issue #48. NOTE the ESTATE path is now the exception: oikosbot-telemetry derive.rs assigns Confidence::Measured to wall_minutes from the GitHub API — the first genuinely Measured quantity in the system." }, - { id = "per-file-collinearity", severity = "design", summary = "estimate_resources() derives energy, duration, carbon and memory from ONE integer (complexity = raw AST node count), so four of five Pareto objectives are scalar multiples of each other and the frontier collapses to a 1-D sort. The dominance maths in oikosbot-pareto is correct; the inputs make it near-vacuous. SOLVED at estate level (telemetry axes measured independent: wall_minutes~size_kb = -0.049 across 381 repos). UNSOLVED at file level. See DEBT.adoc." }, + { id = "enforcement-inert", severity = "resolved-per-file", summary = "RESOLVED for the per-file path on 2026-09-26 (#48): calibration is wired. estimate_resources() maps a detected pattern onto an OperationKind (calibration::operation_for_pattern) and prices it with estimate_operation(), which returns the min/typical/max band AND the confidence that row has earned — Calibrated for HashLookup/Sort/Allocation/MathCompute, Estimated for the host-dependent FileIO/StringOp rows, Estimated on the naive complexity path. So assess().actionable is now true for calibrated drivers and `compare --check` exits 1 on an undocumented calibrated regression; proven by crates/oikosbot-cli/tests/check_gate.rs. Heuristic findings still cannot block, and the refusal stays LOUD (::warning::, #47). STILL OPEN: absolute figures are unvalidated against profiling data, and the eco score's log scale (anchored at 1 J) saturates near 100 under calibrated microjoule-scale estimates, so the eco THRESHOLD discriminates far less than before while the Pareto verdicts (base vs head) are unaffected. ESTATE path remains the genuinely Measured exception: oikosbot-telemetry derive.rs assigns Confidence::Measured to wall_minutes from the GitHub API." }, + { id = "per-file-collinearity", severity = "design", summary = "PARTIALLY SOLVED 2026-09-26 (#48). Units with a recognised pattern are now priced from per-kind calibration rows whose energy/duration/memory formulae are NOT the same function of complexity (Sort is n*log2(n) on energy and duration but linear on memory; Allocation is linear on bytes; MathCompute has no memory term), so those units no longer collapse the frontier. Units with NO recognised pattern still go through the naive path, where energy, duration, carbon and memory all derive from ONE integer (complexity = raw AST node count) and four of five objectives remain scalar multiples of each other. The dominance maths in oikosbot-pareto is correct; the naive inputs make it near-vacuous. SOLVED at estate level (telemetry axes measured independent: wall_minutes~size_kb = -0.049 across 381 repos). See DEBT.adoc." }, { id = "policy-engines-never-execute", severity = "design", summary = "TWO fake gates. (1) policy-engine/datalog/eco_rules.dl has never executed — Souffle is DECLARED in guix/manifest.scm and guix/oikos.scm but invoked nowhere (no match in Justfile, *.just or any workflow); its allocation-waste and debt rules have no Rust counterpart. (2) oikosbot-eclexia's default backend dispatches on the .ecl FILE STEM and never reads file contents; its hardcoded thresholds contradict the files (energy_threshold.ecl says >50 J per function, builtin fires >1000 J total). Mitigated by a loud ::warning:: in #59; still fake." }, { id = "consumer-fleet-dark", severity = "operational", summary = "15 estate repos carry .github/workflows/oikosbot.yml pinned to oikosbot@bb95ab50 (v0.1.0), all merged — but each needs its OWN Actions lockfile before its workflows can start. OikosBot is installed everywhere and running nowhere. Highest-value follow-up." }, { id = "idaptik-ums-repo-wide-startup-failure", severity = "external", summary = "metadatastician/idaptik-ums fails ALL workflows at startup on main (OikosBot, Licence hygiene, CodeQL), with two workflows displayed as PATHS not names — the estate tell for never-parsed. Repo-level, pre-existing, not caused by the OikosBot sweep (the same file succeeded on a branch there)." } @@ -45,7 +45,7 @@ issues = [ [critical-next-actions] actions = [ - { id = "calibration-wiring", priority = "P1", summary = "Issue #48: map detected patterns to calibration OperationKinds so confidence is EARNED per finding, unlocking real enforcement. Changes every resource figure and every downstream score — needs before/after numbers." }, + { id = "calibration-wiring", priority = "P1", status = "landed-2026-09-26", summary = "Issue #48 DONE (pending CI verification): patterns.rs detections mapped onto OperationKind via calibration::operation_for_pattern; estimate_operation() prices recognised units and returns a ResourceRange carrying that row's confidence; naive path retained for unrecognised code and labelled Estimated; ResourceRange propagated on AnalysisResult and emitted in SARIF properties.resource_range; falsifier tests added (calibration confidence ladder, pattern mapping, analyzer Calibrated-vs-Estimated, and CLI tests/check_gate.rs proving --check can exit 1). BEFORE/AFTER (derived by reasoning, NOT measured — no toolchain in the authoring sandbox): a depth-3 nested-loop unit of ~30 AST nodes went from naive 0.1 J/node (~3 J, eco ~89) to the calibrated Sort row (~0.007 J, eco clamped 100); every downstream figure moves with it. The eco-score saturation is documented in STATUS.adoc/QUICKSTART.adoc/README.adoc and flagged for an owner ruling on the score scale." }, { id = "rsr-julia-template-phantom-sha", priority = "P1", summary = "hyperpolymath/rsr-julia-library-template-repo .github/workflows/codeql.yml pins phantom SHA 29b1f65c1f735799893313399435a59f54045865 (no such commit). It is a TEMPLATE, so every minted repo inherits a dead CodeQL gate. Valid replacement: 4187e74d05793876e9989daffde9c3e66b4acd07." }, { id = "bot-affine-webhook-handler", priority = "P2", summary = "Wire the AS-side webhook receiver in bot-integration-affine/ using Http server + Json stdlib externs (ruled: do the upstream AffineScript work, no interim listener)" }, { id = "hpm-json-object-keys", priority = "P3", summary = "Add hpm_json_object_keys export to hpm-json-rsr Zig FFI to close the JObject materialisation gap in stdlib/json.affine to_json" }, diff --git a/CHANGELOG.adoc b/CHANGELOG.adoc index 4c4f98f..cae6ca7 100644 --- a/CHANGELOG.adoc +++ b/CHANGELOG.adoc @@ -17,6 +17,36 @@ https://semver.org/spec/v2.0.0.html[Semantic Versioning]. ==== Added +* feat(analysis): wire the calibration framework into the analyser (issue #48). + `Analyzer::estimate_resources()` maps a detected pattern onto an + `+OperationKind+` via the new `+calibration::operation_for_pattern+` and + prices it with `+estimate_operation()+`, whose `+ResourceRange+` carries the + min/typical/max band **and** the confidence that row has earned — + `+Calibrated+` for the measured operation kinds (`+HashLookup+`, `+Sort+`, + `+Allocation+`, `+MathCompute+`), `+Estimated+` for the host-dependent ones + (`+FileIO+`, `+StringOp+`, `+NetworkCall+`). Unrecognised code keeps the naive + complexity-derived estimate, labelled `+Estimated+`, with no band claimed. + `+redundant-allocation+` is deliberately unmapped so a marginal finding cannot + inherit calibrated confidence. +* feat(metrics): `+ResourceRange+` (min/typical/max + the row's confidence) and + `+AnalysisResult::resource_range+`, `+#[serde(default, + skip_serializing_if = "Option::is_none")]+`, so the uncertainty band is + propagated instead of collapsed to a point estimate. SARIF emits it as + `+properties.resource_range+`. +* test(cli): `+tests/check_gate.rs+` — end-to-end proof that + `+oikosbot compare --check+` can actually block: an undocumented calibrated + regression exits 1, the same regression without `+--check+` exits 0, a + documented `+Pareto-Trade-off:+ is accepted, and a heuristic regression + refuses to block *and says so*. +* docs: `+docs/usage.adoc+` (end-user guide: action inputs, the three modes, + verdicts and the confidence ladder, the SARIF shape, `+.oikos.yml+`, + `+BOT_MODE+`, troubleshooting) and `+docs/ci-runbook.adoc+` (gate inventory, + `+actions.lock+` rules, diagnosing Hypatia and `+Publish Image+` without the + logs, `+just+` target reference, pre-release checklist). Both are indexed in + `+docs/README.adoc+`. +* docs: per-crate guides (`+crates/*/README.adoc+`) and + `+analyzers/code-haskell/README.adoc+`, each stating what the crate owns, its + public surface, and what it does **not** do. * feat(policies): add the OikosBot *finding taxonomy* — three orthogonal axes (`+intent+` / `+maintenance+` / `+locus+`) defined canonically in `+NEUROSYM.a2ml [finding-taxonomy]+`. The confidence-derived `+intent+` @@ -50,6 +80,15 @@ payload extraction, and HTTP-server accept loop are gated on upstream ==== Removed +* chore(policy-engine): delete `+policy-engine/deepproblog/eco_problog.pl+`, + dead since the 2026-07-28 ruling (policy engine → Scallop, Python DeepProbLog + retired with no exemption). `+EXPLAINME.adoc+` no longer cites it. +* refactor(analysis): remove `+calibration::calibrated_estimate()+` and its + duplicated multiplier table — superseded by the pattern → operation-kind + mapping, which left `+patterns.rs+` as the single source of truth for pattern + severity. +* refactor(pareto): remove `+standard_objectives()+`, dead code since the + workspace was extracted; the live objective set is `+result_objectives()+`. * chore(containers): remove the stale ReScript-era `+containers/+` that were ported with the extraction but still built the long-removed `+bot-integration/+` ReScript bot (`+*.res.js+`, `+rescript-runtime/+`) @@ -58,6 +97,16 @@ rebuilt natively when OikosBot is deployable. ==== Fixed +* fix(ci): Hypatia's sole high/critical finding was + `+research_extensions+` RE008 — `+dependabot-automerge.yml+` gated auto-merge + on `+github.actor == 'dependabot[bot]'+`, an identity check built on a context + a fork or rerun event can set. The PR-author gate is kept; the actor conjunct + is removed. +* fix(containers): `+Publish Image+` failed with `+cargo build --release + --locked -p oikosbot-cli+` exit 101 because `+tree-sitter+` 0.27.0 declares + `+rust-version = "1.90"+` / `+edition = "2024"+` while the builder was + `+rust:1.88-slim+`. The builder is now the rolling `+rust:1-slim+`, so the + MSRV floor stays satisfied as dependencies move. * fix(release): restore GHCR publication with a Rust 1.88 builder and the CMake/C++/Make/libclang toolchain required by `+highs-sys+`, publish versioned image tags for release tags, and make the Action default to the verified immutable image digest @@ -94,6 +143,19 @@ step-level (#17) ==== Changed +* refactor(analysis): `+ResourceRange+` moves to `+oikosbot-metrics+` (so it + can be carried on an `+AnalysisResult+`) and gains the `+confidence+` its row + has earned — confidence is a property of the evidence behind an estimate, not + of the estimate. `+estimate_operation()+` keeps its signature and now prices + each row with the confidence that row deserves: `+Calibrated+` for the + measured operation kinds, `+Estimated+` for the host-dependent ones. +* docs: the standing "`+--check+` cannot block a merge" caveats in + `+README.adoc+`, `+QUICKSTART.adoc+`, `+docs/STATUS.adoc+`, `+action.yml+` and + `+docs/COMPARISON-climate-warrior.adoc+` now describe what is true: calibrated + findings may block, heuristic findings may not, and the refusal is loud. The + same documents now name the eco-score saturation that follows from calibrated + (microjoule-scale) estimates, so a passing eco threshold is not mistaken for a + working one. * chore(decouple): sever OikosBot’s dependency on `+gitbot-fleet+`. The default `+cargo+` workspace *excludes* `+crates/oikosbot-fleet+` (the only fleet-aware crate), and the optional `+panic-attacker+` / diff --git a/Containerfile b/Containerfile index 930e2c4..f834c63 100644 --- a/Containerfile +++ b/Containerfile @@ -10,7 +10,12 @@ # Multi-stage, glibc-consistent (rust:slim builder -> debian:slim runtime), # non-root. The oikosbot-fleet bridge is excluded from the default workspace, # so this builds standalone with no gitbot-fleet dependency. -FROM rust:1.88-slim AS builder +# MSRV floor for this workspace is Rust 1.90 (tree-sitter 0.27.0 declares +# rust-version = "1.90", edition = "2024"). Use the rolling stable builder +# rather than a pinned point release so the floor stays satisfied as +# dependencies move; arrow-rs 59.3.0 declares MSRV 1.88, tree-sitter is the +# binding constraint. +FROM rust:1-slim AS builder WORKDIR /build RUN apt-get update \ && apt-get install -y --no-install-recommends cmake g++ libclang-dev make \ diff --git a/DEBT.adoc b/DEBT.adoc index d7db76a..bf990ad 100644 --- a/DEBT.adoc +++ b/DEBT.adoc @@ -29,9 +29,9 @@ Severity is about *consequence*, not effort: | <> | HYGIENE | One orphaned licence text; one doc missing its SPDX header. | <> | STRUCTURAL | Two boilerplate files describe a directory layout this repo does not have; `ARCHITECTURE.adoc` specifies components that were never built. -| <> | BLOCKING | The per-file analyser derives every resource axis from one integer, so its Pareto frontier is a one-dimensional sort. +| <> | BLOCKING | Unrecognised code still derives every resource axis from one integer, so its Pareto frontier is a one-dimensional sort. Recognised patterns are now calibrated (2026-09-26). | <> | BLOCKING | The Datalog policy engine has never executed; the Eclexia policy backend matches on filenames and never reads the files. -| <> | HYGIENE | Repo CI is green; the consumer fleet cannot run until each repo ships a lockfile. +| <> | BLOCKING | Four phantom required contexts block every pull request, so every merge needs `--admin`. The two long-standing red checks are fixed; the consumer fleet cannot run until each repo ships a lockfile. |=== [#licence] @@ -142,10 +142,11 @@ open and are only partly addressed by `docs/README.adoc`. [cols="1,4"] |=== -| `BLOCKING` | *Every resource axis in the per-file analyser is a scalar -multiple of one integer.* `Analyzer::estimate_resources()` -(`crates/oikosbot-analysis/src/analyzer.rs:199`) derives all four axes from -`complexity`, a raw AST **node count**: +| `BLOCKING` | *Every resource axis in the per-file analyser's naive path is a +scalar multiple of one integer.* `Analyzer::naive_resources()` +(`crates/oikosbot-analysis/src/analyzer.rs`) derives all four axes from +`complexity`, a raw AST **node count**. It is still the path taken by any unit +with no recognised pattern: + [source,rust] ---- @@ -164,25 +165,37 @@ rank-inverted, and `Debt` is `100 − 0.5 × complexity`. + *Status:* solved at **estate** level by round one (telemetry gives money, time, energy and carbon as mutually independent axes — measured -`wall_minutes ~ size_kb` correlation is −0.049). *Unsolved at file level.* - -| `BLOCKING` | *`calibration.rs` has zero callers* (issue -https://github.com/hyperpolymath/oikosbot/issues/48[#48]). The module is -written in full — `OperationKind`, `ResourceRange`, pattern multipliers — and -nothing calls `calibrated_estimate()` or `estimate_operation()`. Verified: -`grep -rn "calibrated_estimate\|estimate_operation" crates/` matches only the -defining file. - -| `BLOCKING` | *Consequence: `compare --check` cannot block.* Only -`Measured`/`Calibrated` inputs may fail a run; the analyser emits only -`Estimated`, so `assess().actionable` is always false. Since -https://github.com/hyperpolymath/oikosbot/pull/47[#47] it refuses *loudly* -rather than passing silently — the honest handling of a gate that cannot yet -gate, but the gate still cannot gate. +`wall_minutes ~ size_kb` correlation is −0.049). *Partly solved at file level* +(2026-09-26): a unit with a recognised pattern is priced from a per-kind row +whose formulae differ per axis — `Sort` is `n·log₂ n` on energy and duration but +linear on memory; `Allocation` is linear on bytes; `MathCompute` has no memory +term — so those units no longer collapse the frontier. Units with **no** +recognised pattern still take the naive path below and remain collinear. + +| — | *Resolved 2026-09-26:* *`calibration.rs` had zero callers* (issue +https://github.com/hyperpolymath/oikosbot/issues/48[#48]). +`Analyzer::estimate_resources()` now maps a detected pattern onto +an `OperationKind` via `calibration::operation_for_pattern` and prices it with +`estimate_operation`, whose `ResourceRange` carries the min/typical/max band +*and* the confidence that row has earned. The band is propagated on `AnalysisResult` +(`oikosbot_metrics::ResourceRange`) and emitted in SARIF `properties`. +`calibrated_estimate()` — the unwired, multiplier-based path — was deleted +rather than left as a second, contradictory estimator. + +| — | *Resolved 2026-09-26 (partly):* *`compare --check` could not block.* +Findings built on a recognised pattern now carry +`Calibrated` (measured rows) or `Estimated` (host-dependent rows), so +`assess().actionable` is true for calibrated drivers and `--check` exits 1 on an +undocumented calibrated regression. Proven by +`crates/oikosbot-cli/tests/check_gate.rs` against the real binary. Heuristic +findings still cannot block, and the refusal stays loud (`::warning::`), as +https://github.com/hyperpolymath/oikosbot/pull/47[#47] established. + *Round one changed this partially:* `Confidence::Measured` is now produced for the first time, at `crates/oikosbot-telemetry/src/derive.rs:99`, for -`wall_minutes`. The estate path has Measured data; the per-file path does not. +`wall_minutes`. The estate path has Measured data; the per-file path has +`Calibrated` data (2026-09-26, above) for units with a recognised pattern, and +`Estimated` data for the rest. | `STRUCTURAL` | *`aggregate_point` sums resources but averages quality* (`crates/oikosbot-pareto`), so a `compare` verdict is sensitive to file @@ -194,9 +207,10 @@ improved. weights, while `config/oikos.yaml` advertises a `weights:` block that the loader discards. -| `STRUCTURAL` | *`standard_objectives()` is dead code* — seven axes declared -per `ARCHITECTURE.adoc`, never called. The live set is `result_objectives()` -(five axes, no coverage or debt axis). +| — | *Resolved 2026-09-26:* *`standard_objectives()` was dead code* — seven +axes declared per `ARCHITECTURE.adoc`, never called. The function was deleted; +the live set remains `result_objectives()` (five axes, no coverage or debt +axis). |=== === Language and detection coverage @@ -218,11 +232,24 @@ measurable at all), but per-file analysis cannot see most of the estate. combines AST-kind matching with **substring search on node source text**, producing false negatives on JS and Python. -| `HYGIENE` | *Detected patterns never affect the estimate.* `impact_multiplier` -is computed and attached to findings but is never applied in `analyzer.rs`; the -multiplier logic lives only in the unwired calibration path — and the two -multiplier tables (`patterns.rs` and `calibration.rs`) are duplicated with no -shared source of truth. +| — | *Resolved 2026-09-26:* *detected patterns never affected the estimate.* +`impact_multiplier` was computed and attached to findings but never applied. A +detected pattern now selects the operation category +that prices the unit, and where several patterns map, the largest +`impact_multiplier` wins. The duplicated multiplier table in `calibration.rs` +was deleted with `calibrated_estimate()`; `patterns.rs` is the single source of +truth for pattern severity. + +[NOTE] +==== +What this did **not** fix, and must not be read as fixed: the naive path is +still collinear (four axes, one integer), so units with no recognised pattern +keep the one-dimensional-frontier problem described above. The eco score's log +scale is anchored at 1 J, so calibrated microjoule-scale estimates clamp near +100 and the eco *threshold* discriminates far less than it did; the Pareto +verdicts, which compare base against head, are unaffected. Absolute figures +remain unvalidated against profiling data. +==== |=== === Round-one pipeline (`oikosbot-telemetry` / `-capability` / `-dea`) @@ -304,9 +331,10 @@ behind a feature whose crates (`eclexia_parser`, `eclexia_interp`) are not declared dependencies, so the cfg branch is unbuildable — a seam, not an implementation. -| `HYGIENE` | *`policy-engine/deepproblog/eco_problog.pl` is dead by ruling* -(2026-07-28: policy engine → Scallop, Python DeepProbLog retired with no -exemption) but the asset is still in the tree. +| — | *Resolved 2026-09-26:* *`policy-engine/deepproblog/eco_problog.pl` was +dead by ruling* (2026-07-28: policy engine → Scallop, Python DeepProbLog retired +with no exemption) while the asset was still in the tree. The file is deleted and +`EXPLAINME.adoc` no longer cites it. | `HYGIENE` | *DEA correctness is verified against self-authored analytic cases* — closed-form single-input/output, a strictly-dominated unit, and a strong-duality @@ -326,11 +354,23 @@ downstream. | — | *Resolved:* the estate-wide `startup_failure` epidemic (GitHub's Actions workflow-lockfile enforcement) is cured for this repo — `.github/workflows/actions.lock` shipped in -https://github.com/hyperpolymath/oikosbot/pull/61[#61] and CI now runs green -(Hypatia, Secret Scanner, Governance, Language Policy, CodeQL all `success`). -Reusable-caller permissions fixed in +https://github.com/hyperpolymath/oikosbot/pull/61[#61]. Reusable-caller +permissions fixed in https://github.com/hyperpolymath/oikosbot/pull/63[#63]. +| — | *Resolved 2026-09-26:* two red checks on `main` that had nothing to do +with each other. `Hypatia` was failing on **one** high/critical finding — +`research_extensions` RE008, critical, because +`.github/workflows/dependabot-automerge.yml` gated auto-merge on +`github.actor == 'dependabot[bot]'`, an identity check built on a context a +fork or rerun can set. The PR-author gate is kept; the actor conjunct is gone. +`Publish Image` was failing with `cargo build … exit code 101` because +`tree-sitter` 0.27.0 declares `rust-version = "1.90"` / `edition = "2024"` +while the `Containerfile` builder was `rust:1.88-slim`; CI on `ubuntu-latest` +passed because the runner ships a newer rustc. The builder is now +`rust:1-slim`. Both diagnoses, and the gate mechanics behind them, are written +up in link:docs/ci-runbook.adoc[`docs/ci-runbook.adoc`]. + | `BLOCKING` | *Four phantom required contexts block every pull request, so every merge needs `--admin`.* The `main` ruleset requires 27 status checks. All 27 checks that *can* report do report and pass — but four of the required @@ -378,8 +418,10 @@ the single highest-value follow-up. | `HYGIENE` | *`scorecard.yml` runs per-push*; making it periodic is open as https://github.com/hyperpolymath/oikosbot/pull/64[#64]. -| `HYGIENE` | *`instant-sync.yml` is disabled* (`disabled_manually`) and -undocumented — neither retired nor explained. +| — | *Resolved 2026-09-26:* *`instant-sync.yml` was disabled* +(`disabled_manually`) and undocumented. The workflow's header now records the +settings-level disabled state (which is invisible from the repository contents), +why it is kept, and what re-enabling requires. | `HYGIENE` | *`push-email-notify.yml` is dormant by design* but remains in the tree; the estate's dual-use ruling calls for removing push-email workflows diff --git a/EXPLAINME.adoc b/EXPLAINME.adoc index 5fb7715..3e10edf 100644 --- a/EXPLAINME.adoc +++ b/EXPLAINME.adoc @@ -100,18 +100,19 @@ policy-engine/ — Datalog and DeepProbLog policy rules; policies/ — Eclexia ____ How this is implemented:: -`policy-engine/datalog/eco_rules.dl` holds the deterministic rules and -`policy-engine/deepproblog/eco_problog.pl` the probabilistic (learned) rules. +`policy-engine/datalog/eco_rules.dl` holds the deterministic rule set +(link:#proof[see below] for why it has never executed). The probabilistic +(learned) layer that used to live in `policy-engine/deepproblog/eco_problog.pl` +was retired by owner ruling on 2026-07-28: the policy engine's learning backend +is Scallop, and Python DeepProbLog assets are no longer part of the tree. The Eclexia resource policies in link:policies/[`policies/`] (`carbon_budget.ecl`, `energy_threshold.ecl`, `memory_efficiency.ecl`, `security_sustainability.ecl`) are evaluated by the `oikosbot-eclexia` crate. Caveat:: -The DeepProbLog rules and the praxis (learning) loop are design-stage: the -probabilistic predicates and neural-predicate declarations exist as logic, but -there is no trained model or live feedback ingestion wired in. The knowledge-graph -bindings in `eco_problog.pl` were retargeted from SPARQL-over-Virtuoso to -VCL-over-VeriSimDB at the comment/binding level only (no rule logic changed). +The learning (probabilistic) layer is design-stage: the deterministic rules +exist as logic, but there is no trained model and no live feedback ingestion +wired in. == "Single data layer: VeriSimDB" diff --git a/QUICKSTART.adoc b/QUICKSTART.adoc index 78f63ad..2a56ef5 100644 --- a/QUICKSTART.adoc +++ b/QUICKSTART.adoc @@ -38,18 +38,18 @@ fails the run on an undocumented, measured regression or trade-off. [IMPORTANT] ==== -*`--check` cannot block a merge today, by design.* Only objectives backed by -`Measured` or `Calibrated` inputs may fail a run, and every resource figure -OikosBot currently produces is a heuristic `Estimated` value — the calibration -framework at `crates/oikosbot-analysis/src/calibration.rs` exists but is not -yet wired into the analyzer, which still uses a naive complexity-derived -estimate. - -So `--check` reports and warns, but always exits 0. It says so *loudly* — a -`::warning::` annotation naming the confidence level — rather than passing -silently. A gate that quietly does nothing is indistinguishable from one that -passed, which is precisely the failure mode OikosBot exists to find. -Enforcement becomes real when calibration is wired in. +*`--check` blocks only on evidence, by design.* Only objectives backed by +`Measured` or `Calibrated` inputs may fail a run. The analyser earns that +evidence per finding: a recognised pattern is priced from the calibration table +(`crates/oikosbot-analysis/src/calibration.rs`) and labelled `Calibrated`, while +unclassified code stays on the naive complexity-derived estimate and is labelled +`Estimated`. + +So an undocumented regression whose drivers are calibrated exits 1; the same +regression on heuristic drivers prints a `::warning::` naming the verdict and +the confidence level, and exits 0. A gate that quietly does nothing is +indistinguishable from one that passed, which is precisely the failure mode +OikosBot exists to find — so the refusal is never silent. ==== === Per-repo configuration diff --git a/README.adoc b/README.adoc index 889633e..3bbcb70 100644 --- a/README.adoc +++ b/README.adoc @@ -88,11 +88,13 @@ configuration, and SARIF output that GitHub code scanning ingests. [IMPORTANT] ==== -**Per-file resource figures are heuristic estimates and `--check` cannot block a -merge.** Only `Measured`/`Calibrated` inputs may fail a run, and the analyser -emits only `Estimated` — calibration exists but is not yet wired in (issue -\#48). OikosBot warns loudly rather than passing silently, but treat the -per-file verdict as an advisor, not a regulator. +**Per-file resource figures are still estimates, and only calibrated ones can +block.** The analyser prices a recognised pattern from the calibration table and +labels the finding `Calibrated`; anything it cannot classify stays on the naive +complexity path and is labelled `Estimated`. Only `Measured`/`Calibrated` +objectives may fail a run, so `--check` blocks on a calibrated regression and +warns loudly (exit 0) on a heuristic one. The absolute figures remain static and +unvalidated — trust the direction, not the joules. The *estate* path is the exception: `wall_minutes` comes straight from the GitHub API and is genuinely `Measured`. See link:DEBT.adoc[`DEBT.adoc`] for the diff --git a/action.yml b/action.yml index 1acbede..1f5ffa5 100644 --- a/action.yml +++ b/action.yml @@ -50,11 +50,11 @@ inputs: default: '' check: description: >- - compare mode: fail on an undocumented, measured Pareto - regression/trade-off. NOTE: cannot block today — every resource figure is - a heuristic Estimated value and only Measured/Calibrated inputs may - fail a run, so this warns loudly and exits 0 until calibration is wired - in (see QUICKSTART). + compare mode: fail on an undocumented Pareto regression/trade-off whose + driving objectives are Measured or Calibrated. Recognised patterns are + priced from the calibration table and may block; unrecognised code stays + on the naive estimate (Estimated), which may not — it warns loudly and + exits 0 instead (see QUICKSTART and docs/usage.adoc). required: false default: 'false' image: diff --git a/analyzers/code-haskell/README.adoc b/analyzers/code-haskell/README.adoc new file mode 100644 index 0000000..86943e9 --- /dev/null +++ b/analyzers/code-haskell/README.adoc @@ -0,0 +1,66 @@ +// SPDX-License-Identifier: CC-BY-SA-4.0 +// SPDX-FileCopyrightText: 2026 Jonathan D.A. Jewell += The Haskell analyzer (`analyzers/code-haskell`) +:toc: macro + +A second, independent implementation of the same scoring model, in Haskell. +It exists as a cross-check: two implementations of one specification that +disagree are evidence about the specification, and two that agree on +independently written fixtures are evidence about the implementation. + +== What it computes + +[cols="1,4"] +|=== +| `Eco/Carbon.hs` | Carbon estimates for the resource profiles. + +| `Eco/Energy.hs` | Energy estimates. + +| `Eco/Resource.hs` | The resource profile type and its arithmetic. + +| `Eco/Pareto.hs` | Dominance and frontier logic. + +| `Eco/Analysis.hs` | The top-level analysis entry point. + +| `Quality/Complexity.hs` | Complexity scoring. + +| `Quality/Coupling.hs` | Coupling scoring. + +| `Quality/Coverage.hs` | Coverage scoring. + +| `Quality/Debt.hs` | Technical-debt scoring. + +| `Types/Metrics.hs` | The metric types shared across modules. + +| `Types/Report.hs` | The report/output types. + +| `app/Main.hs` | The executable entry point. + +| `test/Main.hs` | The test suite entry point (`cabal test all`). +|=== + +== How it relates to the Rust workspace + +The Rust workspace under `crates/` is the shipping implementation: it is what +the container image runs, what the GitHub Action calls, and what the estate +pipeline uses. The Haskell analyzer is built and tested by CI +(`just haskell-build` / `just haskell-test`, or the `Haskell Analyzer` job in +`ci.yml`) but is not wired into the action or the CLI. + +That asymmetry is deliberate and worth stating plainly: the two are not kept +in lock-step by a shared schema, so agreement between them is a finding to +check, not a guarantee to rely on. Where they disagree, neither is +automatically right. + +== Build and test + +[source,shell] +---- +just haskell-build +just haskell-test +---- + +== What it does not do + +It does not read GitHub metadata, does not emit SARIF, and does not run in the +published container. It has no dependency on the Rust crates and vice versa. diff --git a/crates/oikosbot-analysis/README.adoc b/crates/oikosbot-analysis/README.adoc new file mode 100644 index 0000000..292b675 --- /dev/null +++ b/crates/oikosbot-analysis/README.adoc @@ -0,0 +1,74 @@ +// SPDX-License-Identifier: CC-BY-SA-4.0 +// SPDX-FileCopyrightText: 2026 Jonathan D.A. Jewell += `oikosbot-analysis` +:toc: macro + +Turns source files into `AnalysisResult`s: parse with tree-sitter, walk the +AST, detect inefficiency patterns, and price the unit. + +== Modules + +[cols="1,4"] +|=== +| `analyzer` | `Analyzer` — parses a file or source string, finds function +nodes, and builds one result per function. + +| `language` | Extension → tree-sitter language. Supports `rs`, `js`/`mjs`/`cjs`, +`ts`/`mts`/`cts` (JavaScript grammar), `py`/`pyw`. + +| `patterns` | `detect_patterns` — nested loops (depth ≥ 3), busy-wait, string +concatenation in a loop, `.clone()` in a loop, unbuffered I/O, large +allocation, redundant `to_string`/`to_owned`. Each match carries an +`impact_multiplier` and a suggestion. + +| `calibration` | `estimate_operation(kind, n)` — the priced operation rows, +each returning a `ResourceRange` that carries the min/typical/max band **and** +the confidence that row has earned (`Calibrated` for the measured kinds, +`Estimated` for the host-dependent ones); `operation_for_pattern` — the +documented pattern → `OperationKind` mapping. + +| `carbon` | `estimate_carbon(energy)` — the single carbon model used everywhere. + +| `dependencies` | Manifest scanning for heavy dependencies +(`oikosbot/heavy-dependency` findings). + +| `directives` | In-source directive parsing. + +| `migration` | Migration helpers for the analysis output. + +| `security` | Security–sustainability correlation. **Feature-gated** +(`panic-attack`): not compiled by default, enabled by +`oikosbot-cli --features security`. +|=== + +== How a resource estimate is chosen + +`Analyzer::estimate_resources` has exactly two paths, chosen by evidence: + +. **Calibrated** — a detected pattern maps to a known operation category. The + estimate comes from `calibration::estimate_operation`, the finding carries + that row's confidence (`Calibrated` for `HashLookup`, `Sort`, `Allocation`, + `MathCompute`; `Estimated` for the host-dependent `FileIO`/`StringOp`/ + `NetworkCall` rows), + and the min/typical/max band is propagated on the result. When several + patterns map, the largest `impact_multiplier` wins. +. **Naive** — no recognised pattern. The historical + `complexity × constant` heuristic (0.1 J per AST node), labelled + `Estimated`, with no band claimed. + +`redundant-allocation` is deliberately unmapped: a marginal finding must not +inherit `Calibrated` confidence from the allocation row. + +== What it does not do + +It does not run the policies under `policies/` (that is `oikosbot-eclexia`), +does not rank units against each other (`oikosbot-pareto`), and does not emit +SARIF (`oikosbot-sarif`). The absolute figures are static estimates: the +calibrated path is better anchored than the naive one, but neither has been +validated against profiling data. + +== Tests + +---- +cargo test -p oikosbot-analysis +---- diff --git a/crates/oikosbot-analysis/src/analyzer.rs b/crates/oikosbot-analysis/src/analyzer.rs index 0486883..df3a7f9 100644 --- a/crates/oikosbot-analysis/src/analyzer.rs +++ b/crates/oikosbot-analysis/src/analyzer.rs @@ -3,9 +3,10 @@ //! Core analysis engine using tree-sitter AST +use crate::calibration::{estimate_operation, operation_for_pattern, OperationKind}; use crate::carbon::estimate_carbon; use crate::language::Language; -use crate::patterns::detect_patterns; +use crate::patterns::{detect_patterns, PatternMatch}; use anyhow::{Context, Result}; use oikosbot_metrics::*; use std::fs; @@ -107,13 +108,16 @@ impl Analyzer { // Estimate resources based on code patterns let complexity = self.estimate_complexity(node); - let resources = self.estimate_resources(complexity); - // Detect problematic patterns + // Detect problematic patterns first: they decide which calibration row + // (if any) the estimate is entitled to use. let pattern_matches = detect_patterns(source, node); let patterns: Vec = pattern_matches.iter().map(|p| p.name.clone()).collect(); let recommendations = self.generate_recommendations(&patterns); + let (resources, confidence, resource_range) = + self.estimate_resources(&pattern_matches, complexity); + // Derive rule_id and suggestion from most significant pattern let (rule_id, suggestion) = if let Some(pm) = pattern_matches.first() { (format!("oikosbot/{}", pm.name), pm.suggestion.clone()) @@ -138,8 +142,9 @@ impl Analyzer { rule_id, suggestion, end_location: Some((end.row + 1, end.column + 1)), - confidence: oikosbot_metrics::Confidence::Estimated, + confidence, pareto: None, + resource_range, }) } @@ -196,8 +201,62 @@ impl Analyzer { count } - fn estimate_resources(&self, complexity: usize) -> ResourceProfile { - // Baseline estimates (will be improved with profiling data) + /// Estimate the resources a unit uses, together with the confidence that + /// estimate is entitled to and the uncertainty band behind it. + /// + /// Two paths, chosen by evidence rather than by preference: + /// + /// * **Calibrated path** — the unit carries a detected pattern that maps to + /// a known operation category (`calibration::operation_for_pattern`). The + /// estimate comes from `calibration::estimate_operation` and the + /// confidence is whatever that row earns (Calibrated for measured rows, + /// Estimated for host-dependent ones). The min/typical/max band is + /// propagated so consumers can see the spread instead of treating the + /// point estimate as exact. + /// * **Naive path** — no recognised pattern, so there is nothing to price. + /// The historical `complexity * constant` heuristic is kept, the finding + /// is labelled `Estimated`, and no band is claimed. + /// + /// When several patterns map, the one with the largest impact multiplier + /// wins (first one on ties, i.e. detection order): the most severe + /// recognised cost driver prices the unit. `redundant-allocation` is + /// intentionally unmapped and therefore never promotes a unit off the naive + /// path. + fn estimate_resources( + &self, + pattern_matches: &[PatternMatch], + complexity: usize, + ) -> (ResourceProfile, Confidence, Option) { + let mut dominant: Option<(OperationKind, f64)> = None; + for pattern in pattern_matches { + let Some(kind) = operation_for_pattern(&pattern.name) else { + continue; + }; + if dominant.is_none_or(|(_, mult)| pattern.impact_multiplier > mult) { + dominant = Some((kind, pattern.impact_multiplier)); + } + } + + match dominant { + Some((kind, _)) => { + let range = estimate_operation(kind, complexity); + // The row that priced the unit also states how much its figure + // is worth; that confidence travels with the finding. + (range.typical.clone(), range.confidence, Some(range)) + } + None => { + let profile = self.naive_resources(complexity); + (profile, Confidence::Estimated, None) + } + } + } + + /// The historical heuristic: four axes derived from the AST node count + /// alone. Kept for units with no recognised pattern — but note that all + /// four axes are the same linear function of `complexity`, so a Pareto + /// frontier computed over them is a one-dimensional sort. That is why the + /// calibrated path exists. + fn naive_resources(&self, complexity: usize) -> ResourceProfile { let energy = Energy::joules(complexity as f64 * 0.1); let duration = Duration::milliseconds(complexity as f64 * 0.5); let carbon = estimate_carbon(energy); @@ -269,3 +328,101 @@ impl Analyzer { recs } } + +#[cfg(test)] +mod tests { + use super::*; + + fn analyze(source: &str) -> Vec { + let mut analyzer = Analyzer::new(Language::Rust).expect("analyzer builds"); + analyzer.analyze_source(source).expect("analysis succeeds") + } + + /// `depth` nested `for` loops — the smallest nest flagged as + /// `nested-loops` is depth 3. The body avoids every other pattern (no + /// `.clone()`, no `File::open`, no `vec![]`, no string `+`, no + /// `to_string`) so exactly one calibrated row prices the unit. + fn nested_loop_work(depth: usize) -> String { + let mut body = + String::from("fn deep_work(items: &[u32]) -> usize {\n let mut total = 0;\n"); + for _ in 0..depth { + body.push_str(" for i in 0..8 {\n"); + } + body.push_str(" total += 1;\n"); + for _ in 0..depth { + body.push_str(" }\n"); + } + body.push_str(" total\n}\n"); + body + } + + /// Straight-line code with no recognised pattern. + fn plain_work(statements: usize) -> String { + let mut body = String::from("fn plain_work() -> usize {\n let mut total = 0;\n"); + for i in 0..statements { + body.push_str(&format!(" total += {};\n", i + 1)); + } + body.push_str(" total\n}\n"); + body + } + + /// Falsifier for the old behaviour: a recognised pattern used to be + /// labelled `Estimated` like everything else, so no finding could ever be + /// `Calibrated` and `--check` could never block. + #[test] + fn recognized_pattern_earns_calibrated_confidence_and_a_band() { + let results = analyze(&nested_loop_work(3)); + assert_eq!(results.len(), 1, "one function, one finding"); + + let result = &results[0]; + assert_eq!(result.confidence, Confidence::Calibrated); + assert_eq!(result.rule_id, "oikosbot/nested-loops"); + + // The band is propagated, not collapsed: min <= typical <= max, and the + // point estimate IS the typical bound. + let range = result + .resource_range + .as_ref() + .expect("a calibrated estimate must carry its band"); + assert!(range.min.energy.0 <= range.typical.energy.0); + assert!(range.typical.energy.0 <= range.max.energy.0); + assert!((range.typical.energy.0 - result.resources.energy.0).abs() < 1e-12); + assert!(range.max.memory.0 >= range.typical.memory.0); + } + + /// Unrecognised code keeps the naive path and is honest about it: no band is + /// claimed where none was computed. + #[test] + fn unrecognized_code_stays_estimated_with_no_band() { + let results = analyze(&plain_work(4)); + assert_eq!(results.len(), 1); + + let result = &results[0]; + assert_eq!(result.confidence, Confidence::Estimated); + assert_eq!(result.rule_id, "oikosbot/general"); + assert!(result.resource_range.is_none()); + + // The naive path is unchanged: a positive, complexity-derived figure + // (0.1 J per AST node) with no band around it. + assert!(result.resources.energy.0 > 0.0); + } + + /// The calibrated path must actually change the numbers, or the wiring + /// would be decoration. Same unit, same size, different evidence. + #[test] + fn calibration_changes_the_estimate() { + let plain = analyze(&plain_work(40)); + let nested = analyze(&nested_loop_work(3)); + + // The naive path charges 0.1 J per node; the calibrated Sort row + // charges microjoules per comparison-bounded operation. A nested-loop + // unit is no longer free just because it is small, and plain code is + // no longer expensive just because it is long. + assert!( + plain[0].resources.energy.0 > nested[0].resources.energy.0, + "naive energy {:.4} J should dwarf calibrated {:.6} J", + plain[0].resources.energy.0, + nested[0].resources.energy.0 + ); + } +} diff --git a/crates/oikosbot-analysis/src/calibration.rs b/crates/oikosbot-analysis/src/calibration.rs index 70699dd..28ffac1 100644 --- a/crates/oikosbot-analysis/src/calibration.rs +++ b/crates/oikosbot-analysis/src/calibration.rs @@ -5,25 +5,15 @@ //! //! Replaces naive `complexity * 0.1 J` with pattern-based resource profiles //! producing ranges (min, typical, max) instead of single numbers. +//! +//! The range type lives in `oikosbot_metrics` so it can be carried on an +//! [`oikosbot_metrics::AnalysisResult`]. A range carries the confidence its row +//! has earned: confidence is a property of the evidence behind an estimate, not +//! of the estimate, and it must be earned per operation kind rather than +//! assigned to a whole file or run. use crate::carbon::estimate_carbon; -use oikosbot_metrics::{Confidence, Duration, Energy, Memory, ResourceProfile}; - -/// Resource estimate with min/typical/max range -#[derive(Debug, Clone)] -pub struct ResourceRange { - pub min: ResourceProfile, - pub typical: ResourceProfile, - pub max: ResourceProfile, - pub confidence: Confidence, -} - -impl ResourceRange { - /// Return the typical estimate as a single profile - pub fn as_typical(&self) -> &ResourceProfile { - &self.typical - } -} +use oikosbot_metrics::{Confidence, Duration, Energy, Memory, ResourceProfile, ResourceRange}; /// Operation categories for calibrated estimates #[derive(Debug, Clone, Copy, PartialEq, Eq)] @@ -50,6 +40,16 @@ pub enum OperationKind { /// /// These are expert estimates that will be refined with profiling data. /// All values assume a modern x86_64 system at ~50W TDP. +/// +/// Returns the min/typical/max envelope, carrying the confidence a finding +/// built on this estimate is entitled to. The ladder is deliberate and +/// per-kind: only `HashLookup`, `Sort`, `Allocation` and `MathCompute` are +/// backed by calibration data, so only those rows are `Calibrated`. `FileIO`, +/// `StringOp` and `NetworkCall` depend on host and workload specifics we have +/// not measured, so they stay `Estimated` even though they use this table. +/// `Generic` is the fallback for code we could not classify and stays +/// `Unknown`. Confidence is therefore earned per finding, never +/// blanket-assigned. pub fn estimate_operation(kind: OperationKind, n: usize) -> ResourceRange { match kind { OperationKind::HashLookup => { @@ -135,31 +135,42 @@ pub fn estimate_operation(kind: OperationKind, n: usize) -> ResourceRange { } } -/// Estimate resources for a function based on its complexity and detected patterns. -pub fn calibrated_estimate(complexity: usize, patterns: &[String]) -> ResourceProfile { - // Start with generic estimate based on complexity - let base = estimate_operation(OperationKind::Generic, complexity); - let mut result = base.typical.clone(); - - // Apply pattern-based adjustments - for pattern in patterns { - let multiplier = match pattern.as_str() { - "nested-loops" => 3.0, - "busy-wait" => 5.0, - "string-concat-in-loop" => 2.0, - "clone-in-loop" => 1.5, - "unbuffered-io" => 3.0, - "large-allocation" => 2.0, - "redundant-allocation" => 1.2, - _ => 1.0, - }; - - result.energy = Energy::joules(result.energy.0 * multiplier); - result.duration = Duration::milliseconds(result.duration.0 * multiplier); - result.carbon = estimate_carbon(result.energy); +/// Map a detected pattern (see `crate::patterns`) onto the operation category +/// that best explains its cost, if any. +/// +/// The mapping is an *approximation* and is documented as such: a pattern tells +/// us what kind of work a unit repeats, and [`estimate_operation`] prices that +/// kind of work. It does not tell us the exact operation count, so `n` stays +/// the AST node count — a size proxy, not an instruction count. +/// +/// Returns `None` for patterns we decline to map. `redundant-allocation` is +/// deliberately unmapped: its impact multiplier (1.2) is marginal and mapping +/// it would let a cosmetic finding inherit `Calibrated` confidence from the +/// allocation row, which would be a blanket assignment rather than an earned +/// one. Unmapped patterns leave the unit on the naive path, labelled +/// `Estimated`. +pub fn operation_for_pattern(pattern: &str) -> Option { + match pattern { + // Repeated comparison-bounded work. A depth-d loop is superlinear in + // its bound the way n·log n is, and comparison-bound is the dominant + // cost of sorting, so Sort is the closest priced category. + "nested-loops" => Some(OperationKind::Sort), + // A spin loop burns CPU without yielding: repeated arithmetic with no + // allocation and no I/O, i.e. the MathCompute row. Note the row's + // per-op cost is small, so a busy-wait is still *relatively* worse + // than plain arithmetic only through its pattern multiplier — the + // absolute figure is the weak part of this mapping. + "busy-wait" => Some(OperationKind::MathCompute), + // Per-iteration copy plus reallocation of a growing buffer. + "string-concat-in-loop" => Some(OperationKind::StringOp), + // Per-iteration heap copy. + "clone-in-loop" => Some(OperationKind::Allocation), + // Unbuffered syscalls, one per iteration. + "unbuffered-io" => Some(OperationKind::FileIO), + // A single large heap allocation. + "large-allocation" => Some(OperationKind::Allocation), + _ => None, } - - result } fn profile(energy_j: f64, duration_ms: f64, memory_bytes: usize) -> ResourceProfile { @@ -184,18 +195,102 @@ mod tests { assert_eq!(range.confidence, Confidence::Calibrated); } - #[test] - fn test_calibrated_estimate_with_patterns() { - let base = calibrated_estimate(10, &[]); - let with_pattern = calibrated_estimate(10, &["nested-loops".to_string()]); - // Pattern should increase energy - assert!(with_pattern.energy.0 > base.energy.0); - } - #[test] fn test_network_call_energy() { let range = estimate_operation(OperationKind::NetworkCall, 1); // Network calls should be significant energy users assert!(range.typical.energy.0 >= 0.1); + // ...but we have not measured them, so they stay Estimated. + assert_eq!(range.confidence, Confidence::Estimated); + } + + /// The confidence ladder must be earned per kind, never blanket-assigned. + #[test] + fn test_confidence_is_per_operation_kind() { + for kind in [ + OperationKind::HashLookup, + OperationKind::Sort, + OperationKind::Allocation, + OperationKind::MathCompute, + ] { + let range = estimate_operation(kind, 10); + assert_eq!( + range.confidence, + Confidence::Calibrated, + "{kind:?} is calibrated data and must report Calibrated" + ); + } + for kind in [ + OperationKind::FileIO, + OperationKind::StringOp, + OperationKind::NetworkCall, + ] { + let range = estimate_operation(kind, 10); + assert_eq!( + range.confidence, + Confidence::Estimated, + "{kind:?} is host-dependent and must stay Estimated" + ); + } + let range = estimate_operation(OperationKind::Generic, 10); + assert_eq!(range.confidence, Confidence::Unknown); + } + + /// Every priced row must keep min <= typical <= max on all four axes, or + /// the propagated band would be meaningless. + #[test] + fn test_range_is_ordered_on_every_axis() { + for kind in [ + OperationKind::HashLookup, + OperationKind::Sort, + OperationKind::FileIO, + OperationKind::NetworkCall, + OperationKind::Allocation, + OperationKind::StringOp, + OperationKind::MathCompute, + OperationKind::Generic, + ] { + let range = estimate_operation(kind, 500); + assert!(range.min.energy.0 <= range.typical.energy.0); + assert!(range.typical.energy.0 <= range.max.energy.0); + assert!(range.min.duration.0 <= range.typical.duration.0); + assert!(range.typical.duration.0 <= range.max.duration.0); + assert!(range.min.memory.0 <= range.typical.memory.0); + assert!(range.typical.memory.0 <= range.max.memory.0); + assert!(range.min.carbon.0 <= range.typical.carbon.0); + assert!(range.typical.carbon.0 <= range.max.carbon.0); + } + } + + #[test] + fn test_operation_for_pattern_mapping() { + assert_eq!( + operation_for_pattern("nested-loops"), + Some(OperationKind::Sort) + ); + assert_eq!( + operation_for_pattern("busy-wait"), + Some(OperationKind::MathCompute) + ); + assert_eq!( + operation_for_pattern("string-concat-in-loop"), + Some(OperationKind::StringOp) + ); + assert_eq!( + operation_for_pattern("clone-in-loop"), + Some(OperationKind::Allocation) + ); + assert_eq!( + operation_for_pattern("unbuffered-io"), + Some(OperationKind::FileIO) + ); + assert_eq!( + operation_for_pattern("large-allocation"), + Some(OperationKind::Allocation) + ); + // Deliberately unmapped: a marginal finding must not inherit + // Calibrated confidence from the allocation row. + assert_eq!(operation_for_pattern("redundant-allocation"), None); + assert_eq!(operation_for_pattern("not-a-pattern"), None); } } diff --git a/crates/oikosbot-analysis/src/dependencies.rs b/crates/oikosbot-analysis/src/dependencies.rs index 7f60fe6..adf2f5c 100644 --- a/crates/oikosbot-analysis/src/dependencies.rs +++ b/crates/oikosbot-analysis/src/dependencies.rs @@ -253,6 +253,7 @@ fn dep_finding_to_result(finding: DepFinding) -> AnalysisResult { end_location: None, confidence: Confidence::Estimated, pareto: None, + resource_range: None, } } diff --git a/crates/oikosbot-analysis/src/security.rs b/crates/oikosbot-analysis/src/security.rs index 2b83d0f..ef3da5a 100644 --- a/crates/oikosbot-analysis/src/security.rs +++ b/crates/oikosbot-analysis/src/security.rs @@ -125,6 +125,7 @@ mod inner { end_location: None, confidence: Confidence::Estimated, pareto: None, + resource_range: None, }; security_findings.push(finding); diff --git a/crates/oikosbot-capability/README.adoc b/crates/oikosbot-capability/README.adoc new file mode 100644 index 0000000..e984d90 --- /dev/null +++ b/crates/oikosbot-capability/README.adoc @@ -0,0 +1,34 @@ +// SPDX-License-Identifier: CC-BY-SA-4.0 +// SPDX-FileCopyrightText: 2026 Jonathan D.A. Jewell += `oikosbot-capability` +:toc: macro + +Verified capability from run history alone. A gate with N successes and zero +failures in its entire history cannot be shown to be *able* to fail, so its +green is not evidence and its successes are excluded from verified output. +This is the operationalisation of the estate's fake-gate pathology. + +== Public surface + +`assess(runs, releases, min_n) -> Vec`, where each row carries +run totals, startup failures, dead (never-parsed) runs, workflow counts, and +`infallible_gate_candidates`. + +`infallible_gate_candidates` is keyed on **workflow path**, not the free-text +workflow name. Two different files can share a name (two workflows both titled +"CI"), and grouping by name merges their run counts — which can hide a +fake-gate candidate behind a same-named failing workflow. Measured on +2026-08-07: by path 1482 candidates across 313 repos, by name 1270. The +212-candidate gap was the concealment, and it is why the shipped key is the +path. + +== What it does not do + +It does not decide that a gate is fake — it lists candidates that need a human +look, and it excludes their successes from verified output. + +== Tests + +---- +cargo test -p oikosbot-capability +---- diff --git a/crates/oikosbot-cli/README.adoc b/crates/oikosbot-cli/README.adoc new file mode 100644 index 0000000..a4a45cd --- /dev/null +++ b/crates/oikosbot-cli/README.adoc @@ -0,0 +1,64 @@ +// SPDX-License-Identifier: CC-BY-SA-4.0 +// SPDX-FileCopyrightText: 2026 Jonathan D.A. Jewell += `oikosbot-cli` +:toc: macro + +The `oikosbot` binary — the product surface. Every other crate is a library; +this is what the container image runs and what the composite action invokes. + +== Subcommands + +[cols="1,4"] +|=== +| `check ` | Analyse a directory and compare each finding's eco score +against the floor. Below-floor findings are listed; the run exits 1 only when +the config makes enforcement blocking (`mode: regulator` or +`enforcement: blocking`). + +| `report ` | Same analysis, emit SARIF/JSON/text. This is the action's +default mode. + +| `compare ` | Pareto verdict for a diff: improvement, regression, +trade-off or neutral, with per-objective deltas and the confidence that backs +them. `--check` enforces the trade-off doctrine; `--pr-body` supplies the PR +description to check for a documented `Pareto-Trade-off:`. + +| `estate collect|analyse|report` | Estate telemetry, capability and DEA +analysis. Read-only by design. + +| `analyze ` | Analyse a single file. + +| `self-analyze` | Dogfooding: run the analyser over this repository. + +| `fleet ` | Prints how to run the optional `oikosbot-fleet` bridge, +which is deliberately *not* a dependency of this binary. +|=== + +== Configuration + +`--config ` or an auto-discovered `.oikos.yml` next to the analysed +directory. Keys that take effect: `mode`, `thresholds.eco_minimum.carbon` / +`.energy`, `enforcement`, `exclude`, `analysis.languages`. Everything else in +`config/oikos.yaml` is parsed and discarded (documented as aspirational there). + +== The enforcement doctrine + +`compare --check` exits 1 on an undocumented regression or trade-off **only +when every driving objective is `Measured` or `Calibrated`**. When it cannot +enforce, it prints a `::warning::` naming the verdict and the confidence, and +exits 0 — never silently. A gate that quietly does nothing is indistinguishable +from one that passed. + +=== Integration tests + +`tests/check_gate.rs` proves the doctrine end to end against the real binary: +an undocumented calibrated regression exits 1, the same regression without +`--check` exits 0, a documented trade-off is accepted, and a heuristic +regression refuses to block *and says so*. + +== Tests + +---- +cargo test -p oikosbot-cli +cargo test -p oikosbot-cli --test check_gate +---- diff --git a/crates/oikosbot-cli/src/main.rs b/crates/oikosbot-cli/src/main.rs index 4ab4464..a8dbb04 100644 --- a/crates/oikosbot-cli/src/main.rs +++ b/crates/oikosbot-cli/src/main.rs @@ -409,7 +409,7 @@ fn main() -> Result<()> { excluded from the default workspace so OikosBot builds standalone.\n\n\ To run it (with hyperpolymath/gitbot-fleet checked out as a sibling):\n \ cargo run --manifest-path crates/oikosbot-fleet/Cargo.toml -- {} {}\n\n\ - See crates/oikosbot-fleet/README.md and DISAMBIGUATION.adoc.", + See crates/oikosbot-fleet/README.adoc and DISAMBIGUATION.adoc.", path.display(), ctx, ); diff --git a/crates/oikosbot-cli/tests/check_gate.rs b/crates/oikosbot-cli/tests/check_gate.rs new file mode 100644 index 0000000..d1410d7 --- /dev/null +++ b/crates/oikosbot-cli/tests/check_gate.rs @@ -0,0 +1,171 @@ +// SPDX-License-Identifier: MPL-2.0 +// SPDX-FileCopyrightText: 2025 Jonathan D.A. Jewell + +//! End-to-end proof that `oikosbot compare --check` can actually block. +//! +//! Before the calibration wiring every finding was `Estimated`, so `--check` +//! could only ever print the inert-enforcement warning: the gate could not +//! fail. These tests pin the behaviour that makes the trade-off doctrine real — +//! +//! * a regression whose driving objectives are calibrated **blocks**; +//! * the same regression without `--check` stays advisory; +//! * a documented trade-off is accepted; +//! * a heuristic (unrecognised) regression still refuses to block, and says so +//! loudly rather than passing silently. +//! +//! Fixtures are sized so both sides land on a calibrated row: base is a +//! depth-3 loop nest, head the same nest one level deeper. Both are therefore +//! `Calibrated`, and the deeper nest is worse on every objective, so the +//! verdict is a genuine `Regression` with real drivers behind it. + +use std::fs; +use std::path::Path; +use std::process::{Command, Output}; +use tempfile::TempDir; + +/// Write a throwaway source tree and return its handle. +fn tree(files: &[(&str, &str)]) -> TempDir { + let dir = TempDir::new().expect("temp dir"); + for (name, body) in files { + fs::write(dir.path().join(name), body).expect("write fixture"); + } + dir +} + +/// A Rust function whose body is `depth` nested `for` loops. +/// +/// Depth 3 is the smallest nest `detect_patterns` flags as `nested-loops`, +/// which maps onto the calibrated `Sort` row. The body avoids every other +/// pattern — no `.clone()`, no `File::open`, no `vec![]`, no string `+`, no +/// `to_string` — so exactly one calibrated row prices the unit and nothing +/// drags its confidence back to `Estimated`. +fn nested_loop_work(depth: usize) -> String { + let mut body = String::from("fn deep_work(items: &[u32]) -> usize {\n let mut total = 0;\n"); + for _ in 0..depth { + body.push_str(" for i in 0..8 {\n"); + } + body.push_str(" total += 1;\n"); + for _ in 0..depth { + body.push_str(" }\n"); + } + body.push_str(" total\n}\n"); + body +} + +/// A function with no recognised pattern: `statements` straight-line +/// increments. Both sides of a comparison using this stay on the naive, +/// complexity-derived path and are therefore `Estimated`. +fn plain_work(statements: usize) -> String { + let mut body = String::from("fn plain_work() -> usize {\n let mut total = 0;\n"); + for i in 0..statements { + body.push_str(&format!(" total += {};\n", i + 1)); + } + body.push_str(" total\n}\n"); + body +} + +fn compare(base: &Path, head: &Path, extra: &[&str]) -> Output { + Command::new(env!("CARGO_BIN_EXE_oikosbot")) + .arg("compare") + .arg(base) + .arg(head) + .args(extra) + .output() + .expect("spawn oikosbot") +} + +fn stdout(out: &Output) -> String { + String::from_utf8_lossy(&out.stdout).to_string() +} + +fn stderr(out: &Output) -> String { + String::from_utf8_lossy(&out.stderr).to_string() +} + +#[test] +fn undocumented_calibrated_regression_blocks_under_check() { + let base = tree(&[("base.rs", &nested_loop_work(3))]); + let head = tree(&[("head.rs", &nested_loop_work(4))]); + + let out = compare(base.path(), head.path(), &["--check"]); + + assert_eq!( + out.status.code(), + Some(1), + "an undocumented calibrated regression must exit 1; stdout: {}", + stdout(&out) + ); + let stdout = stdout(&out); + assert!( + stdout.contains("PARETO REGRESSION"), + "expected a regression verdict; stdout: {stdout}" + ); + assert!( + stdout.contains("Confidence: Calibrated"), + "the drivers must be calibrated, not heuristic; stdout: {stdout}" + ); +} + +#[test] +fn same_regression_is_advisory_without_check() { + let base = tree(&[("base.rs", &nested_loop_work(3))]); + let head = tree(&[("head.rs", &nested_loop_work(4))]); + + let out = compare(base.path(), head.path(), &[]); + + assert_eq!( + out.status.code(), + Some(0), + "without --check the run is advisory; stderr: {}", + stderr(&out) + ); + assert!(stdout(&out).contains("PARETO REGRESSION")); +} + +#[test] +fn documented_regression_is_accepted_under_check() { + let base = tree(&[("base.rs", &nested_loop_work(3))]); + let head = tree(&[ + ("head.rs", &nested_loop_work(4)), + ( + "pr.md", + "Summary\n\nPareto-Trade-off: depth-4 nest kept for cache locality.\n", + ), + ]); + // The CLI resolves --pr-body relative to its own working directory, so the + // fixture must be addressed absolutely. + let pr_body = head.path().join("pr.md"); + + let out = compare( + base.path(), + head.path(), + &["--check", "--pr-body", pr_body.to_str().unwrap()], + ); + + assert_eq!( + out.status.code(), + Some(0), + "a documented trade-off must be accepted; stdout: {}", + stdout(&out) + ); +} + +#[test] +fn heuristic_regression_refuses_to_block_and_says_so() { + let base = tree(&[("base.rs", &plain_work(2))]); + let head = tree(&[("head.rs", &plain_work(8))]); + + let out = compare(base.path(), head.path(), &["--check"]); + + assert_eq!( + out.status.code(), + Some(0), + "heuristic estimates advise, they do not block; stdout: {}", + stdout(&out) + ); + assert!( + stderr(&out).contains("NOT enforced"), + "a gate that cannot fire must say so unmissably; stderr: {}", + stderr(&out) + ); +} diff --git a/crates/oikosbot-dea/README.adoc b/crates/oikosbot-dea/README.adoc new file mode 100644 index 0000000..de2fbea --- /dev/null +++ b/crates/oikosbot-dea/README.adoc @@ -0,0 +1,34 @@ +// SPDX-License-Identifier: CC-BY-SA-4.0 +// SPDX-FileCopyrightText: 2026 Jonathan D.A. Jewell += `oikosbot-dea` +:toc: macro + +Data Envelopment Analysis over estate entities: allocative and productive +efficiency from observed inputs and outputs, solved as paired linear programs +(`good_lp` + HiGHS). + +== Public surface + +* `Dmu { name, inputs, outputs }` — one decision-making unit (typically one + repository). +* `dea(dmus) -> Vec` — per unit: `theta_ccr` (constant returns to + scale), `theta_bcc` (variable returns), `theta_mult` (the multiplier form), + `peers` (the reference set with positive envelopment weight), and the input + and output shadow prices (`v_i`, `u_r`). +* Strong duality is asserted, not assumed: `|theta_ccr - theta_mult|` must agree + within tolerance, which is the implementation's own correctness check. + +== What it does not do + +DEA is *relative*: a score is a unit's standing against the observed peer set, +so a new snapshot can move a score without the unit changing. The solver has +been validated against self-authored analytic cases (a closed-form single +input/output, a strictly dominated unit, and the duality assertion) — not yet +against a published worked example with a citation, which is the standard that +would make it externally defensible. + +== Tests + +---- +cargo test -p oikosbot-dea +---- diff --git a/crates/oikosbot-eclexia/README.adoc b/crates/oikosbot-eclexia/README.adoc new file mode 100644 index 0000000..3c5daeb --- /dev/null +++ b/crates/oikosbot-eclexia/README.adoc @@ -0,0 +1,42 @@ +// SPDX-License-Identifier: CC-BY-SA-4.0 +// SPDX-FileCopyrightText: 2026 Jonathan D.A. Jewell += `oikosbot-eclexia` +:toc: macro + +Policy adapter: evaluates the Eclexia (`.ecl`) policies under `policies/` +against a set of analysis results and returns `PolicyDecision`s. + +[IMPORTANT] +==== +**This backend is a seam, not an implementation.** With the default features it +shells out to an `eclexia` binary if one is on `PATH`; if none is found it falls +back to a builtin evaluator that matches policies by **file stem** and applies +hardcoded thresholds — the `.ecl` contents are never parsed. The hardcoded +thresholds can and do disagree with the policy text +(`energy_threshold.ecl` declares `> 50.0 J` per function; the builtin fires at +`> 1000 J` total). Every builtin evaluation emits a `::warning::` saying exactly +that (`builtin_warning`). It is loud, and it is still fake. Tracked as debt. +==== + +== Public surface + +* `evaluate_policies(policy_dir, results)` — one `PolicyDecision` per `.ecl` + file, plus a `Warn` decision for any policy that errors. +* `decisions_to_results(decisions)` — converts decisions into synthetic + `AnalysisResult`s (`oikosbot/policy-`) so they join the Pareto and SARIF + paths like any other finding. +* `builtin_warning(policy_stem)` — the loud notice quoted above. +* `EXAMPLE_POLICY` — a documented example `.ecl` for tests and docs. + +== The `eclexia-native` feature + +Reserved and **unbuildable**: the `cfg` branch calls into `eclexia_parser` / +`eclexia_interp`, which are not declared dependencies. It fails with an explicit +"unavailable" error rather than compiling a stub, so a successful build can +never claim native integration. + +== Tests + +---- +cargo test -p oikosbot-eclexia +---- diff --git a/crates/oikosbot-eclexia/src/lib.rs b/crates/oikosbot-eclexia/src/lib.rs index 0c3c32b..08cad5d 100644 --- a/crates/oikosbot-eclexia/src/lib.rs +++ b/crates/oikosbot-eclexia/src/lib.rs @@ -409,6 +409,7 @@ pub fn decisions_to_results(decisions: &[PolicyDecision]) -> Vec end_location: None, confidence: oikosbot_metrics::Confidence::Estimated, pareto: None, + resource_range: None, } }) .collect() @@ -531,6 +532,7 @@ mod tests { end_location: None, confidence: oikosbot_metrics::Confidence::Estimated, pareto: None, + resource_range: None, } } } diff --git a/crates/oikosbot-fleet/src/lib.rs b/crates/oikosbot-fleet/src/lib.rs index 46cdcd8..18c0e85 100644 --- a/crates/oikosbot-fleet/src/lib.rs +++ b/crates/oikosbot-fleet/src/lib.rs @@ -384,6 +384,7 @@ mod tests { end_location: Some((10, 2)), confidence: Confidence::Estimated, pareto: None, + resource_range: None, } } diff --git a/crates/oikosbot-metrics/README.adoc b/crates/oikosbot-metrics/README.adoc new file mode 100644 index 0000000..d0fdeea --- /dev/null +++ b/crates/oikosbot-metrics/README.adoc @@ -0,0 +1,42 @@ +// SPDX-License-Identifier: CC-BY-SA-4.0 +// SPDX-FileCopyrightText: 2026 Jonathan D.A. Jewell += `oikosbot-metrics` +:toc: macro + +The shared vocabulary of the workspace: units, scores, and the record every +other crate passes around. No analysis logic lives here — if a number has a +unit or a confidence, its type is defined in this crate. + +== What it owns + +* **Units** — `Energy` (J), `Duration` (ms), `Carbon` (gCO2e), `Memory` + (bytes), each a newtype over the base unit with constructors + (`Energy::joules`, `Carbon::kilograms_co2e`, …). Constructors, not bare + floats, so a unit mix-up cannot compile. +* **`ResourceProfile`** — the four-axis profile for one code unit, plus + `Add` and `cost(&ShadowPrices)` for the Eclexia-style shadow-price + valuation. +* **`ResourceRange`** — the min/typical/max envelope behind a calibrated + estimate. +* **Scores** — `EcoScore`, `EconScore` (both clamped 0–100) and + `HealthIndex::compute(eco, econ, quality)`, whose `overall` is + `0.4·eco + 0.3·econ + 0.3·quality`. +* **`Confidence`** — the ladder `Measured > Calibrated > Estimated > + `Unknown`, ordered by `PartialOrd`. Only `Measured` and `Calibrated` may + drive a blocking decision anywhere in the system. +* **`AnalysisResult`** / **`CodeLocation`** / **`ParetoInfo`** — the finding + record. Every field after `resources` is `#[serde(default)]`, so a JSON + document written by an older version still loads; `resource_range` is + additionally skipped when absent, keeping the wire format stable. + +== What it does not do + +It does not compute anything from source. `EcoScore::new` clamps; it does not +judge. Producing a number is `oikosbot-analysis`'s job; interpreting a set of +them is `oikosbot-pareto`'s. + +== Tests + +---- +cargo test -p oikosbot-metrics +---- diff --git a/crates/oikosbot-metrics/src/lib.rs b/crates/oikosbot-metrics/src/lib.rs index bab7fd3..325a5ff 100644 --- a/crates/oikosbot-metrics/src/lib.rs +++ b/crates/oikosbot-metrics/src/lib.rs @@ -132,6 +132,27 @@ pub struct ResourceProfile { pub memory: Memory, } +/// Min/typical/max envelope for a resource estimate, with the confidence the +/// row that produced it has earned. +/// +/// [`AnalysisResult::resources`] collapses an estimate to a single point (the +/// `typical` bound). The envelope is kept alongside it so consumers can see the +/// uncertainty band instead of treating the point estimate as exact. It is +/// present only when the estimate came from a calibrated operation profile (see +/// `oikosbot_analysis::calibration::estimate_operation`); the naive +/// `complexity`-derived path produces a point estimate with no band and leaves +/// this `None`. +#[derive(Debug, Clone, Serialize, Deserialize)] +pub struct ResourceRange { + pub min: ResourceProfile, + pub typical: ResourceProfile, + pub max: ResourceProfile, + /// Confidence is a property of the evidence behind an estimate, so it rides + /// with the row that produced it — and is earned per operation kind, never + /// assigned to a whole file or run. + pub confidence: Confidence, +} + impl ResourceProfile { pub fn zero() -> Self { ResourceProfile { @@ -282,6 +303,11 @@ pub struct AnalysisResult { /// Pareto pass; None when analyzed in isolation). #[serde(default)] pub pareto: Option, + /// Uncertainty band behind `resources`, when the estimate is calibrated + /// (see [`ResourceRange`]). `None` on the naive `complexity`-derived path, + /// where `resources` is a point estimate with no band. + #[serde(default, skip_serializing_if = "Option::is_none")] + pub resource_range: Option, } /// Source code location diff --git a/crates/oikosbot-pareto/README.adoc b/crates/oikosbot-pareto/README.adoc new file mode 100644 index 0000000..1f2ebfc --- /dev/null +++ b/crates/oikosbot-pareto/README.adoc @@ -0,0 +1,38 @@ +// SPDX-License-Identifier: CC-BY-SA-4.0 +// SPDX-FileCopyrightText: 2026 Jonathan D.A. Jewell += `oikosbot-pareto` +:toc: macro + +Multi-objective comparison: dominance, frontiers, scores, and the base-vs-head +verdict that the `compare` subcommand reports. + +== Public surface + +* **Objectives** — `result_objectives()` is the live five-axis set the Rust + analyser actually measures: carbon, energy, duration, memory (minimise) and + the quality score as the maintainability proxy (maximise). +* **Dominance and frontiers** — `dominates`, `frontier_indices`, + `dominated_by_counts`, `normalize`, `pareto_scores`, `refactor_candidates`, + all ε-tolerant (`DEFAULT_EPSILON = 1e-6`). +* **Verdicts** — `compare` classifies head vs base as `Improvement`, + `Regression`, `TradeOff` or `Neutral`; `assess` adds the per-objective + deltas, the list of *drivers* (objectives that actually moved), and + `actionable`. +* **Confidence gating** — `actionable` is true only when **every driving** + objective is backed by `Measured` or `Calibrated` input. Heuristic estimates + advise; they do not block. `aggregate_confidence` and `weaker_confidence` + combine confidences across a result set. +* **Wiring** — `apply_to_results` runs the intra-repo pass (frontier + membership, `ParetoScore`, and the full `EconScore` composition); + `result_point`, `aggregate_point` project results into objective space. + +== What it does not do + +It never decides whether a finding is *true* — it compares numbers it is given. +`actionable` is a statement about evidence, not about severity. + +== Tests + +---- +cargo test -p oikosbot-pareto +---- diff --git a/crates/oikosbot-pareto/src/lib.rs b/crates/oikosbot-pareto/src/lib.rs index 72e8009..303ef67 100644 --- a/crates/oikosbot-pareto/src/lib.rs +++ b/crates/oikosbot-pareto/src/lib.rs @@ -73,20 +73,6 @@ impl Objective { } } -/// The seven standard OikosBot objectives, as specified in -/// `Eco.Pareto.standardObjectives` (ARCHITECTURE.adoc). -pub fn standard_objectives() -> Vec { - vec![ - Objective::new("carbon_intensity", Direction::Minimize, 0.20), - Objective::new("energy_consumption", Direction::Minimize, 0.15), - Objective::new("execution_time", Direction::Minimize, 0.15), - Objective::new("memory_usage", Direction::Minimize, 0.10), - Objective::new("maintainability", Direction::Maximize, 0.15), - Objective::new("test_coverage", Direction::Maximize, 0.10), - Objective::new("technical_debt", Direction::Minimize, 0.15), - ] -} - /// The objective set used when treating per-unit [`AnalysisResult`]s as points /// in objective space. Only axes the Rust analyzer actually measures today: /// the four `ResourceProfile` axes (minimize) plus the quality score diff --git a/crates/oikosbot-sarif/README.adoc b/crates/oikosbot-sarif/README.adoc new file mode 100644 index 0000000..a0a765a --- /dev/null +++ b/crates/oikosbot-sarif/README.adoc @@ -0,0 +1,42 @@ +// SPDX-License-Identifier: CC-BY-SA-4.0 +// SPDX-FileCopyrightText: 2026 Jonathan D.A. Jewell += `oikosbot-sarif` +:toc: macro + +Emits SARIF 2.1.0 from `AnalysisResult`s, so GitHub code scanning can ingest +OikosBot's findings. + +== Shape + +* One run, driver `oikosbot`, `semantic_version` from the caller. +* `tool.driver.rules` carries the builtin rule catalogue (`oikosbot/nested-loops`, + `oikosbot/busy-wait`, `oikosbot/string-concat-in-loop`, `oikosbot/clone-in-loop`, + `oikosbot/unbuffered-io`, `oikosbot/large-allocation`, + `oikosbot/redundant-allocation`, `oikosbot/eco-threshold`, + `oikosbot/carbon-intensity`, `oikosbot/security-sustainability`) with default + levels. Findings whose `rule_id` is not in the catalogue still emit, with no + `ruleIndex`. +* `level` is derived from the eco score: `error` below 30, `warning` below 60, + else `note`. +* `region` carries start **and** end line/column, so the whole function is + annotated rather than its first line. +* `properties` is the machine-readable payload: `eco_score`, `econ_score`, + `quality_score`, `overall_health`, `energy_joules`, `carbon_gco2e`, + `duration_ms`, `memory_bytes`, `confidence`, optional `suggestion`, optional + `pareto_status` / `pareto_score` / `pareto_dominated_by`, and — for + calibrated estimates — `resource_range` with `min`/`typical`/`max` per axis. +* Suggestions ride in the message and in `properties`. There is deliberately no + SARIF `fix` object: OikosBot does not rewrite code. + +== What it does not do + +`oikosbot/general` records are per-function telemetry, not defects, and are +**filtered out of the SARIF results** — analysing a repository must not raise a +code-scanning alert for every function in it. They still appear in `json` and +`text` output. + +== Tests + +---- +cargo test -p oikosbot-sarif +---- diff --git a/crates/oikosbot-sarif/src/lib.rs b/crates/oikosbot-sarif/src/lib.rs index 46ff184..a0e7a33 100644 --- a/crates/oikosbot-sarif/src/lib.rs +++ b/crates/oikosbot-sarif/src/lib.rs @@ -391,6 +391,33 @@ fn convert_result(result: &AnalysisResult, rule_ids: &[&str]) -> SarifResult { properties["pareto_score"] = serde_json::json!(pareto.score); properties["pareto_dominated_by"] = serde_json::json!(pareto.dominated_by); } + // Propagate the calibrated uncertainty band so consumers can see the spread + // instead of treating the point estimate as exact. Absent on the naive + // path, where there is no band to report. + if let Some(ref range) = result.resource_range { + properties["resource_range"] = serde_json::json!({ + "energy_joules": { + "min": range.min.energy.0, + "typical": range.typical.energy.0, + "max": range.max.energy.0, + }, + "carbon_gco2e": { + "min": range.min.carbon.0, + "typical": range.typical.carbon.0, + "max": range.max.carbon.0, + }, + "duration_ms": { + "min": range.min.duration.0, + "typical": range.typical.duration.0, + "max": range.max.duration.0, + }, + "memory_bytes": { + "min": range.min.memory.0, + "typical": range.typical.memory.0, + "max": range.max.memory.0, + }, + }); + } SarifResult { rule_id: rule_id.clone(), @@ -430,6 +457,7 @@ mod tests { end_location: Some((25, 2)), confidence: Confidence::Estimated, pareto: None, + resource_range: None, } } diff --git a/crates/oikosbot-telemetry/README.adoc b/crates/oikosbot-telemetry/README.adoc new file mode 100644 index 0000000..3f8426b --- /dev/null +++ b/crates/oikosbot-telemetry/README.adoc @@ -0,0 +1,41 @@ +// SPDX-License-Identifier: CC-BY-SA-4.0 +// SPDX-FileCopyrightText: 2026 Jonathan D.A. Jewell += `oikosbot-telemetry` +:toc: macro + +Estate-level telemetry: collect GitHub Actions run history, derive energy, +carbon and cost from it, and write Parquet + JSON snapshots. This is the one +part of OikosBot that measures rather than estimates. + +== Modules + +[cols="1,4"] +|=== +| `collect` | `collect_runs`, `list_repos`, `collect_releases` — read-only +collectors that shell out to the `gh` CLI through the `GhRunner` trait +(`GhCli` is the production implementation, so tests can substitute a fake). + +| `rows` | `RunRow`, `RepoRow`, `ReleaseRow` — the flat, serialisable records +everything downstream consumes. + +| `derive` | `derive_per_repo(runs, assumptions)` — per-repo aggregates, and +`confidence_of(metric)` which states the confidence of each derived figure: +`wall_minutes` is `Measured` (it comes straight from the GitHub API), while +energy, carbon and cost are `Calibrated`/`Estimated` from `Assumptions`. + +| `snapshot` | `write_runs`/`read_runs`, `write_repos`/`read_repos`, +`write_releases`/`read_releases` — Parquet staging. +|=== + +== What it does not do + +It never gates anything: `oikosbot estate` reports. Enforcement lives in the +per-file `check`/`compare` path. Snapshots are written outside this repository +(`hyperpolymath/oikosbot-estate`) so the tool never silently measures a corpus +that contains itself. + +== Tests + +---- +cargo test -p oikosbot-telemetry +---- diff --git a/docs/COMPARISON-climate-warrior.adoc b/docs/COMPARISON-climate-warrior.adoc index 416b430..51b2df2 100644 --- a/docs/COMPARISON-climate-warrior.adoc +++ b/docs/COMPARISON-climate-warrior.adoc @@ -87,10 +87,13 @@ Warrior to store documents with an SCI ledger, OikosBot to review code. == What OikosBot does not do (honesty section) -* Estimates are **static heuristics** today (`Confidence::Estimated`), not - measurements. Runtime calibration and the VeriSimDB-backed praxis loop are - roadmap items. This is exactly why the confidence gate exists: heuristics - advise; only measured or calibrated figures may ever block a merge. +* Estimates are **static** today, not measurements: a recognised pattern is + priced from the calibration table (`Calibrated`), and anything unclassified + stays on the naive complexity heuristic (`Estimated`). Neither has been + validated against profiling data, and runtime calibration plus the + VeriSimDB-backed praxis loop are roadmap items. This is exactly why the + confidence gate exists: heuristics advise; only measured or calibrated + figures may ever block a merge. * The GitHub **App** (webhook, PR comments without CI) is not yet live — CI Action mode is the generation-1 surface. The App is gated on upstream AffineScript stdlib work by deliberate ruling (no interim listener). diff --git a/docs/README.adoc b/docs/README.adoc index 88ee941..844c1fb 100644 --- a/docs/README.adoc +++ b/docs/README.adoc @@ -28,20 +28,44 @@ fleet slot (see link:../DISAMBIGUATION.adoc[DISAMBIGUATION]). * link:../policies/README.adoc[policies/README] — Eclexia (`.ecl`) policies and the finding taxonomy. * link:../bot-integration-affine/README.adoc[bot-integration-affine] — the AffineScript webhook receiver (scaffold). * link:../.github/CONTRIBUTING.md[CONTRIBUTING] — fork / branch / PR flow. -* _Planned (link:https://github.com/hyperpolymath/oikosbot/issues/16[#16]): per-crate READMEs, the Haskell analyzer guide, an end-to-end build walkthrough._ +* link:ci-runbook.adoc[CI runbook] — every gate, what it enforces, and how to + read a red check when the logs are unavailable. +* _Crate guides_ — one README per workspace crate, each stating what it owns, + its public surface, and what it does **not** do: + link:../crates/oikosbot-metrics/README.adoc[metrics], + link:../crates/oikosbot-analysis/README.adoc[analysis], + link:../crates/oikosbot-pareto/README.adoc[pareto], + link:../crates/oikosbot-sarif/README.adoc[sarif], + link:../crates/oikosbot-eclexia/README.adoc[eclexia], + link:../crates/oikosbot-telemetry/README.adoc[telemetry], + link:../crates/oikosbot-capability/README.adoc[capability], + link:../crates/oikosbot-dea/README.adoc[dea], + link:../crates/oikosbot-cli/README.adoc[cli], + plus the optional + link:../crates/oikosbot-fleet/README.adoc[fleet bridge] (excluded from the + default workspace). +* link:../analyzers/code-haskell/README.adoc[Haskell analyzer] — what + `analyzers/code-haskell` computes and how it relates to the Rust workspace. +* _Still open (link:https://github.com/hyperpolymath/oikosbot/issues/16[#16]): an end-to-end build walkthrough._ == For maintainers * link:../GOVERNANCE.adoc[GOVERNANCE] + link:../MAINTAINERS.adoc[MAINTAINERS] — governance model and roles. * link:../ROADMAP.adoc[ROADMAP] — phases and milestones. * link:../.machine_readable/descriptiles/README.adoc[.machine_readable/descriptiles] — the machine-readable state/runbook (see below). -* _Planned (link:https://github.com/hyperpolymath/oikosbot/issues/16[#16]): CI/workflow runbook, `just`-target reference, pre-release checklist._ +* link:ci-runbook.adoc[CI runbook] — the gate inventory, `actions.lock` rules, + how to diagnose Hypatia and `Publish Image` failures, the `just` target + reference, and the pre-release (GitHub Marketplace) checklist. +* _Still open (link:https://github.com/hyperpolymath/oikosbot/issues/16[#16]): the `README` → `docs/` split._ == For end users * link:GITHUB_APP_SETUP.adoc[GitHub App setup] — install and configure the App. * link:../SECURITY.md[SECURITY] — vulnerability reporting and data handling. -* _Planned (link:https://github.com/hyperpolymath/oikosbot/issues/17[#17]): interpreting findings/SARIF, `BOT_MODE`, policy customization, troubleshooting, and a production deploy runbook (today link:../DEPLOY.adoc[DEPLOY] is a stub)._ +* link:usage.adoc[Using OikosBot] — action inputs, the three modes, what each + verdict and confidence level means, the SARIF shape, `.oikos.yml` + configuration, `BOT_MODE`, and a troubleshooting table. +* _Still open (link:https://github.com/hyperpolymath/oikosbot/issues/17[#17]): a production deploy runbook — link:../DEPLOY.adoc[DEPLOY] is still a stub, gated on AffineScript operational parity._ == Machine-readable docs (A2ML) @@ -56,8 +80,9 @@ The canonical project state and decisions live in `.machine_readable/descriptile == Current state (read first) * link:STATUS.adoc[**STATUS**] — what works, what does *not*, and how each claim - was verified. Includes the standing caveat that resource figures are heuristic - estimates and `--check` cannot yet block a merge. + was verified. Includes the standing caveat that resource figures are static + estimates, that only calibrated findings may block a merge, and that the eco + score's absolute scale saturates under calibrated magnitudes. * link:COMPARISON-climate-warrior.adoc[**OikosBot vs Climate Warrior**] — positioning against the nearest GitHub Marketplace neighbour. @@ -69,4 +94,7 @@ The canonical project state and decisions live in `.machine_readable/descriptile * **VeriSimDB** octad data layer (deferred) — link:verisimdb-client.adoc[verisimdb-client] + `META.a2ml` ADR-001. * **Provenance & the PMPL license** — link:../PALIMPSEST.adoc[PALIMPSEST]. -_Audience-depth gaps are tracked in link:https://github.com/hyperpolymath/oikosbot/issues/16[#16] (dev/maintainer) and link:https://github.com/hyperpolymath/oikosbot/issues/17[#17] (end-user); the README→`docs/` split is also in #16._ +_Remaining audience-depth gaps: the README→`docs/` split and the end-to-end +build walkthrough in link:https://github.com/hyperpolymath/oikosbot/issues/16[#16], +and the production deploy runbook in +link:https://github.com/hyperpolymath/oikosbot/issues/17[#17]._ diff --git a/docs/STATUS.adoc b/docs/STATUS.adoc index ac1b61e..0f3a88c 100644 --- a/docs/STATUS.adoc +++ b/docs/STATUS.adoc @@ -101,20 +101,26 @@ The economic core (`crates/oikosbot-pareto`) is real and executable: [IMPORTANT] ==== -**Every resource figure is a heuristic estimate, and enforcement is inert.** - -`Analyzer::estimate_resources()` is still `complexity × 0.1 J` — the naive form -that `calibration.rs` was written to replace but is not yet wired to. So every -result carries `Confidence::Estimated`, and since only `Measured` or -`Calibrated` inputs may fail a run, **`--check` always exits 0**. - -It says so loudly — a `::warning::` annotation naming the verdict and the -confidence that blocked enforcement — rather than passing silently. A gate that -quietly does nothing is indistinguishable from one that passed, which is exactly -the failure mode OikosBot exists to find. - -Tracked in issue #48. Until it lands, treat OikosBot as an advisor, not a -regulator. +**Resource figures are estimates, and enforcement is earned per finding.** + +`Analyzer::estimate_resources()` has two paths. A unit with a recognised +pattern is priced from `calibration::estimate_operation()` and carries that +row's confidence — `Calibrated` for the measured operation kinds, `Estimated` +for the host-dependent ones. A unit with no recognised pattern stays on the +naive `complexity × 0.1 J` form and is labelled `Estimated`. + +Since only `Measured` or `Calibrated` inputs may fail a run, `--check` blocks +on an undocumented regression whose drivers are calibrated, and refuses to +block on heuristic drivers — saying so loudly in a `::warning::` rather than +passing silently. A gate that quietly does nothing is indistinguishable from one +that passed, which is exactly the failure mode OikosBot exists to find. + +Wired in issue #48. What is *not* earned yet: the absolute figures have never +been validated against profiling data, and the eco score's log scale is anchored +at 1 J, so calibrated (microjoule-scale) estimates clamp near 100 — the eco +threshold therefore discriminates far less than the Pareto verdicts, which +compare base against head. Treat OikosBot as an advisor with a working +regulator, not the reverse. ==== Also incomplete: diff --git a/docs/ci-runbook.adoc b/docs/ci-runbook.adoc new file mode 100644 index 0000000..3f00e02 --- /dev/null +++ b/docs/ci-runbook.adoc @@ -0,0 +1,227 @@ +// SPDX-License-Identifier: CC-BY-SA-4.0 +// SPDX-FileCopyrightText: 2026 Jonathan D.A. Jewell += OikosBot CI Runbook +:toc: macro +:icons: font + +Audience: maintainers. This is the operational companion to +link:../ARCHITECTURE.adoc[ARCHITECTURE] (what the system is meant to be) and +link:../DEBT.adoc[DEBT] (what it is not yet): what runs on every change, what +each gate actually enforces, how to diagnose a red check when the logs are not +available, and what has to be true before a release. + +Most push-triggered gates are branch-filtered to `main`, so **a pull request +into `main` is the only way to exercise CI, Hypatia, Governance, CodeQL and the +linters from a working branch**. `Publish Image` runs on pushes to `main`, on +`v*` tags, and on manual dispatch only. + +== Gate inventory + +[cols="2,1,4",options="header"] +|=== +| Workflow | Triggers | What it enforces + +| `CI` +| push/PR → `main` +| Three jobs: `Haskell Analyzer` (`cabal build`/`cabal test` under +`analyzers/code-haskell`), `Rust Workspace` (`cargo fmt --check`, +`cargo test --workspace --locked`, the Eclexia adapter shell test, then +`cargo clippy --workspace`), and `AffineScript` (scaffold check for +`bot-integration-affine`). The Rust job is the gate for every change under +`crates/`. + +| `CodeQL Security Analysis` +| push/PR → `main`, weekly +| Static analysis for Rust and Actions. `codeql.yml` pins +`github/codeql-action@v4.38.1` — the same ref as `ci.yml`, deliberately, so the +lock list stays single-valued. + +| `Governance` +| push/PR → `main`, dispatch +| Delegates to the standards reusable workflow: `actions.lock` verification, +licence/SPDX hygiene, and (informational) baseline validation. A lock mismatch +fails closed here. + +| `Hypatia Security Scan` +| push → `main`/`master`/`develop`, PR → `main`/`master`, weekly, dispatch +| Delegates to `hypatia-scan-reusable.yml` in `hyperpolymath/standards` with +`block-on-high: true`. SARIF findings at `critical` or `high` become +`level: error` and fail the step. Pinned by SHA in `actions.lock`. + +| `Workflow Security Linter` +| any PR, push touching `.github/workflows/**` +| Re-runs the SPDX-header check (line 1 must be the SPDX line) and the +pinned-action check, which probes the API for SHA *existence* — a phantom SHA +is a failure, not a warning. + +| `Secret Scanner` +| any PR, push → `main` +| Delegates to the standards secret scanner. Findings are advisory +annotations, not merge blocks. + +| `Language Policy Enforcement` +| push → `main`, PR +| Blocks new `.py`, `.rb`, `.pl`, `.java`, `.kt` files (repo language policy). +Existing assets are grandfathered; new ones are not. + +| `Publish Image` +| push → `main`, `v*` tags, dispatch +| Builds `Containerfile` and pushes `ghcr.io/hyperpolymath/oikos`. Path-filtered +to `crates/**`, the manifests, the `Containerfile`, `.dockerignore` and its own +file, so a docs-only change does not rebuild the image. + +| `Mirror to Git Forges` +| push → `main`, dispatch +| Pushes the repo to the configured mirrors. `Instant Sync` (fan-out +propagation to the farm) is disabled at the repository-settings level and +documented as such in its own header. + +| `GitHub Pages` +| push → `main`, dispatch +| Builds and deploys the AsciiDoc site. A cold GHC build is the usual cause of +a long first run. + +| `OSSF Scorecard` +| weekly, dispatch +| Supply-chain self-assessment. Schedule-only by design; it does not gate +pushes. + +| `Dependabot Auto-Merge` +| PRs authored by `dependabot[bot]` +| Approves and auto-merges security updates at patch/minor. Gated on the PR +*author*, never on `github.actor`. + +| `Labels` / `Label Triage` +| dispatch / push on `.github/labels.json`; issue events +| Label set synchronisation and triage. Non-gating. +|=== + +== Reading a failure without the logs + +In restricted environments (no artifact download, `gh run view --log-failed` +failing with EOF) the check-run **annotations** API is the reliable channel: + +[source,shell] +---- +# Resolve the run for a commit, then the job, then its annotations +gh api repos/hyperpolymath/oikosbot/actions/runs?head_sha= \ + --jq '.workflow_runs[] | "\(.name) \(.id) \(.conclusion)"' + +# The failing job's id (pick the job whose NAME matches the failing check) +gh api repos/hyperpolymath/oikosbot/actions/runs//jobs \ + --jq '.jobs[] | "\(.name) \(.id) \(.conclusion)"' + +gh api repos/hyperpolymath/oikosbot/check-runs//annotations \ + --jq '.[].message' +---- + +Two gotchas, both learned the hard way: + +* The annotations endpoint takes the **job** id, not the run id (a run id + returns 404). Pick the job whose `name` matches the failing check — a run can + contain several jobs, and `jobs[].id | head -1` is not necessarily the one + that failed. +* Reusable-workflow checks (`Hypatia`, `Governance`, `Secret Scanner`) surface + their diagnostics as a single annotation plus a retained artifact + (`hypatia-scan-findings`, 90-day retention). When the artifact cannot be + downloaded, the annotation text is the whole diagnosis. + +== `actions.lock` + +`.github/workflows/actions.lock` is the estate lockfile GitHub's workflow +enforcement requires. Rules that keep it green: + +* A workflow's `uses:` and its lock entry must agree on ref **and** SHA. The + classic drift is bumping the workflow but not the lock (or vice versa), which + shows up as a `Governance` failure, not a workflow failure. +* Regenerate in the same change as the `uses:` edit + (`scripts/update-actions-lock.sh` upstream), and never prune entries by hand: + several entries exist only to cover actions used *inside* the standards + reusable workflows (checkout, cache, upload-artifact, setup-beam, + scorecard-action, ssh-agent, rust-toolchain, editorconfig-checker). +* Verify locally before pushing: + +[source,shell] +---- +bash tools/ci/lockcheck.sh # actions.lock ↔ workflows consistency +bash tools/ci/linter-verify.sh # runs the linter's checks verbatim +---- + +`tools/ci/lockcheck.sh` re-implements the verifier's consistency rules; it +reported `VALID: 15 workflows, 29 dependencies` on 2026-09-25. + +== Hypatia + +Hypatia is the estate's neurosymbolic workflow scanner. What matters for +diagnosis: + +* Severity maps to SARIF level as `critical`/`high` → `error`. `block-on-high: + true` counts `error`s, so exactly one high-or-critical finding is enough to + fail the step. +* Findings of kind `code_scanning_alerts` are rejected at render time — they + never reach the SARIF, so a "finding" you can see in the Security tab is not + necessarily what failed the build. +* Without `.hypatia-baseline.json`, every current finding counts. A baseline + (schema: `.machine_readable/hypatia-baseline.schema.json` in standards, plus + `scripts/apply-baseline.sh`) suppresses *baselined* findings only, so the + gate stays blocking for anything new. **Never baseline a real vulnerability + to go green** — the baseline is for accepted, evidenced findings. + +Two root causes already diagnosed in this repo, kept here because both are +recurring estate-wide patterns: + +[cols="1,3"] +|=== +| Finding | Cause and fix + +| `research_extensions` RE008 (critical) on + `.github/workflows/dependabot-automerge.yml` +| The rule flags `github.actor == 'dependabot[bot]'` as an identity check + built on a spoofable context: `github.actor` is the dispatch actor, which + fork and rerun events can set without the PR being Dependabot's. The fix is + to gate on the PR author (`github.event.pull_request.user.login`) only. The + rule was wired into the scan on 2026-09-10, which is why a byte-identical + caller went from green to red. + +| `Publish Image` → `Build and push` exit 101 +| A dependency MSRV drift: `tree-sitter` 0.27.0 declares + `rust-version = "1.90"` and `edition = "2024"`, while the `Containerfile` + builder was `rust:1.88-slim`. CI on `ubuntu-latest` passed because the + runner ships a newer rustc; the container did not. The builder now uses the + rolling `rust:1-slim`. Diagnose this class of failure by checking the + *builder* toolchain against the workspace's highest declared `rust-version`, + not by trusting CI's green tick. +|=== + +== Local re-verification + +[source,shell] +---- +just check # rust-build + rust-test + haskell-build + haskell-test +just rust-fmt-check # cargo fmt --check, as CI runs it +bash tools/ci/linter-verify.sh +bash tools/ci/lockcheck.sh +---- + +There is no Rust, Haskell or Elixir toolchain in every environment this repo is +maintained from. Where that is true, Rust changes are verified in CI only: +keep them small and reviewable, say so plainly in the PR, and prefer a PR over +`main` so the branch-filtered gates actually run. + +== Pre-release checklist (GitHub Marketplace) + +The Action is the release unit: `action.yml` plus the container image it pins. + +1. `main` is green across CI, CodeQL, Governance, Hypatia, Workflow Security + Linter, Secret Scanner and Language Policy. +2. `Publish Image` has succeeded on the release commit, and + `ghcr.io/hyperpolymath/oikos@sha256:` resolves. +3. `action.yml`'s pinned `image` digest is updated to that digest (never to a + mutable tag), and the version in the root `Cargo.toml` matches the release. +4. The inputs documented in link:usage.adoc[the end-user guide] match + `action.yml` exactly, including defaults. +5. `CHANGELOG.adoc` has an entry for the release; `DEBT.adoc` reflects what is + still unearned. +6. The Marketplace listing is published from `main` by a repository owner, and + the `DEPLOY.adoc` gate (AffineScript operational parity) is either met or + explicitly still open. diff --git a/docs/usage.adoc b/docs/usage.adoc new file mode 100644 index 0000000..7f74a09 --- /dev/null +++ b/docs/usage.adoc @@ -0,0 +1,394 @@ +// SPDX-License-Identifier: CC-BY-SA-4.0 +// SPDX-FileCopyrightText: 2026 Jonathan D.A. Jewell += Using OikosBot +:toc: macro +:icons: font + +Audience: anyone consuming OikosBot — as a GitHub Action in a workflow, as the +published container, or as the CLI in a pipeline. This page is about *running* +it and *reading what it says*. For what the numbers mean and how far they can +be trusted, start with link:../EXPLAINME.adoc[EXPLAINME] and +link:../DEBT.adoc[DEBT]; the short version is repeated in +<> below. + +== Two ways to run it + +[cols="1,4"] +|=== +| Composite Action | `uses: hyperpolymath/oikosbot@main` (or a pinned SHA). Runs the +published CLI container, so a consumer needs no Rust toolchain. See +`examples/oikosbot-ci.yml` for a complete workflow. + +| Container | `docker run --rm ghcr.io/hyperpolymath/oikos@sha256: `. +Same binary, same flags; useful for local runs and for CI systems that are not +GitHub Actions. +|=== + +The CLI's entrypoint is `oikosbot`, with four subcommands: `report`, `check`, +`compare` and `estate` (plus `self-analyze` for dogfooding). + +== Action inputs + +Every input is optional; the defaults are what `examples/oikosbot-ci.yml` uses. + +[cols="2,1,3,2",options="header"] +|=== +| Input | Default | Meaning | Maps to CLI + +| `mode` +| `report` +| `report` (write SARIF), `check` (threshold gate), `compare` (Pareto verdict +against a base). Anything else is a hard error. +| subcommand + +| `path` +| `.` +| Directory to analyse; in `compare` mode this is the *head* directory. +| positional arg + +| `base` +| _(empty)_ +| Base directory for `compare` mode — e.g. a checkout of the target branch. +Required when `mode: compare`. +| positional arg + +| `format` +| `sarif` +| `sarif` \| `json` \| `text`. +| `--format` + +| `output` +| `results.sarif` +| Output file; empty means stdout. The written path is exposed as the action's +`output-file` output. +| `--output` + +| `config` +| _(empty)_ +| Path to an `.oikos.yml`. Empty means auto-discover one under `path`. +| `--config` + +| `eco-threshold` +| _(empty)_ +| Eco-score floor, 0–100. Empty means "take it from the config, else 50". +| `--eco-threshold` + +| `pr-body` +| _(empty)_ +| File holding the PR description, for the trade-off documentation check in +`compare` mode. +| `--pr-body` + +| `check` +| `false` +| In `compare` mode, fail the run on an undocumented regression or trade-off +_whose drivers are measured or calibrated_. Heuristic findings still cannot +block; see <>. +| `--check` + +| `image` +| pinned digest +| The container image to run. Pinned by digest, never by tag — override it only +to test a candidate build. +| (action-level) +|=== + +[NOTE] +==== +`eco-threshold` is an *eco-score* floor, not an energy or carbon figure. The +config keys `thresholds.eco_minimum.carbon` and `.energy` both express that +same 0–100 floor today; `carbon` wins if both are set. +==== + +== What each mode does + +`report`:: Analyse the directory and emit findings. With `format: sarif` the +output is uploaded to GitHub code scanning by the caller's workflow; `json` is +the same results with the SARIF envelope removed; `text` prints a summary. + +`check`:: Analyse, then compare each finding's eco score against the floor. +Below-floor findings are listed, and the run **fails (exit 1) when the config +says enforcement is blocking** — either `enforcement: blocking` or +`mode: regulator`. With an `advisor`/`consultant` config the run reports and +exits 0, printing which config made it advisory. + +`compare`:: Analyse base and head, project both onto the five Pareto +objectives, and classify the change: + +[cols="2,4"] +|=== +| Verdict | Meaning + +| `Improvement` +| Head is better on at least one objective and worse on none. + +| `Regression` +| Head is worse on at least one objective and better on none. + +| `TradeOff` +| Better on some, worse on others. Must be documented with a +`Pareto-Trade-off:` trailer (or a heading containing "pareto trade-off") in the +PR body passed via `pr-body`. + +| `Neutral` +| No objective moved by more than ε (1e-6). +|=== + +[#enforcement] +=== Enforcement: what can actually fail a run + +`compare --check` fails on an undocumented `Regression` or `TradeOff` **only +when every driving objective is backed by `Measured` or `Calibrated` input**. +That gate exists because a heuristic number must never be able to block a +merge, and it is enforced by the code, not by convention. + +What that means in practice: + +* A function whose detected pattern maps to a known operation category is + priced from the calibration table + (`crates/oikosbot-analysis/src/calibration.rs`) and carries + `Confidence::Calibrated` for the measured rows (`HashLookup`, `Sort`, + `Allocation`, `MathCompute`) — so a regression driven by those findings *can* + block. +* A function with no recognised pattern stays on the naive complexity-derived + estimate and carries `Confidence::Estimated`, which can never block. +* When `--check` is requested but the drivers are heuristic, the run **exits 0 + and prints a `::warning::` naming the verdict and the confidence level**. + A gate that quietly does nothing is indistinguishable from one that passed, + which is the failure mode OikosBot exists to find. + +The estate path (`oikosbot estate`) is the one place with genuinely `Measured` +input: `wall_minutes` comes straight from the GitHub API. + +== Reading a finding + +Every finding is one analysed function (or one synthetic policy/security +record) with: + +* **location** — file, line/column, and the end position used for range + annotations, plus the function name; +* **rule_id** — `oikosbot/` (see the rule table below) or + `oikosbot/general` for a function with no detected pattern; +* **suggestion** — the concrete fix, also printed into the SARIF message; +* **resources** — energy (J), duration (ms), carbon (gCO2e), memory (bytes); +* **health** — eco (0–100), econ (0–100), quality (0–100), and the composite + `overall = 0.4·eco + 0.3·econ + 0.3·quality`; +* **confidence** — the ladder below; +* **resource_range** — when the estimate is calibrated, the min/typical/max + band behind `resources`. `resources` *is* the typical bound. Absent on the + naive path, where no band was computed. + +[cols="1,3"] +|=== +| Confidence | Meaning + +| `Measured` +| Read from an instrument (the GitHub API's wall-clock minutes). The only + level that needs no calibration. + +| `Calibrated` +| Priced from a measured operation profile. Earned per finding, per operation + kind — never assigned to a whole file or run. + +| `Estimated` +| Heuristic: a pattern we recognise but have not measured (host-dependent I/O + and string work), or the naive complexity path. Advises; never blocks. + +| `Unknown` +| No estimate at all (the generic fallback row). +|=== + +[IMPORTANT] +==== +Absolute figures are still small, static and unvalidated against profiling +data. Treat the *relative* signal — which function is worse, which direction a +diff moves — as the product, and the absolute joules as an order-of-magnitude +hint. See link:../DEBT.adoc[DEBT] for the current list of unearned claims. +==== + +=== Rules + +[cols="2,1,4",options="header"] +|=== +| Rule id | Default level | Fires when + +| `oikosbot/nested-loops` +| warning +| Loop nesting depth ≥ 3. + +| `oikosbot/busy-wait` +| warning +| A `loop`/`while` whose body has no sleep/await/yield/recv, no I/O and no + legitimate iteration. + +| `oikosbot/string-concat-in-loop` +| warning +| A `+`/`+=` string concatenation inside a loop body. + +| `oikosbot/clone-in-loop` +| note +| `.clone()` inside a loop body. + +| `oikosbot/unbuffered-io` +| warning +| `File::open`/`File::create` with no `BufReader`/`BufWriter` in scope. + +| `oikosbot/large-allocation` +| note +| An allocation with a numeric literal above 1,000,000. + +| `oikosbot/redundant-allocation` +| note +| Five or more `.to_string()`/`.to_owned()` in one function. + +| `oikosbot/eco-threshold`, `oikosbot/carbon-intensity`, + `oikosbot/security-sustainability` +| warning +| Emitted by the threshold, policy and security-correlation paths rather than + by pattern detection. +|=== + +=== SARIF shape + +`format: sarif` emits SARIF 2.1.0 with one run, driver name `oikosbot`, and the +rules above in `tool.driver.rules`. Per result: + +* `ruleId` is the finding's `rule_id`; `level` comes from the eco score + (`error` below 30, `warning` below 60, else `note`). +* `locations[0].physicalLocation.region` carries start/end line and column, so + GitHub can annotate the whole function. +* `properties` carries the machine-readable payload: `eco_score`, + `econ_score`, `quality_score`, `overall_health`, `energy_joules`, + `carbon_gco2e`, `duration_ms`, `memory_bytes`, `confidence`, optional + `suggestion`, optional `pareto_status` / `pareto_score` / + `pareto_dominated_by`, and — for calibrated estimates — `resource_range` + with `min`/`typical`/`max` per axis. + +`oikosbot/general` records are per-function telemetry, not defects: they appear +in `json`/`text` output and are **filtered out of the SARIF results**, so +analysing a repository does not raise a code-scanning alert for every function +in it. + +== Configuring it (`.oikos.yml`) + +Place an `.oikos.yml` in the analysed directory (or pass `--config`). The keys +that take effect today: + +[source,yaml] +---- +mode: regulator # consultant | advisor (default) | regulator +thresholds: + eco_minimum: + carbon: 50 # the eco-score floor, 0-100 (energy: is the fallback) + enforcement: blocking # fail the run when below the floor +exclude: # globs, matched against paths relative to the root + - "**/target/**" + - "**/node_modules/**" +analysis: + languages: [rust, javascript, python] +---- + +* `mode: regulator` (or `enforcement: blocking`) is what makes a below-floor + finding fail the run. Without either, `check` reports and exits 0, naming the + config that made it advisory. +* `analysis.languages` is intersected with what the tree-sitter analyser can + actually parse (`rust`, `javascript`, `python`). Naming only unsupported + languages is warned about loudly and falls back to the default set, because + analysing nothing would be a silent no-scan. +* Other keys in `config/oikos.yaml` (`eco_standard`, `eco_excellence`, + `complexity`, and the datastore block) are parsed and **discarded** by the + current loader. They are documented for the target design and take no effect; + the file says so inline. + +Policy files under `policies/` (`*.ecl`, Eclexia) are a separate mechanism — +see link:../policies/README.adoc[policies/README]. They are evaluated by +`oikosbot-eclexia` when `--policy-dir` is passed. + +== `BOT_MODE` (the App) vs `mode` (the CLI) + +`BOT_MODE` configures the AffineScript webhook receiver in +`bot-integration-affine/` — it is read in `src/Config.affine`, lower-cased, and +anything unrecognised becomes `advisor`. + +[cols="1,3"] +|=== +| Value | Behaviour (from `src/Types.affine` and `src/Report.affine`) + +| `consultant` +| Answers questions and offers alternatives. PR comments open with + "Oikos consultant — analysis". + +| `advisor` (default) +| Proactive suggestions on pull requests. Comments open with + "Oikos advisor — suggestions". + +| `regulator` +| Enforces policy compliance. Comments open with + "Oikos regulator — policy review". +|=== + +The CLI's `.oikos.yml` `mode:` key is the same three values, but it only +decides enforcement (regulator ⇒ blocking); it does not change the report +wording. `GET /health` on the receiver reports the active mode, which is the +quickest way to confirm what a deployment thinks it is. + +[NOTE] +==== +The GitHub App receiver is a scaffold: the webhook handler and HMAC +verification exist, the HTTP listener is gated on upstream AffineScript stdlib +work, and comments are not yet posted from a live deployment. Until then, CI is +the supported path. +==== + +== Troubleshooting + +[cols="2,3",options="header"] +|=== +| Symptom | Cause and fix + +| `unknown mode '' (report | check | compare)` +| The action's `mode` input is misspelled. The container exits 1 immediately. + +| `compare mode requires the 'base' input` +| `mode: compare` without `base`. Check out the target branch into a directory + and pass it. + +| `no analyzable files under and/or ` +| The directories contain no file with a supported extension, or everything + was excluded. Check `exclude` globs and `analysis.languages`. + +| `Unsupported file extension: ` +| A file was analysed directly (not via a directory) and its extension is not + `rs`/`js`/`py`. + +| Run reports but never fails +| Expected in `advisor`/`consultant` mode. Set `mode: regulator` or + `enforcement: blocking` to make below-floor findings fail. + +| `--check` warns "NOT enforced" and exits 0 +| The verdict is real but its drivers are heuristic (`Estimated`), and only + measured/calibrated inputs may block. This is the gate refusing to fake a + decision, not a bug. + +| Every function is an alert in code scanning +| It should not be: `oikosbot/general` records are filtered out of SARIF. If + you see per-function alerts, the caller is uploading `format: json` output as + SARIF. + +| Eco scores look uniformly high +| The eco score is a log scale anchored at 1 J; calibrated estimates are + microjoule-scale, so most units clamp at 100. The eco *threshold* therefore + discriminates far less than the Pareto verdicts, which compare base against + head. This is a known, tracked limitation, not a silently passing gate. +|=== + +[#honesty] +== What to trust + +* Direction and ranking (which unit is worse, which way a diff moves) are the + product. They are deterministic for a given input tree. +* Absolute joules, grams and milliseconds are static estimates. The calibrated + path is better anchored than the naive one, but neither has been validated + against profiling data on real hardware. +* `Confidence` tells you which of the two you are looking at, per finding. Read + it before acting on a number. diff --git a/policy-engine/deepproblog/eco_problog.pl b/policy-engine/deepproblog/eco_problog.pl deleted file mode 100644 index 09fce9f..0000000 --- a/policy-engine/deepproblog/eco_problog.pl +++ /dev/null @@ -1,208 +0,0 @@ -%% SPDX-License-Identifier: MPL-2.0 -%% SPDX-FileCopyrightText: 2024-2025 hyperpolymath -%% -%% Oikos Bot Policy Engine - DeepProbLog Rules -%% =========================================== -%% Probabilistic logic programming rules that learn from practice. -%% These rules complement the deterministic Datalog rules by handling -%% uncertainty and learning patterns from observed outcomes. -%% -%% Reference: https://github.com/ML-KULeuven/deepproblog - -%% ============================================================================= -%% NEURAL PREDICATES -%% ============================================================================= - -%% Neural network for estimating carbon intensity from code features -%% Input: Code feature vector (complexity, loop depth, allocations, etc.) -%% Output: Probability of high carbon intensity -nn(carbon_estimator, [CodeFeatures], CarbonProb) :: high_carbon_prob(Code, CarbonProb) :- - code_features(Code, CodeFeatures). - -%% Neural network for energy pattern classification -nn(energy_classifier, [CodeFeatures], EnergyClass) :: energy_pattern(Code, EnergyClass) :- - code_features(Code, CodeFeatures). - -%% Neural network for predicting refactoring success -nn(refactor_predictor, [CodeFeatures, RefactorType], SuccessProb) :: - refactor_success_prob(Code, RefactorType, SuccessProb) :- - code_features(Code, CodeFeatures). - -%% Neural network for technical debt estimation -nn(debt_estimator, [CodeFeatures, HistoricalChanges], DebtScore) :: - predicted_debt(Code, DebtScore) :- - code_features(Code, CodeFeatures), - change_history(Code, HistoricalChanges). - -%% ============================================================================= -%% PROBABILISTIC ECO RULES -%% ============================================================================= - -%% Probabilistic eco-friendliness based on learned patterns -%% P(eco_friendly | carbon_prob, energy_pattern) -P :: eco_friendly_prob(Code) :- - high_carbon_prob(Code, CarbonP), - energy_pattern(Code, EnergyClass), - P is (1 - CarbonP) * energy_efficiency_factor(EnergyClass). - -%% Energy efficiency factors (can be learned) -energy_efficiency_factor(efficient) := 0.9. -energy_efficiency_factor(moderate) := 0.6. -energy_efficiency_factor(inefficient) := 0.3. -energy_efficiency_factor(unknown) := 0.5. - -%% ============================================================================= -%% PROBABILISTIC PARETO RULES -%% ============================================================================= - -%% Probability that a change will improve Pareto position -%% Learned from historical refactoring outcomes -P :: pareto_improvement_likely(Code, ChangeType) :- - dominated(Code), - refactor_success_prob(Code, ChangeType, P), - P > 0.6. - -%% Probability of maintaining Pareto optimality after change -P :: maintains_pareto(Code, ChangeType) :- - pareto_optimal(Code), - refactor_success_prob(Code, ChangeType, BaseP), - % Penalty for changes to already optimal code - P is BaseP * 0.8. - -%% ============================================================================= -%% LEARNING FROM PRACTICE (Praxis Loop) -%% ============================================================================= - -%% Evidence from past refactoring outcomes -%% These facts are updated as we observe real outcomes -observed_improvement(code_id_1, carbon_reduction, 0.15). -observed_improvement(code_id_1, energy_reduction, 0.20). -observed_no_improvement(code_id_2, complexity_reduction). - -%% Learn policy effectiveness -%% P(policy_effective | policy, outcomes) -P :: policy_effective(Policy) :- - policy_application(Policy, Code, Outcome), - outcome_positive(Outcome), - policy_success_rate(Policy, P). - -%% Update success rates based on observations (simplified) -policy_success_rate(Policy, Rate) :- - findall(1, (policy_application(Policy, _, positive)), Successes), - findall(1, (policy_application(Policy, _, _)), Total), - length(Successes, S), - length(Total, T), - T > 0, - Rate is S / T. - -%% ============================================================================= -%% RECOMMENDATION CONFIDENCE -%% ============================================================================= - -%% Confidence in refactoring recommendations -%% Combines rule-based reasoning with learned patterns -confidence(recommendation(Code, Action, Reason), Confidence) :- - base_confidence(Reason, BaseConf), - refactor_success_prob(Code, Action, SuccessProb), - Confidence is BaseConf * SuccessProb. - -base_confidence(eco_improvement, 0.8). -base_confidence(debt_reduction, 0.7). -base_confidence(quality_improvement, 0.75). -base_confidence(pareto_optimization, 0.65). - -%% ============================================================================= -%% ADAPTIVE THRESHOLDS -%% ============================================================================= - -%% Thresholds that adapt based on project context and history -%% Start with defaults, adjust based on outcomes - -adaptive_threshold(eco_minimum, carbon, Threshold) :- - project_baseline(carbon, Baseline), - learned_improvement_rate(carbon, Rate), - % Gradually increase expectations - Threshold is max(50, Baseline * (1 + Rate)). - -adaptive_threshold(eco_minimum, energy, Threshold) :- - project_baseline(energy, Baseline), - learned_improvement_rate(energy, Rate), - Threshold is max(50, Baseline * (1 + Rate)). - -%% Default thresholds when no history -adaptive_threshold(eco_minimum, carbon, 50) :- \+ project_baseline(carbon, _). -adaptive_threshold(eco_minimum, energy, 50) :- \+ project_baseline(energy, _). - -%% ============================================================================= -%% KNOWLEDGE GRAPH INTEGRATION (VeriSimDB) -%% ----------------------------------------------------------------------------- -%% Retargeted from the retired ArangoDB + Virtuoso pair to a single VeriSimDB -%% identity-consonance store. The live path is VCL (VeriSim Consonance Language) -%% over VeriSimDB's *semantic* witness (was Virtuoso/RDF/SPARQL) and *graph* -%% witness (was ArangoDB/AQL). The SPARQL/AQL strings below are retained verbatim -%% as legacy bindings / porting reference ONLY — they are facts (data), not rules, -%% so no rule logic changes here. They are re-expressed as VCL at wire-up time -%% (deferred; see ROADMAP.adoc). -%% ============================================================================= - -%% Legacy binding — semantic witness (was Virtuoso/RDF). Port to VCL at wire-up. -sparql_query(eco_best_practices, " - PREFIX eco: - PREFIX sw: - - SELECT ?practice ?description ?impact - WHERE { - ?practice a eco:BestPractice ; - eco:description ?description ; - eco:carbonImpact ?impact . - FILTER (?impact > 0.1) - } - ORDER BY DESC(?impact) -"). - -%% Legacy binding — graph witness (was ArangoDB/AQL). Port to VCL at wire-up. -aql_query(dependency_impact, " - FOR v, e, p IN 1..3 OUTBOUND @startNode GRAPH 'code_dependencies' - LET impact = SUM(p.vertices[*].carbon_score) - RETURN { path: p, total_impact: impact } -"). - -%% ============================================================================= -%% COACHING AND SUGGESTIONS -%% ============================================================================= - -%% Generate coaching suggestions with confidence levels -P :: coaching_suggestion(Code, Suggestion, Priority) :- - needs_attention(Code, Reason), - suggestion_for(Reason, Suggestion), - priority_for(Reason, Priority), - confidence(recommendation(Code, Suggestion, Reason), P), - P > 0.5. - -suggestion_for(high_carbon, "Consider memoization for repeated computations"). -suggestion_for(high_carbon, "Evaluate algorithm complexity - can you reduce from O(n^2)?"). -suggestion_for(energy_inefficient, "Replace busy-waiting with event-driven patterns"). -suggestion_for(energy_inefficient, "Use connection pooling instead of creating new connections"). -suggestion_for(high_debt, "Extract duplicated logic into shared functions"). -suggestion_for(low_coverage, "Add tests for edge cases in critical paths"). - -priority_for(high_carbon, high) :- !. -priority_for(energy_inefficient, high) :- !. -priority_for(high_debt, medium) :- !. -priority_for(low_coverage, medium) :- !. -priority_for(_, low). - -%% ============================================================================= -%% TRAINING DATA GENERATION -%% ============================================================================= - -%% Generate training examples for neural networks -%% Format: input features -> expected output -training_example(carbon_estimator, Features, Label) :- - historical_analysis(Code, Features, carbon_score, ActualScore), - (ActualScore < 40 -> Label = high ; Label = normal). - -training_example(refactor_predictor, [Features, Action], Label) :- - historical_refactor(Code, Action, Outcome), - code_features(Code, Features), - (Outcome = success -> Label = 1.0 ; Label = 0.0).