Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .agents/skills/augustus/SKILL.md

Large diffs are not rendered by default.

2 changes: 1 addition & 1 deletion .claude-plugin/marketplace.json
Original file line number Diff line number Diff line change
Expand Up @@ -7,7 +7,7 @@
"description": "Design judgment for placing typed probabilistic judgment (Jev-class System One) using math, logic, and algorithmic mental models — across AI, SWE, business, knowledge work, and life. Jev is the exemplar, not the monopoly (Laya, kev, OpenJev, TypeAR, GLiNER, encoder ZS). Ranking ≠ calibration; soft Noul ≠ hard gate. Formal methods are one pillar. Never launder a Noul as a proof. Not a TypeSafe product — load typesafe-ai for Jev contracts. Named for Augustus De Morgan, mentor of W. S. Jevons.",
"name": "augustus",
"source": "./.agents/skills/augustus",
"version": "0.4.0"
"version": "0.5.0"
}
]
}
132 changes: 71 additions & 61 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -16,74 +16,84 @@ folds: `research/notes.md`.

## [Unreleased]

Merged after v0.4.0: 0843 (§114 / 289–302 / #97), NanoJev (§115 /
303–308 / #98), jcr (§116 / 309–316 / #99), SemIf (§117 / 330–336 /
#100), llm-to-jev (§118 / 322–329 / #101), plus this hygiene pass.
Does **not** bump the 0.4.0 pin. Uniqueness dumps live in
## [0.5.0] - 2026-09-20

Pegged against [`typesafe-ai/skills` v0.5.7](https://github.com/typesafe-ai/skills/tree/v0.5.7)
(`65a39f3`, 2026-09-12). Live HEAD of that repo is still this commit,
the only tagged official-skill revision.

Nine commits on `main` after the v0.4.0 tag (merged #33 README map,
#31 Harbor-jevals, #34 Pages layout, #35 hysteresis/ECE, #36 NanoJev,
#38 jcr, #37 SemIf, #40 llm-to-jev, #39 0843 hygiene). Open #41
(0947 fold) is in flight on another branch and is not part of this
release. Verbose hourly locks stay in
[`research/changelog-hourly.md`](research/changelog-hourly.md).
Do not reopen or amend PR #23–#40. Do not push onto open #41.

### Added

- **User-provided 0940 HIGH (`notes.md` §118).** Migration /
question-design on-ramp: [alexwestco/llm-to-jev](https://github.com/alexwestco/llm-to-jev)
turns decision-shaped LLM prompts into proposed Choice/Score/Noul.
Conversion assistant, not an automatic guarantee of equivalent
behavior. Deterministic heuristics, not an LLM. Partial
convertibility; prose stays with the LLM. Heuristic conversion ≠
calibrated Noul. Companion to the decision-design card +
validation gate. Composition items 322–329 / batch #101.
- **SemIf densify (PRIMARY, `notes.md` §117).** [TheoLeeCJ/SemIf](https://github.com/TheoLeeCJ/SemIf)
MIT; homepage openjev.com; default **master**; live REST
**2282★** / **140** forks; HEAD `ca3ba65f1429`. SemIf was
formerly OpenJev; independent; not affiliated with Jev or
TypeSafe. Direct option logits; 0 output tokens; MLX backend
(`--backend mlx`). Speed *theirs* Qwen3.5-4B 3090: direct
1.023s vs AR JSON 5.332s (**5.21×**); argmax agree 18/21
(systems comparison ≠ semantic equivalence). Browser ladder
authored BA 0.813, pert 0.766, TypeSafe subset 0.845 vs
Published Jev 0.883 *theirs*. Softmax over options ≠ calibrated
Noul. Rename is densify not a second census. JevBench 74.6 is
§78 not this ladder. Composition items 330–336 / batch #100.
- **NiazMorshed2007/jcr (`notes.md` §116).** Capability-tree
lookup returns context, **does not execute**. skills vs
capabilities. 0.6 band is application policy. routing ≠
permission. docs ≠ authority to run. sol-vs-opus5-20
lookup+explain only; n=1; wall-time mixed; Not Harbor
task-execution. Live REST **4★**; HEAD `138b3832`; size
**14850**. Composition items 309–316 / batch #99.
- **NanoJev unified-games-v1 densify (`notes.md` §115).**
TianyuCodings/NanoJev (Python MIT; **1289★** / **158** forks;
HEAD `618cea6d`). Quote *theirs*: A 0.6B parallel decision
model; zero output-token decoding; one model, four games.
Game success ≠ calibrated Noul. not TypeSafe Jev. Composition
items 303–308 / batch #98.
- **Hourly 0843 HIGH (`notes.md` §114).** Measurement / judgment
fold. PRIMARY: pretrained Qwen2.5 base ECE already low;
instruct-tuning wrecks honesty (acc flat, mean conf
74.1%→96.7%). Equal-width ECE ≠ quantile ECE (0.113 vs 0.076
*theirs*). Calibration does not compose; hop-ECE is
permutation-invariant (soundness theater as a trajectory
audit); Deferred Crispification; 25–60× headline withdrawn.
Hysteresis `{enter:0.8, exit:0.6}` is policy attached to a
probability. Ranking ≠ calibration. g0runmezadam/what-is-jev
**is** tunahansahin897/what-is-jev. Qwen2.5 / Qwen 3.8 /
Qwen/Qwen3.8-27B ≠ Archer. Archer still promised_not_landed.
Composition items 289–302 / batch #97.
- **Hourly 0743 HIGH (`notes.md` §113).** Harbor-jevals PRIMARY
(ywchiu/jev_benchmark). Verdict-open-jev linear ECE floor ≠
TypeSafe replica. DecisionOps ACT / REVIEW / FALLBACK
(provider failure is **not** a policy outcome). Lock dump:
[`research/changelog-hourly.md`](research/changelog-hourly.md).
- **Pages / onboarding.** Custom site layout (nav, comparison, install,
pillars). With vs without Augustus on the homepage and README
(`docs/assets/with-without-augustus.svg`): call-then-act vs
place-then-judge. Jev is the exemplar, not the monopoly.

- **README.** Scannable skill map (one line per file). The living catalog
stays in the reference cards and `research/notes.md`, not a README wall.

- **Recipes (class, not Jev-only).** Same with/without split for any
Choice/Score/Noul-style or typed probabilistic judgment tool. Full
cards: [`docs/release-notes-v0.5.0.md`](docs/release-notes-v0.5.0.md).
Shape, not invented scores:
- **Encoder (GLiNER / GLiClass):** without — swap locate/categorize for
a decision head and hard-gate spans. With — species map; remainder
after extractive spans. Measure span quality separately from ECE.
- **Open heads (Laya, SemIf, kev, Jeff-1):** without — treat wire-compat
or argmax agree as a replica. With — softmax over options ≠ calibrated
Noul; systems timing ≠ semantic equivalence. Measure ECE/Brier on
held-out, not only speed or top-1.
- **NanoJev:** without — game wins as calibration. With — specialist
gameplay S1; local boolean ≠ TypeSafe noul. Measure held-out game
success separately from ECE.
- **llm-to-jev:** without — ship converted prompts as equivalent
behavior. With — heuristic on-ramp; review Score rubric; prose stays
with the LLM. heuristic conversion ≠ calibrated Noul.
- **jcr:** without — run what the capability tree found. With — lookup
returns context and **does not execute**. Routing ≠ permission;
docs ≠ authority to run.
- **localjev / prompted JSON:** without — parse generated JSON as a
Noul. With — schema-valid ≠ picked-right; prompted JSON ≠ structured
logit read.

- **Class / migration.** llm-to-jev conversion on-ramp (`notes.md` §118).
SemIf rename + MLX densify (`§117`; formerly OpenJev, independent).
NanoJev unified-games densify (`§115`). jcr capability resolver
(`§116`; docs ≠ execute).

- **Measurement honesty.** 0843: hysteresis `{enter, exit}` is policy
attached to a probability, not a model property; instruct-tuning can
wreck ECE while accuracy stays flat; equal-width ECE ≠ quantile ECE;
hop-ECE is permutation-invariant (trajectory soundness theater);
ranking ≠ calibration. 0743: Harbor-jevals practice (schema-pass ≠
joint fields; skip-and-call-a-tool); Verdict linear ECE floor ≠
TypeSafe replica; DecisionOps ACT / REVIEW / FALLBACK (a provider
failure is **not** a policy outcome). Evaluator reports both ECEs,
AUC, accuracy@0.5, cost-optimal threshold, hysteresis, and hop-ECE
invariance.

### Changed

- `evaluate_decisions.py` hop-ECE self-test covers reverse **and**
even/odd interleave (permutation invariance is not reverse-only).
- `uniqueness_gate.py` checks Pages strings (`LICENSE` lives in
the #34 layout) and refuses CHANGELOG/README dump walls. The
0843 uniqueness lock includes merged #34 and #35.
- `docs/ecosystem.md` 0843 blurb cites `notes.md` §114.
- Marketplace plugin version and SKILL YAML pin: 0.4.0 → 0.5.0.
- Homepage / README / CITATION.cff / SECURITY supported line follow 0.5.0.
- CHANGELOG Unreleased dump folded into this cut. Verbose hourly locks
remain in `research/changelog-hourly.md`.
- Evaluator hop-ECE self-test covers reverse and even/odd interleave
(permutation invariance is not reverse-only). uniqueness_gate checks
Pages strings (`LICENSE` in the #34 layout) and refuses CHANGELOG/README
dump walls. `docs/ecosystem.md` 0843 blurb cites `notes.md` §114.

### Security

- `SECURITY.md`: supported line is 0.5.x. Report via GitHub Security
Advisories. No invented Scorecard number.

## [0.4.0] - 2026-09-20

Expand Down
2 changes: 1 addition & 1 deletion CITATION.cff
Original file line number Diff line number Diff line change
Expand Up @@ -24,5 +24,5 @@ keywords:
- formal-methods
- agentic-ai
license: MIT
version: 0.4.0
version: 0.5.0
date-released: "2026-09-20"
2 changes: 1 addition & 1 deletion CONTRIBUTING.md
Original file line number Diff line number Diff line change
Expand Up @@ -45,7 +45,7 @@ re-opened as "new." Before folding:
- Do not re-fold an already-landed section as a new beat
- Do not reopen or amend a merged fold PR (#23–#40)
- Do not push onto an in-flight fold PR (open #41)
- Pre-0.4.0 uniqueness dump: `research/changelog-hourly.md` (archive,
- Hourly uniqueness dump: `research/changelog-hourly.md` (archive,
not release notes)

## Secrets
Expand Down
19 changes: 17 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -34,7 +34,22 @@ never launder a Noul as a proof.

![Jev-class models with vs without Augustus. Without: call the model, act on the score, then quiet failure modes (soft Noul treated as hard gate, GPT bakeoff framing, no falsifier, polarity unchosen). With Augustus: state, pillar and family map, question design, fail-open vs fail-closed, typed Choice Score Noul, code owns effects, named falsifying experiment.](docs/assets/with-without-augustus.svg)

Jev-class: Jev, kev, Laya, OpenJev. Call and act, or place the judgment.
Jev-class: Jev, kev, Laya, OpenJev, GLiNER, SemIf, NanoJev, Jeff-1, localjev.
Call and act, or place the judgment. Same split for any typed probabilistic
judgment tool, not Jev-only.

## Recipes

Class-wide, not a TypeSafe how-to. Full cards:
[`docs/release-notes-v0.5.0.md`](docs/release-notes-v0.5.0.md) ·
[Pages recipes](https://24601.github.io/Augustus/#recipes).

- **Encoder (GLiNER / GLiClass).** Without: treat locate/categorize as a decision head and hard-gate spans. With: species map; remainder after extractive spans. Measure span quality separately from ECE.
- **Open heads (Laya, SemIf, kev, Jeff-1).** Without: wire-compat or argmax agree as replica. With: softmax ≠ calibrated Noul; systems timing ≠ semantic equivalence. Measure ECE/Brier on held-out, not only speed.
- **NanoJev.** Without: game wins as calibration. With: specialist gameplay S1; local boolean ≠ TypeSafe noul. Measure held-out game separately from ECE.
- **llm-to-jev.** Without: ship converted prompts as equivalent behavior. With: heuristic on-ramp; review the Score rubric. heuristic conversion ≠ calibrated Noul.
- **jcr.** Without: run what the tree found. With: lookup returns context; **does not execute**. Routing ≠ permission; docs ≠ authority to run.
- **localjev / prompted JSON.** Without: parse generated JSON as a Noul. With: schema-valid ≠ picked-right.

## The skill

Expand Down Expand Up @@ -113,7 +128,7 @@ GPT's instructions or a Project's knowledge and it will follow the protocol.
## Versioning

See [CHANGELOG.md](CHANGELOG.md) and
[releases](https://github.com/24601/Augustus/releases). Current: **0.4.0**,
[releases](https://github.com/24601/Augustus/releases). Current: **0.5.0**,
written against [`typesafe-ai/skills` v0.5.7](https://github.com/typesafe-ai/skills/tree/v0.5.7)
(`65a39f3`; live HEAD still this commit). Re-read live TypeSafe docs
before treating that pin as current API behavior.
Expand Down
6 changes: 3 additions & 3 deletions SECURITY.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,9 +4,9 @@

| Version | Supported |
| ------- | --------- |
| 0.4.x | Yes |
| 0.3.x | No |
| < 0.3 | No |
| 0.5.x | Yes |
| 0.4.x | No |
| < 0.4 | No |

This repo is a design-judgment skill plus offline scripts. There is no
hosted API and no runtime that accepts untrusted input by default.
Expand Down
14 changes: 11 additions & 3 deletions docs/_includes/comparison.html
Original file line number Diff line number Diff line change
Expand Up @@ -4,13 +4,18 @@
<h2 id="compare-title">With vs without Augustus</h2>
<p>
Call the model and act on the score, or place the judgment. Jev is the
exemplar. kev, Laya, and OpenJev sit in the same class. Not TypeSafe-only.
exemplar. Encoders, open heads, gameplay specialists, and conversion
on-ramps sit in the same class. Not TypeSafe-only.
</p>
<ul class="class-chips" aria-label="Models in this class">
<li>Jev</li>
<li>kev</li>
<li>Laya</li>
<li>OpenJev</li>
<li>GLiNER</li>
<li>SemIf</li>
<li>NanoJev</li>
<li>Jeff-1</li>
</ul>
</div>
<div class="compare-grid">
Expand Down Expand Up @@ -82,7 +87,10 @@ <h3>Place, then judge</h3>
</article>
</div>
<p class="compare-caption">
Shared class: Jev, kev, Laya, OpenJev. Augustus owns placement. The model
owns narrow judgment. Code owns the effect.
Shared class: Jev, kev, Laya, OpenJev, GLiNER, SemIf, NanoJev, Jeff-1,
localjev. Augustus owns placement. The model owns narrow judgment. Code
owns the effect.
<a href="#recipes">Class recipes</a>
(problem → without → with → measure).
</p>
</section>
104 changes: 104 additions & 0 deletions docs/_includes/recipes.html
Original file line number Diff line number Diff line change
@@ -0,0 +1,104 @@
<section class="section recipes" id="recipes" aria-labelledby="recipes-title">
<div class="section-head">
<p class="kicker">Recipes</p>
<h2 id="recipes-title">Same split, whole class</h2>
<p>
Not Jev-only. Problem → without → with → what to measure. Full cards
in the
<a href="{{ '/release-notes-v0.5.0.html' | relative_url }}">v0.5.0 notes</a>.
No invented scores.
</p>
</div>
<div class="recipe-grid">
<article class="recipe panel">
<p class="kicker">Encoder</p>
<h3>GLiNER / GLiClass</h3>
<dl>
<dt>Problem</dt>
<dd>Locate or categorize treated as a decision head.</dd>
<dt>Without</dt>
<dd>Swap, hard-gate spans, bake off a chat LLM.</dd>
<dt>With</dt>
<dd>Species map. Extractive remainder. Soft scores ≠ hard gates.</dd>
<dt>Measure</dt>
<dd class="recipe-measure">Span quality separately from ECE. Softmax ≠ Noul.</dd>
</dl>
</article>
<article class="recipe panel">
<p class="kicker">Open heads</p>
<h3>Laya, SemIf, kev, Jeff-1</h3>
<dl>
<dt>Problem</dt>
<dd>Wire-compat or argmax agree treated as a replica.</dd>
<dt>Without</dt>
<dd>Drop-in swap. Ship the speedup. Skip OOD.</dd>
<dt>With</dt>
<dd>Softmax ≠ calibrated Noul. Systems timing ≠ semantic equivalence.</dd>
<dt>Measure</dt>
<dd class="recipe-measure">Held-out ECE/Brier and accuracy. In-distribution vs OOD.</dd>
</dl>
</article>
<article class="recipe panel">
<p class="kicker">Gameplay</p>
<h3>NanoJev</h3>
<dl>
<dt>Problem</dt>
<dd>Game success treated as a calibrated Noul.</dd>
<dt>Without</dt>
<dd>Quote a win rate as a production gate.</dd>
<dt>With</dt>
<dd>Specialist S1. Local boolean ≠ TypeSafe noul.</dd>
<dt>Measure</dt>
<dd class="recipe-measure">Held-out game metrics on one ledger. ECE on another.</dd>
</dl>
</article>
<article class="recipe panel">
<p class="kicker">Conversion</p>
<h3>llm-to-jev</h3>
<dl>
<dt>Problem</dt>
<dd>A chat prompt assumed equivalent to Choice/Score/Noul.</dd>
<dt>Without</dt>
<dd>Paste, convert, ship.</dd>
<dt>With</dt>
<dd>Heuristic on-ramp. Review the Score rubric. Prose stays with the LLM. heuristic conversion ≠ calibrated Noul.</dd>
<dt>Measure</dt>
<dd class="recipe-measure">Suitability labels and human review. Not equivalent behavior.</dd>
</dl>
</article>
<article class="recipe panel">
<p class="kicker">Lookup</p>
<h3>jcr</h3>
<dl>
<dt>Problem</dt>
<dd>Finding a documented command treated as permission to run it.</dd>
<dt>Without</dt>
<dd>The agent executes whatever the tree returned.</dd>
<dt>With</dt>
<dd>Returns context. <strong>Does not execute.</strong> Routing ≠ permission.</dd>
<dt>Measure</dt>
<dd class="recipe-measure">Lookup+explain only. Not Harbor task-execution.</dd>
</dl>
</article>
<article class="recipe panel">
<p class="kicker">Readout</p>
<h3>localjev / prompted JSON</h3>
<dl>
<dt>Problem</dt>
<dd>Parsed JSON from a generator treated as a Noul.</dd>
<dt>Without</dt>
<dd>Schema-valid taken as picked-right.</dd>
<dt>With</dt>
<dd>Prompted JSON ≠ structured logit read. Schema-valid ≠ picked-right.</dd>
<dt>Measure</dt>
<dd class="recipe-measure">Schema pass separately from calibration and from the effect.</dd>
</dl>
</article>
</div>
<p class="compare-caption">
Measurement recipe (hysteresis, equal-width vs quantile ECE, hop-ECE,
Harbor schema-pass ≠ joint, DecisionOps FALLBACK):
<a href="{{ '/release-notes-v0.5.0.html' | relative_url }}">v0.5.0 notes</a>.
Ranking ≠ calibration. Soft Noul ≠ hard gate.
</p>
</section>
2 changes: 1 addition & 1 deletion docs/_layouts/default.html
Original file line number Diff line number Diff line change
Expand Up @@ -36,7 +36,7 @@
<a href="{{ '/ecosystem.html' | relative_url }}"{% if page.url contains 'ecosystem' %} aria-current="page"{% endif %}>Ecosystem</a>
</li>
<li>
<a href="{{ '/release-notes-v0.4.0.html' | relative_url }}"{% if page.url contains 'release-notes' %} aria-current="page"{% endif %}>v0.4.0</a>
<a href="{{ '/release-notes-v0.5.0.html' | relative_url }}"{% if page.url contains 'release-notes' %} aria-current="page"{% endif %}>v0.5.0</a>
</li>
<li>
<a href="https://github.com/24601/Augustus" rel="noopener noreferrer">GitHub</a>
Expand Down
Loading
Loading