Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
11 changes: 8 additions & 3 deletions CONTRIBUTING.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,7 @@ Gloss improves when a benchmark claim becomes easier to inspect, reproduce, or c
- Add a PowerPoint construct the corpus does not cover.
- Reproduce a grader false positive or false negative.
- Improve a candidate assertion and its mutation fixture.
- Improve a deterministic comparative generation path and prove the score change.
- Review a prompt requirement, evidence record, or release gate.
- Make the documentation easier to run from a clean checkout.

Expand All @@ -29,7 +30,7 @@ Install [uv](https://docs.astral.sh/uv/), Python 3.12, LibreOffice Impress, and
git clone https://github.com/aronchick/gloss.git
cd gloss/acidslide-v1/grader
uv sync --extra dev --locked
uv run python ../benchmark/validate_corpus.py
uv run ../benchmark/validate_corpus.py
uv run pytest tests/test_mutation_fixtures.py -q
```

Expand Down Expand Up @@ -84,7 +85,10 @@ CI must pass. Maintainers may ask for a smaller fixture, clearer provenance, or

## Launch surface

The static site lives in `site/`. Its public counts come from `site/evidence/preview-v1.json`, which is verified against the generated mutation fixture reports by:
The static site lives in `site/`. Harness counts come from
`site/evidence/preview-v1.json`; comparative bars come from the byte-equivalent
public copy of `acidslide-v1/benchmark/comparative-v1/summary.json`. Verify both
sources, local assets, copy, and media metadata with:

```bash
node launch/verify-launch.mjs
Expand All @@ -96,7 +100,8 @@ Rebuild the silent 21-second launch film with:
node launch/render-video.mjs
```

The renderer requires `rsvg-convert` and `ffmpeg`. Do not hand-edit a published number in the site or video without updating the evidence bundle and its verification source.
The renderer requires `rsvg-convert` and `ffmpeg`. It reads the comparative
summary directly. Do not hand-edit a published number in the site or video.

## Conduct and security

Expand Down
18 changes: 18 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -23,6 +23,24 @@ Every launch count is recorded in [`site/evidence/preview-v1.json`](site/evidenc
node launch/verify-launch.mjs
```

## Frozen comparative bundle

[`acidslide-v1/benchmark/comparative-v1`](acidslide-v1/benchmark/comparative-v1)
contains four repository-owned generation paths, three public seeds per path,
and 12 editable 20-slide decks. The canonical Linux/amd64 grader completed all
240 slide renders.

The current local artifact scores are 67.68% for the native paths and 62.32%
for the visual paths. These are reproducible workflow baselines, not model
rankings. Every report says `local artifact score; self-reported` and carries no
model attribution.

Reproduce every deck, report, hash, and public bar:

```bash
./acidslide-v1/benchmark/comparative-v1/reproduce.sh
```

Release mode intentionally fails until independent prompt convergence, assertion provenance and evidence, reviewer approvals, baselines, environment manifests, and signed release indexes are complete.

## Score provenance
Expand Down
4 changes: 4 additions & 0 deletions acidslide-v1/.dockerignore
Original file line number Diff line number Diff line change
Expand Up @@ -22,3 +22,7 @@ benchmark/release-index.json
benchmark/release-index-chain.json
benchmark/baselines/*.json
benchmark/control-handoffs/*.json

# Comparative decks and local self-reported grading outputs are mounted inputs,
# never part of the renderer image they identify.
benchmark/comparative-v1/
52 changes: 52 additions & 0 deletions acidslide-v1/benchmark/comparative-v1/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,52 @@
# Gloss comparative v1

This frozen bundle compares four deterministic, repository-owned PowerPoint
generation paths:

- `native-precise`
- `native-fast`
- `visual-precise`
- `visual-fast`

Each path has three seeded runs (`1103`, `2207`, `3301`). These are reproducible
workflow baselines, not model rankings. Every result is labeled:

> local artifact score; self-reported

Every generation record is labeled:

> repository-owned path; no model attribution

## Reproduce every bar

From the repository root:

```bash
./acidslide-v1/benchmark/comparative-v1/reproduce.sh
```

The command regenerates all 12 editable 20-slide decks, builds the pinned
Linux/amd64 grader image, grades all 240 slides, freezes artifact hashes, and
recomputes every published metric.

## Published artifacts

- `cohort.json` binds the local scoring cohort to the exact grader source,
container, prompts, checklist, schemas, fonts, and assets.
- `manifest.json` freezes the hashes and metrics for all 12 runs.
- `summary.json` is the only source used by the public chart and launch video.
- `runs/<path>/run-<n>/` contains `deck.pptx`, `generation.json`,
`artifact-context.json`, `report.json`, and `artifact-sha256.json`.
- `verify_bundle.py` independently checks archive structure, all hashes,
disclosure labels, report completion, cohort identity, and summary math.

Metrics:

- **Local fidelity** — AcidSlide’s severity-weighted artifact score.
- **Visual SSIM** — mean full-slide similarity to the gold renders.
- **Native pass** — severity-weighted checklist pass rate with every
`visual_ssim` check excluded.

Local reports are intentionally ineligible for an official leaderboard. They do
not assert that a hosted service verified the score or that any model generated
the artifact.
33 changes: 33 additions & 0 deletions acidslide-v1/benchmark/comparative-v1/cohort.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,33 @@
{
"disclosure": {
"attribution": "repository-owned path; no model attribution",
"comparison_scope": "reproducible workflow baselines, not model rankings",
"verification_label": "local artifact score; self-reported"
},
"grader_revision": "unavailable",
"provenance": {
"environment_attestation_sha256": "sha256:62695d02ec683aa3680f2db72ff6c17fcce14fccfb0b5367c75704d92db04602",
"grader_source_tree_sha256": "sha256:685ce0505346cf3e844d49ab25c51e906ec917d2bf6b6ea6bdb95e65d4e23d28",
"scoring_cohort_id": "sha256:13b3015b9e346d8eb4cd7f26fd13f3c551941a8b64a5bbd7e483063567bfff1f",
"scoring_manifest_sha256": "sha256:97a4df0aa45d198e0a1494d6c2501d8c9fb375c617aff1012588d843da2b20b8"
},
"scoring_manifest": {
"bundle_hashes": {
"asset_manifest_sha256": "sha256:813a36b7e34fec686863e0d94ec4c040344b1c1a27270cdb103e424f988f4e8c",
"checklist_bundle_sha256": "sha256:bcfe2c7c3647742eeb90edbef74c1de09651c05b4cdbc1b5cc4430b67f5e9d83",
"font_manifest_sha256": "sha256:a770fe1f2a127a532da82dd8f9063c43dcc51523b91575c676db4fac80bc7c55",
"grader_source_tree_sha256": "sha256:685ce0505346cf3e844d49ab25c51e906ec917d2bf6b6ea6bdb95e65d4e23d28",
"prompt_bundle_sha256": "sha256:1579cd0439040042b091e276b3d67d0fbb9dfbf7087bea80651fd40835c7b818",
"schema_bundle_sha256": "sha256:f9a55826487a67023137ef2319e9a99724469e38cf9ff5e78932801c4d9616a5",
"scored_assertion_inventory_sha256": "sha256:62be3978e59e177c819b6900964f0d543a4e39fec0b89fd5c45404ccf214fc47"
},
"disclosure": {
"attribution": "repository-owned path; no model attribution",
"comparison_scope": "reproducible workflow baselines, not model rankings",
"verification_label": "local artifact score; self-reported"
},
"environment_attestation_sha256": "sha256:62695d02ec683aa3680f2db72ff6c17fcce14fccfb0b5367c75704d92db04602",
"image_digest": "sha256:29f5fe733d762d2633f28824a18807c3273d02c040db56c0cda12affd246732d",
"schema_version": "1.0"
}
}
Loading
Loading