AIQ records fixed-fixture AI and agent benchmark results. The repository contains a Rust runner, a Rust verifier, a Next.js application, the public AIQ Core catalog, and one declarative PostgreSQL schema.
AIQ production is live at aiq.wiki. The personal Vercel
scope acgbox hosts project aiq. The personal Supabase organization ACG Box
hosts project aiq on PostgreSQL 17.6 with reference
xxnszykaeapolqdnhalx. The personal Cloudflare account that owns the
aiq.wiki zone owns DNS handoff. Production uses the private Storage buckets
aiq-submission-packages and aiq-runner-artifacts.
The only supported production tuple is AIQ Core 1.1.0, task scorer 1.0.6, aggregate
scoring 1.0.8, and measurement 2.0.0. Do not publish, preserve online,
migrate, or display a legacy tuple as production evidence. Production must
remain without an Official AIQ 2.0 publication until a fresh complete,
non-synthetic, signed 17-by-72 calibration is replayed under policy v2 to
establish the fixed item bank and admission v3, and
a separate fresh 17-by-72 Official package passes native verifier replay and
all release gates.
Formal calibration and Official model and evaluator work has no benchmark-enforced wall-time, step, tool-call, aggregate-evaluator, or per-check deadline. The runner separately measures model and evaluator elapsed time, agent steps, tool calls by type, tokens, and estimated cost. These values are auxiliary evidence only and cannot change task score, AIQ, quality, strict pass, interval, eligibility, or ranking. Functional preflight and hard safety boundaries remain separate. A safety, runtime, provider, or infrastructure termination produces a null semantic score, never a semantic zero.
All earlier bounded or deadline-bearing runs remain immutable failed release evidence. They cannot be relabeled, composed with selected reruns, or published under the active tuple.
- Repository source targets AIQ Core
1.1.0, with 72 private controlled tasks in ten domains. Task evaluation stays at1.0.6; aggregate scoring is1.0.8. - Every formal task encodes
wall_seconds: null,max_steps: null, andmax_tool_calls: null. Controlled evaluator configuration usesaiq.evaluator-config.v2withcompletion_policy: natural_completionand no aggregate or per-check deadline. - The candidate.20 catalog is the deterministic source candidate for the
1.1.0task set. A fresh independent review and seal, complete 17-by-72 calibration, policy-v2 fixed-bank admission, separate complete Official run, publication, and deployment are required. No earlier publication is a fallback. - The public catalog contains metadata and commitments, not private task content.
- Task scores use committed weighted binary checks. A failed hard gate or structural check sets the score to zero; otherwise the evaluator divides passed positive weight by total positive weight. The runner commits only one semantic evaluator result for each sealed response and workspace. A retryable evaluator process failure keeps that evidence pending and reruns only the evaluator on resume. The independent verifier executes the evaluator once and compares the parsed result and exact raw output digest. A failed verifier invocation releases the claim for a later model-free replay. A first successful replay with different output also requires a later confirmation attempt. Publication remains blocked until one exact replay matches.
- The source-head AIQ measurement contract is
2.0.0: the Official ranking score is100 × logistic(theta)from the admitted fixed Rasch item bank; theta and its conditional Wald interval are reported separately from the raw equal-domainqualityScorediagnostic. This contract is not an IQ norm or a 150-point scale. - Calibration policy
aiq.official-calibration-policy.v2reports the binary informative-task rate and its 0.50 descriptive target, but does not use that count as a release cliff. Complete semantic coverage, non-uniformity, universal floor and ceiling limits, domain checks, and model and latent spread remain hard gates. - Strict pass is strict successes divided by all attributable tasks with a
valid semantic task score. Partial scores remain in that denominator; only
missing, infrastructure-invalid, runtime-failed, and unscored tasks are
excluded.
invalid_tasksrecords observed runtime or infrastructure failures, whilemissing_tasksis reserved for an expected cell with no result record. Runtime failures are not semantic zeros. The Wilson interval uses the same sample. - The model matrix contains 17 configurations: six Sol, six Terra, and five Luna.
- The runner performs capability preflight, executes tasks, scores results, and
creates signed
aiq.result-package.v4envelopes. - Every result keeps separate runner-observed model and evaluator elapsed time
as
latency.wall_msandlatency.evaluator_msand, when Codex reports it, token usage and a versioned Standard API-equivalent cost estimate. - AIQ, Rasch ability, quality, strict pass, ranking, and intervals use only evaluator-backed semantic task scores. Elapsed time, tokens, tool use, and estimated cost are independent efficiency evidence and never change a score.
- Public evidence labels time as
runner_observed, provider token source asprovider_reported, and verifier-checked token and cost evidence asverifier_recomputed. Unavailable evidence remains null, not zero. - The verifier reconstructs submitted workspaces and replays deterministic
evaluators before it signs
aiq.verifier-attestation.v4evidence. - The verifier also provides an offline
diagnose-rescoreaudit. It first verifies and replays one source package, then scores the preserved cells with a candidate source, task, evaluator, runtime, and toolchain set. Its create-new report is permanently non-Official and non-ranking. It cannot publish or create an attestation. - Production uses three distinct identities: runner, verifier, and publisher.
- The Web application reads public database views and sends controlled writes through server routes.
The source-head ordered task-metadata catalog digest is:
sha256:3580555315d49a62b28b6947491819276dca5b261ade802f10b33808569d1708
Its public source-release digest is
sha256:5b651845280ea0b27a8cfc2aec4efe4c149bbbfdeeb0e5a9c883174938b58d69.
The production task-set identity is aiq-core/1.1.0; the retained catalog source
identity is aiq-core/1.1.0-candidate.20. Do not infer any controlled
identity from these public digests. The reviewed evaluator identity is
sha256:748e0a6c07eb7e3407cc22d50b65eb6d055305cb6e1d719ca3cfd3a109bec809.
The current no-deadline public-safe database task-set identity is
sha256:c7481e46c64dbf5ff9f50a85c83608d48390a03cbf9e94a1d89ab36aeb6df89a,
and its task-commitment manifest identity is
sha256:d8dddd1bc496a1609c3268068fdfdfa4562c589ddfdfec365a6a49caadefe96b.
These are checked-in bindings derived from the reviewed candidate.15 seal.
Production activation still requires the exact private corpus, calibration,
admission, Official package, and verifier evidence.
The checked Core schema
requires runner.identity_kind to remain source_only and
runner.built_binary_sha256 to remain null. The shared Rust validator now fails
closed on this runner subtree for both Core and Contrast. Contrast does not have
a separate checked-in JSON schema. Each corpus also binds the Node.js and ripgrep
identities. The source-only corpus rule and signed per-run runner and complete
Codex runtime provenance are the executable product contracts. The Codex runtime
is one private directory that contains exactly the codex executable and its
codex-code-mode-host sibling. After the final clean build, the operator retains
a private, unsigned audit receipt with the exact source commit and tree identity
and SHA-256 values for the native runner, verifier, Codex executable, and Codex
code-mode host. The offline native verifier validates this receipt against an
independently supplied receipt digest. It is not a database input or published
artifact. Node.js and ripgrep remain bound by the corpus commitment. Do not infer
a runtime hash from a generated-task tree digest. The accepted AIQ 2.0 publication
will be one batch of 17
configuration runs and 1,224 task-level executions.
Elapsed time, provider-token usage, and Standard API-equivalent cost are
reported separately from AIQ.
The Web application is a professional analysis workbench. Official evidence presents calibrated ability with its conditional 95% interval. Synthetic fixtures present descriptive quality with task-mix sensitivity and never appear as Official. Scientific context also reports strict pass with a Wilson interval, sample count, coverage, missing cells, runtime state, scoring method, and provenance. It keeps semantic task outcomes separate from runtime, invalid, and missing cells. Cost remains an estimated Standard API-equivalent comparison, not an actual ChatGPT or Codex subscription bill. Charts use ECharts with SVG rendering and ARIA descriptions. Users can select system, light, or dark color themes. Production views must use only real evidence for the sole production tuple, not synthetic or legacy data.
The 1.1.0 source candidate is candidate.20 at
benchmarks/candidates/aiq-core-1.1.0/. It preserves candidate.19 except for one
scored, digest-bound final-response reconciliation decision in tool-use-02.
Candidate.19 completed all 1,224 calibration cells but failed policy v2 because
the Tool Use domain mean facility was 0.949580, above the 0.90 ceiling.
Candidate.20 keeps the policy, 72 tasks, 17-model matrix, supplied tool, and
workspace receipt semantics unchanged. It still requires fresh independent
review, sealing, calibration, admission, and Official evidence.
benchmarks/candidates/aiq-core-1.1.0/catalog.json uses
aiq.catalog.v2. Candidate.20 keeps 71 response contracts and advances only
tool-use-02 from workspace-only scoring to a final-response reconciliation
decision while retaining its workspace and receipt checks. It keeps the
seven distinct tool-use constructs. It records 29 private tasks as retained and
43 as repaired after candidate.15 calibration. It has 72 distinct
within-domain clusters. Every task requires gold, alternate_correct, partial,
adversarial_format, and empty; timeout is not_applicable under natural
completion. The catalog is the sole expected-class authority.
Its canonical catalog digest is
sha256:00e555904daa023f0f7731a0e7e66e12d833641f62eb4aaa59447be6417807d7.
Its ordered task-metadata digest is
sha256:3580555315d49a62b28b6947491819276dca5b261ade802f10b33808569d1708.
Its public release digest is
sha256:5b651845280ea0b27a8cfc2aec4efe4c149bbbfdeeb0e5a9c883174938b58d69.
Candidates.1 through .18 are immutable predecessor evidence. Candidate.5
remains the durable source for the seven distinct disclosed scenario, operation,
result, evaluator, metamorphic, and cross-task substitution contracts. Its
source integration was rejected because
its catalog task-metadata identity was
sha256:cfac96630c9efe3153d80ed43effd6e541bef751e1e7f766a52cfb2910fa3fc4,
while the Rust commitment consumer and public v3 schema still required
sha256:393cb2563b2161ccb42dd5a50ea63a7827f4d5c485ca0a98103e80eef3d0fbe6.
Candidate.6 removed that duplicate Rust identity authority. Candidate.7 kept the
catalog-derived commitment authority and repaired the trusted execution and
qualification-evidence bridge. Candidate.7 is rejected because its completed-run,
recovery, and package paths returned to the active 1.0.7 validator after candidate
preparation. Candidate.8 carries one provenance-bound validation context through
those paths, but its package command derived that context from the saved record and
its private runtime used Node.js 24.19.0. Candidate.9 requires independent task,
corpus, and source inputs for candidate packaging. It binds the complete validated
context through signed-payload serialization and uses Node.js 24.18.0. The active
1.0.7, Contrast, historical, and Official validators remain
unchanged. The commitment validator derives the expected candidate identity from
the validated embedded catalog and rejects candidate.14 and older identities.
Candidate.9 is rejected because debugging-04 declared src/task.mjs while its
prompt, workspace bindings, and weighted evaluator import src/task.ts, and
instruction-following-05 declared the non-schema field type undefined for
calculation_note. Candidate.10 corrects those values to src/task.ts and
string, but its location check is rejected because candidate.3's versioned
response contract selected the source locations used to validate candidate.3
through candidate.5. Candidate.11 preserves the corrected leaves and all task,
evaluator, fixture, and tool semantics. Its separately tracked
benchmarks/candidates/aiq-core-1.1.0/task-response-authority.json file owns the
public-safe response mode and locations for every task. The generator validates
all versioned contracts against that unchanged projection. Candidate.11 is
rejected because its tracked private validator treated protected
expected_file_sha256 inputs as response outputs and inferred final response
from workspace-policy absence. Candidate.12 corrects that existing owner only.
It requires one hard-gate complete_workspace_policy, detects response_*
checks for final response, excludes protected inputs from the mutable allowlist,
and derives workspace locations from progress files plus evaluator path
targets, with evaluator-source fallback only when both are empty. The immutable
tasks resolve to 71 workspace responses and one final response: progress is exact
for 66 workspace tasks, empty for four, and a strict subset for one.
Candidate.12 is rejected before model invocation because file loading used lexical
task order while candidate validation required checked-catalog order. Candidate.13
uses the existing catalog-order owner during candidate-only preparation. It also
selects the v3 commitment schema only on that same candidate preflight route; the
active and standalone preflight routes continue to require v2.
Candidate.13 is rejected operational evidence. Its isolated Jordan qualification
completed all 17 preflight probes and started 88 task cells, but only 56 of 1,224
cells completed before the five-hour subscription quota reached 100%; seven-day
usage was 17%. No package, verifier stage, attestation, or qualification artifact
exists. Candidate.14 keeps every task-facing semantic unchanged and replaces only
the quota-infeasible candidate release-qualification shape. Its one 216-cell run
and package completed, but the real candidate verifier loaded ordinary filenames
lexically and rejected the catalog-ordered evaluator identity before replay. No
stage, attestation, or qualification artifact exists. Candidate.15 reuses the
existing checked-catalog ordering owner in that candidate-only verifier load path.
Its complete 1,224-cell calibration had no runtime failures, but policy v2
rejected 38 universal full-credit tasks, ten universal semantic-zero tasks, only
22 non-uniform tasks, and six degenerate domains. Candidate.16 keeps policy v2,
repairs those task-bank failures, and normalizes only the exact Codex zsh
transport before hashing the logical ToolUse command.
Independent review rejects candidate.16 because all 43 revised tasks reject a
publicly declared optional field; six Documentation tasks also score optional
next_steps as required. Candidate.17 repairs that existing evaluator owner and
the independently identified 1.1.0/v3/13-view runbook drift.
Independent source review rejects candidate.17 because one Security model
sentence still says 12 public views and the focused regression did not match
that wording. Candidate.18 corrects that final documentation owner. Its isolated
Morgan calibration reached 288 checkpoint results before Codex returned
Selected model is at capacity; task output containing authentication caused
the adapter to misclassify that temporary capacity event as terminal authentication.
Candidate.19 routes the exact structured capacity event through existing resumable
backpressure and limits provider-failure classification to stderr or structured
error and turn.failed messages.
Its complete calibration then failed only the Tool Use domain facility ceiling;
candidate.20 adds the bounded tool-use-02 response decision described above.
The exact 42 task-issue closures remain unchanged. The catalog records the unauthenticated candidate execution and qualification-evidence bridge as one separate source-integrity closure. It records the candidate.7 end-to-end validation failure as a second source-only closure. Neither closure counts as a task issue. The package-input, Node.js runtime, public response-contract, and private response-source-owner repairs are four additional source-only closures. None counts as a task issue.
For each tool-use task, the hard gate requires exactly one total tool call and
one command_execution call. It also requires one completed command-line digest for
node bin/task-tool.mjs:
sha256:6763cc80f8294b52c6494f1c9891e41a8e3cd1c466ca622377c59643a0466319.
The separate receipt command_sha256 continues to identify the supplied tool
file bytes. The runner retains digest counts only for task-declared required
command identities; exact total and per-tool counts still expose undeclared or
extra calls. One three-configuration qualification matrix can contain at most
21 declared digest entries. The runner removes command text from provider
stdout and stderr evidence, but it
extracts and preserves the exact semantic final response before that log
redaction. Candidates.14 through .17 remain inactive and not production-publishable.
Candidate.18 is the current source candidate; production cutover requires its
fresh review, seal, complete calibration, admission, and Official evidence.
AIQ Core 1.1.0 sealing also requires one independently supplied
aiq.leakage-review.v2 record for every task. Each record binds the reviewer,
the reviewer task or thread, review time, source commit and tree, source
manifest, task definition, catalog entry, verdict, method, scope, and notes.
The sealer copies and hashes the supplied record. It rejects missing, extra,
rejected, stale, or mismatched records. It never creates a completed review
from task-authored notes. Recorded process separation is review evidence. It is
not cryptographic proof that a human reviewer was independent. The v1 review
contract remains isolated to frozen 1.0.7 compatibility.
The shared aiq-runner library owns
aiq.benchmark-qualification-policy.v2. Both existing executables use that
same in-process implementation:
cargo run -p aiq-runner -- qualify-candidate --help
cargo run -p aiq-verifier -- verify-qualification --helpQualification accepts exactly one predeclared, complete, non-synthetic 3-by-72 Calibration stage: Sol medium, Terra medium, and Luna medium in that order over all 72 catalog-ordered tasks. Its verifier-derived 216 cells must all be semantically complete, and its attestation must come from the predeclared verifier. The manifest fixes the candidate, corpus, source, model selection, run, and verifier identities before execution. The final artifact additionally binds the exact signed package, runner, stage, attestation, provenance, and matrix digests.
This is an end-to-end execution and identity qualification only. It makes no prediction-interval, Spearman-correlation, run-variance, or precise-rank claim. Those stability-only fields are absent from the v3 manifest and artifact. The active/default production matrix remains all 17 configurations. Any task, evaluator, policy, or source-identity revision requires a new reviewed source identity and fresh release evidence.
| Path | Purpose |
|---|---|
apps/aiq/ |
Scheduled observation orchestration, release validation, and cleanup |
apps/aiq-runner/ |
Capability checks, task execution, scoring, packaging, and submission |
apps/aiq-verifier/ |
Queue claims, artifact reconstruction, evaluator replay, and attestations |
apps/web/ |
Public Next.js site and controlled server gateways |
benchmarks/ |
Public catalog, schemas, and synthetic examples |
databases/ |
Desired database state, fresh initializer, and disposable SQL checks |
openwiki/ |
Architecture, method, operations, and deployment handoff |
Private tasks, expected outputs, controlled evaluators, signing keys, Codex authentication, and production data must stay outside Git.
Use Node.js 24.15.0 or newer, the npm 11.17.0 version pinned by package.json,
the stable Rust toolchain selected by rust-toolchain.toml, and the locked
dependencies. cargo make fmt
also requires a separately managed nightly rustfmt toolchain.
npm ci --ignore-scripts
cargo run -p aiq-runner -- demo
npm run devOpen http://localhost:3000. When both public Supabase variables are absent in
development, the site uses checked-in synthetic data. Production fails closed
when its configuration is incomplete.
Useful runner commands:
cargo run -p aiq-runner -- matrix
cargo run -p aiq-runner -- validate --public-tasks benchmarks/examples/tasks
cargo run -p aiq-runner -- validate-core-corpus --help
cargo run -p aiq-runner -- validate-contrast-corpus --help
cargo run -p aiq-runner -- --help
cargo run -p aiq-verifier -- --help
cargo run -p aiq-verifier -- diagnose-rescore --helpInstall the Playwright browsers once on a fresh host, then run the complete local browser gate:
npm exec --workspace @aiq/web -- \
playwright install --with-deps chromium firefox webkit
cargo make check
cargo make verifycheck is the complete read-only source gate. It checks the database schema,
TypeScript, Rust and TypeScript lint rules, vstyle, and all Rust and TypeScript
tests. verify extends check with one Web build and every local browser
acceptance suite. fmt is an independent mutating action; it is not a dependency
of either gate. Native release builds, production browser checks, database
runtime checks, deployment, and publication remain separate contracts. Do not run
component tasks again in the same validation pass. Coverage instrumentation is
opt in with cargo make test-typescript-coverage.
The two subscription smokes are ignored and opt in. Each consumes one Codex subscription attempt.
cargo make smoke-subscription
cargo make smoke-controlled-subscriptionThe public-task smoke validates a fixed example. The controlled-task smoke needs operator-supplied private task, evaluator, corpus, runtime, workspace, and Codex inputs. Neither smoke creates a benchmark result.
databases/schema.sql is the sole desired database state.
databases/init.ts is the only production initialization entry point. There is
no migration chain. It opens
one PostgreSQL connection and applies the schema plus public reference data in
one transaction. It accepts the direct host or exact port-5432 session pooler
identity for personal Supabase project xxnszykaeapolqdnhalx. An explicit
test/development override accepts
only a loopback target and cannot apply in production. It rejects a database
that already contains the AIQ schema, gateway roles, or either exact AIQ
Storage bucket identity. Apply this one greenfield desired state to the existing
target project only after its AIQ namespace is empty. If AIQ residue exists, the
operator must
remove only aiq_private, the two AIQ gateway roles, and the exact AIQ-owned
public views and RPC overloads. Preserve all Supabase-managed and non-AIQ
objects. This cleanup is a deployment prerequisite, not a migration or
compatibility path. The schema creates the aiq-submission-packages and
aiq-runner-artifacts Storage buckets as private. The preflight rejects either
existing bucket identity. Do not create the buckets in a separate operator step.
The preflight enumerates the 13 canonical public view names and all public RPC
names from the desired state. It rejects every overload of those exact RPC
names without matching unrelated public objects.
AIQ_DATABASE_URL='<direct-or-session-pooler-url>' \
AIQ_PRODUCTION_REFERENCE=/controlled/production-reference.json \
cargo make init-databaseFor an empty AIQ namespace, the production reference must contain the real
controlled, non-synthetic AIQ Core 1.1.0 v3 corpus commitment, its real canonical
published_at timestamp, and
exactly three public identities: runner, verifier, and publisher. Prepare it
only after the controlled corpus passes model-free validation, the operator
verifies the final native build, and one real signed non-synthetic 17-by-72
package passes native verifier replay; the repository contains no substitute
production reference. Retain the private final-build audit receipt separately.
Database initialization does not accept or validate that receipt.
A successful initialization receipt must report aggregate scoring 1.0.8, both public
catalog identities, 72 tasks, 17 model configurations, and three nodes.
Use one initialized disposable database for production-shape smoke and calibration publication checks:
cargo make smoke-database
AIQ_DATABASE_URL='<direct-or-session-pooler-url>' cargo make smoke-calibration-databaseUse a separate fresh PostgreSQL 17 database for the deterministic synthetic flow:
psql "$AIQ_DATABASE_URL" -X --set ON_ERROR_STOP=1 \
--file databases/schema.sql
psql "$AIQ_DATABASE_URL" -X --set ON_ERROR_STOP=1 \
--file databases/synthetic-demo.sql
psql "$AIQ_DATABASE_URL" -X --set ON_ERROR_STOP=1 \
--file databases/integration.sqlDo not apply the synthetic flow to the initialized production-shape database or to production.
- The runner validates the controlled corpus, toolchain, and capability manifest.
- It executes the selected tasks and runs each formal evaluator against the sealed model evidence. A retryable evaluator process failure keeps the model evidence pending for evaluator-only recovery. The runner writes content-addressed artifacts.
- It scores the run, records efficiency evidence, and signs one v4 result package.
POST /api/submissionsstores the exact package bytes and queues the package as unverified.- The verifier claims the package, reconstructs the workspaces, and executes each deterministic evaluator once for that claim attempt. An operational replay failure releases the claim for retry without changing the retained package or invoking a model.
POST /api/verificationsstages the normalized batch and records the signed verifier attestation.- A distinct publisher identity completes publication through the gateway.
- Public security-invoker views supply the Web application.
Official means a complete, non-synthetic 17-by-72 run with valid task-set
1.1.0, task-scorer 1.0.6, aggregate-scorer 1.0.8, and measurement 2.0.0
bindings that completed this flow and was published as
trusted_verified. A complete synthetic fixture uses the
synthetic_complete classification, has no Official AIQ value, and is never
ranking eligible. There is one submission, native verification, and publication
path.
The only Official execution and publication path runs aiq-runner and
aiq-verifier natively on the controlled Apple Silicon macOS host with direct
network access. Use the release binaries in this order:
admit-permissions, preflight, run, score, package, submit, and then
verifier replay. admit-permissions is model-free; preflight is the first paid
step. Only its exact configuration probes and runnable task cells in run
invoke models. Scoring, packaging, submission, verifier replay, and publication
do not invoke models. The same private admission receipt binds preflight through
package.
Provide the runner signing key only to package, the submission token only to
submit, and verifier credentials only to the verifier command.
The source runner targets native macOS. Linux and Docker remain future
deployment targets. A frozen aiq release built from this source starts the
observation scheduler at 03:00 and 15:00 UTC. It selects one canonical
12-hour slot, holds a global nonblocking lock, and does not start a second
scheduler. An installed frozen release keeps its existing behavior until an
operator replaces it. A self-contained release
stores the pinned runner and verifier binaries and an exact Git source bundle.
aiq restores the clean detached source at a stable per-slot path below the
private state_root/scratch directory. This path stays outside the macOS
platform-minimal roots that model processes can read. The command does not use a
repository worktree at run time. On a host fixed to the America/New_York time
zone, the macOS launchd template wakes at 11:05, 11:35, 23:05, and 23:35 local
time. The four wakes cover EST and EDT with one bounded retry for each UTC slot.
Official task dispatch must begin during the first two hours of a slot, and the
v2 configuration requires all 32 supported workers for the fixed 1,224-cell
matrix. A late wake does not start a new matrix. Temporary selected-model capacity and
subscription quota, usage, or rate limits are persisted as non-terminal backpressure:
completed cells stay in the same checkpoint, rejected cells remain pending, and later
scheduled wakes resume the oldest blocked slot before they can start newer paid work. The
scheduler starts Official and Speed as sibling publication paths for the same
slot. Official keeps its two-hour model-dispatch grace. Speed has an independent
12-hour slot window. After the scheduler grants dispatch, neither path waits for
the other path. A slow or failed path cannot block the other path's dispatch or
publication. Each path writes retained status below its own slot directory, and
aiq status composes both outcomes. A completed run
with a non-semantic infrastructure result is retained as unpublished evidence.
It is not retried or presented as an AIQ score. Provider-capacity backpressure is not a
completed result and is therefore the sole exception to that terminal rule.
The subscription runner uses a protected copy of ~/.codex/auth.json in an
isolated per-release CODEX_HOME; it does not reuse the interactive Codex home
as its writable runtime directory. It also uses a private two-file copy of the
ChatGPT app's codex and codex-code-mode-host executables. Capability
preflight succeeds only after Codex completes one command and writes the exact
content-bound marker in a fresh disposable workspace.
See Operations and Validation for the native command
contract. Repository support does not prove that private inputs, credentials,
or live model capabilities are configured.
Normal/Fast transport measurements are auxiliary evidence. observe-speed
reads the live Codex model catalog before any paid turn, records an exact
available, unsupported, or unavailable state for each selected configuration,
and runs paired Normal/Fast fixed-response trials only for advertised modes.
It records completion, total elapsed time, aggregate output throughput, token
usage, tool use, and estimated ChatGPT credits. It does not calculate or modify
AIQ. The current Codex JSONL stream does not expose a trustworthy first-token
timestamp, so TTFT and post-first-token throughput remain explicit unavailable
values instead of estimates.
cargo run -p aiq-runner -- observe-speed --help
cargo run -p aiq-runner -- submit-speed --help
cargo install --locked --path apps/aiq
aiq status --config /absolute/private/path/to/continuous-observation.json
aiq doctor --config /absolute/private/path/to/continuous-observation.json
aiq run --config /absolute/private/path/to/continuous-observation.json
aiq run --config /absolute/private/path/to/continuous-observation.json \
--slot 2026-08-12T03-00ZUse --slot only for one known canonical UTC slot. Official task dispatch can
start only during the first two hours of its current slot. The frozen runner
can resume an unchanged checkpoint during the same slot only when it contains
no indeterminate in-flight cell. A checkpoint with explicit subscription
backpressure can also resume after the dispatch grace or 12-hour slot window;
it reuses the exact admitted preflight and never replaces completed cells. Other
checkpoints with sealed pending evaluator work can also resume that work after
the window without another task-model invocation. This rule includes a retryable
evaluator process failure. An indeterminate model cell
still fails closed after all sealed pending evaluator work is recovered. Other
late slots can continue only when the complete Official run output already
exists and only scoring or publication remains. aiq recognizes that output
only as an aiq.run.v4 document with all 1,224 results; the runner's
create-once reservation is not a completed run. Otherwise, aiq
records a terminal missed or unpublished state without new model work. Speed
model dispatch can start during its own 12-hour slot. An existing Speed batch
can resume submission after that window without new model work. A
terminal slot remains a no-op. If another
observation owns the global lock, a scheduled run coalesces successfully
without starting another model process; doctor reports the contention instead.
Each runner, verifier, and evaluator step runs below an internal supervisor that
owns a separate process session. Runner-created model and evaluator process
groups remain in that session. A private pipe binds the supervisor to the
user-facing aiq parent. If that parent exits or is killed, the pipe closes and
the supervisor repeatedly sends SIGTERM, then SIGKILL, to every remaining
session process before it exits. This no-orphan boundary does not depend on
launchd process-group cleanup.
Start from config/continuous-observation.example.json and
config/com.acgbox.aiq.continuous-observations.plist.example. Keep the concrete
configuration and launchd plist outside Git. Use aiq install-release once to
copy the minimal frozen release, create its source bundle, and print the release
manifest digest. Install the release in a versioned directory outside the
repository. The private v2 configuration contains stable runtime paths, limits,
the endpoint, the manifest digest, and optional non-secret unattended provider
metadata. It does not contain a source worktree path, worker executable path,
provider credential, or consumer secret.
cargo install is sufficient for local operator use. An unattended service
must pin apps/aiq/package.nix in the host configuration so an unrelated Cargo
install cannot replace the scheduled executable.
Set official_jobs to 32; lower values are rejected before model work. The
aiq run accepts either all four explicitly supplied consumer variables or no
consumer variables. Partial ambient delivery fails closed. When all four are
absent, aiq requires the complete unattended_secrets metadata, reads the
exact Keychain bootstrap, performs one Universal Auth login, and retrieves only
the four fixed prod:/aiq keys. It removes the provider session before it starts
a downstream step. Provider credentials and tokens do not reach workers. The
orchestrator gives the signing key only to package, the submission token only
to submission steps, and verifier credentials only to the verifier. Each owner
uses a fresh isolated CODEX_HOME directory for the slot. A retryable slot
retains checkpoints and raw artifacts. Checkpoint v10 distinguishes
indeterminate model work from sealed pending evaluator work. The latter resumes
from the same model response and workspace without another model invocation. A
retryable evaluator process failure stays in this pending state and cannot
create a terminal run. Provider-declared temporary model capacity and subscription limits
leave the affected cells pending under aiq.subscription-backpressure.v1; the runner
also migrates v9 checkpoints to the v10 evaluator-resume shape and legacy v8
checkpoints that incorrectly committed those limits as terminal results. On resume, aiq revalidates the permission
admission, complete Official run, submission receipts, and verifier receipt
before it reuses them. It stores non-success verifier records in a private
append-only attempt log. A create-once success receipt is valid only when its
package SHA-256 and idempotency identity match the exact local package. Copied
credentials are removed after each invocation. A terminal slot keeps both
owner-status records, the compact batch,
package, score, attestation, and receipts. It removes the detached source, raw
local artifacts, replay scratch, checkpoints, and disposable workspaces.
launchd invokes the pinned aiq run --config ... command directly. Use
absolute AIQ and configuration paths, supply HOME, USER, LOGNAME, and the
pinned execution PATH, and do not set a repository working directory. The
provider identity must grant only the four fixed source keys. The runtime keeps
the Keychain bootstrap and short-lived provider token inside AIQ.
The already-provisioned provider target is external frozen state. Do not run
setup to reconcile, rotate, or replace it. For a new exact target only, the
hidden aiq operator provision-unattended --config ... command uses
config/unattended-provider-provision.example.json. It refuses an existing
Keychain account or provider identity, creates only the fixed identity,
four-key privilege, Universal Auth method, and Keychain bootstrap, and rolls
back only known intermediate writes.
- Keep runner, verifier, and publisher credentials separate.
- Keep privileged Supabase values in server-only environment variables.
- Keep
aiq-submission-packagesandaiq-runner-artifactsprivate. - Use RLS and the narrow database RPCs; do not write private tables from the browser.
- Put authentication, request limits, and a WAF in front of write routes.
- Run the Storage reconciliation worker before the deletion worker.
- Treat readiness responses as bounded dependency evidence, not deployment proof.
See OpenWiki quickstart, operations, and deployment handoff for the maintained details.
aiq.wiki is canonical, and www.aiq.wiki returns a permanent 308 redirect
that preserves the request path. Automatic Vercel project and branch aliases can
be removed only transiently because a later deployment can recreate or reassign
them. A deployment-specific URL is intrinsic to its retained deployment. The
current generated Vercel surfaces emit noindex.