Problem
Today ExaMLOps has no answer to "what model could I start from?" — only to "what have we
already trained and promoted?" (the MLflow-backed model registry: versions, Staging/Canary/
Production aliases, drift and traffic history, populated by a training run). An operator starting
a new project has to already know a model definition to write one; there is no curated, browsable,
provenance-tracked list of deployable model definitions — base models or validated recipes/configs
— to pull in as a starting point.
The distinction (this is the design's central point)
|
Model Catalog (proposed) |
Model Registry (exists today) |
| Answers |
"What could I start from?" |
"What did we produce, and what's live?" |
| Populated by |
a curated, signed/reviewed publish |
a training run |
| Identity |
name@catalog_version (immutable, content-hashed) |
name/version with movable aliases |
| Mutable after publish |
never — a correction is a new version |
yes — aliases move, versions archive |
| Its only verb toward a project |
pull → materializes a new model definition, trains/deploys nothing |
promote, serve, drift/traffic control |
Conflating the two — e.g. treating a catalog entry as already-deployable, or storing catalog
entries as registry versions — is the most likely design mistake, so the proposal states this
distinction explicitly and keeps the two systems structurally separate.
Proposed data model — CatalogEntry
An immutable, content-addressed, versioned record (same shape discipline as this project's
existing agent-version registry: canonical-JSON content hash, insert-only rows):
name, catalog_version, entry_hash (sha256 over canonical JSON of the rest)
kind: base_model | recipe
source: either a reference into the project's existing external-data-connector layer, or a
pinned external URI (an unpinned/floating reference — no "latest" tag, no branch — is refused at
publish time, never accepted and flagged later)
license (SPDX identifier, required)
resource_hint (a hardware/resource-profile hint)
eval_summary_ref (a pointer into the existing evaluation-suite-results store — no duplicate
score storage)
provenance.supplychain_ref + provenance.trust_tier (T1_signed | T1_unsigned) — reusing
the project's existing model-signing/verification module exactly; no second signing mechanism
recipe.model_yaml_template (for kind: recipe) — renders into the platform's existing
per-model YAML format, the single source of truth for a model definition, unmodified
Trust is always visible: an unsigned source is flagged in every listing and in the pull
confirmation, never silently treated as equal-trust to a signed one.
Proposed CLI surface
exa catalog list [--kind base_model|recipe] [--license <spdx>] [--evaluated-only] [--trust-tier ...]
exa catalog show <entry>[@<version>]
exa catalog pull <entry>[@<version>] --project <project> [--as <model-name>] [--dry-run]
exa catalog publish <path-to-entry.yaml> # admin-tier, no self-service publish
exa catalog withdraw <entry>@<version> # destructive-tier, audited
pull renders the entry's recipe into a new per-model YAML, assigns it to the target project via
the existing project-resource-membership mechanism, and records a lineage edge via the existing
lineage/provenance module — no new lineage system, no new project-membership table.
Phased plan with acceptance criteria
Phase 1 — static/curated catalog, read-only.
Validation module + storage table + exa catalog list|show|publish.
Acceptance: every listed entry has a populated trust tier; entry_hash round-trips on show;
publishing an entry with an unpinned source is refused with a named reason.
Phase 2 — pull into a project + lineage.
exa catalog pull materializes the per-model YAML, assigns project membership, emits a lineage
edge, records an append-only pull row.
Acceptance: a pulled entry's YAML passes the existing registry-integrity CI guard unmodified; the
lineage graph shows the catalog entry upstream of the new model; --dry-run writes nothing.
Phase 3 — signature/provenance verification wired to the existing supply-chain module.
exa catalog publish --sign and a verify-before-materialize gate on pull (enforce/warn modes),
reusing the existing model-artifact signing mechanism for the entry manifest itself.
Acceptance: a tampered entry fails verification in enforce mode with a named reason.
Phase 4 — dashboard catalog browser.
Read API + admin-gated pull endpoint (same code path as the CLI, no parallel implementation) +
a Catalog console page with a per-project "pull into this project" action.
Acceptance: capability-gated (viewer read-only, admin pull); a dashboard pull produces an
identical database row and lineage edge to the CLI path.
Likely modules touched
- New:
examlops/catalog/ (manifest validation, storage, pull logic)
- New:
examlops/cli/commands/catalog_cmd.py
- Reused, unmodified: the model-supply-chain signing/verification module, the lineage/provenance
module, the evaluation-suite-results store, the project-resource-membership mechanism, the
external-data-connector layer (for source.kind: dataplane-style refs)
- New (Phase 4): a dashboard catalog router + frontend console
This is a design proposal only — nothing described above is implemented yet.
Problem
Today ExaMLOps has no answer to "what model could I start from?" — only to "what have we
already trained and promoted?" (the MLflow-backed model registry: versions,
Staging/Canary/Productionaliases, drift and traffic history, populated by a training run). An operator startinga new project has to already know a model definition to write one; there is no curated, browsable,
provenance-tracked list of deployable model definitions — base models or validated recipes/configs
— to pull in as a starting point.
The distinction (this is the design's central point)
name@catalog_version(immutable, content-hashed)name/versionwith movable aliasespull→ materializes a new model definition, trains/deploys nothingpromote,serve, drift/traffic controlConflating the two — e.g. treating a catalog entry as already-deployable, or storing catalog
entries as registry versions — is the most likely design mistake, so the proposal states this
distinction explicitly and keeps the two systems structurally separate.
Proposed data model —
CatalogEntryAn immutable, content-addressed, versioned record (same shape discipline as this project's
existing agent-version registry: canonical-JSON content hash, insert-only rows):
name,catalog_version,entry_hash(sha256 over canonical JSON of the rest)kind:base_model|recipesource: either a reference into the project's existing external-data-connector layer, or apinned external URI (an unpinned/floating reference — no "latest" tag, no branch — is refused at
publish time, never accepted and flagged later)
license(SPDX identifier, required)resource_hint(a hardware/resource-profile hint)eval_summary_ref(a pointer into the existing evaluation-suite-results store — no duplicatescore storage)
provenance.supplychain_ref+provenance.trust_tier(T1_signed|T1_unsigned) — reusingthe project's existing model-signing/verification module exactly; no second signing mechanism
recipe.model_yaml_template(forkind: recipe) — renders into the platform's existingper-model YAML format, the single source of truth for a model definition, unmodified
Trust is always visible: an unsigned source is flagged in every listing and in the pull
confirmation, never silently treated as equal-trust to a signed one.
Proposed CLI surface
pullrenders the entry's recipe into a new per-model YAML, assigns it to the target project viathe existing project-resource-membership mechanism, and records a lineage edge via the existing
lineage/provenance module — no new lineage system, no new project-membership table.
Phased plan with acceptance criteria
Phase 1 — static/curated catalog, read-only.
Validation module + storage table +
exa catalog list|show|publish.Acceptance: every listed entry has a populated trust tier;
entry_hashround-trips onshow;publishing an entry with an unpinned source is refused with a named reason.
Phase 2 — pull into a project + lineage.
exa catalog pullmaterializes the per-model YAML, assigns project membership, emits a lineageedge, records an append-only pull row.
Acceptance: a pulled entry's YAML passes the existing registry-integrity CI guard unmodified; the
lineage graph shows the catalog entry upstream of the new model;
--dry-runwrites nothing.Phase 3 — signature/provenance verification wired to the existing supply-chain module.
exa catalog publish --signand a verify-before-materialize gate onpull(enforce/warn modes),reusing the existing model-artifact signing mechanism for the entry manifest itself.
Acceptance: a tampered entry fails verification in enforce mode with a named reason.
Phase 4 — dashboard catalog browser.
Read API + admin-gated pull endpoint (same code path as the CLI, no parallel implementation) +
a Catalog console page with a per-project "pull into this project" action.
Acceptance: capability-gated (viewer read-only, admin pull); a dashboard pull produces an
identical database row and lineage edge to the CLI path.
Likely modules touched
examlops/catalog/(manifest validation, storage, pull logic)examlops/cli/commands/catalog_cmd.pymodule, the evaluation-suite-results store, the project-resource-membership mechanism, the
external-data-connector layer (for
source.kind: dataplane-style refs)This is a design proposal only — nothing described above is implemented yet.