Problem
Every surface in ExaMLOps that requests compute today specifies its shape with raw, surface-local
flags, and each surface has grown its own vocabulary independently:
exa pipeline run --gpus N --cluster <name|auto> passes a bare GPU count into a
ResourceAsk(gpus=..., cpus=..., nodes=...).
exa pipeline distributed launch <model> --nodes N --gpus-per-node N --strategy fsdp|zero|megatron
builds a distributed-training launch command from the same underlying shape, with different flag
names.
exa workbench create <name> --project <p> --cpu <cores> --memory-gb <gb> writes cpu/memory_gb
straight into the workbench record — with no GPU dimension at all today.
- Ray Serve replicas get a fractional GPU only via an internal autoscaling config value; no
per-model config expresses a GPU/CPU/memory shape declaratively.
None of these compose. An operator who has a resource shape they like — say, "2× A100, 8 CPU
cores, 32GB RAM, CUDA 12.4" — cannot name it once and reuse it across a workbench, a training run,
and a serving deployment. They re-type the equivalent flags at every call site, in a different
flag vocabulary each time, with no record of why a given shape was chosen or that two call sites
were even meant to match. This is straightforward operational drift, and it gets worse as the
platform grows more compute surfaces (HPC fleet placement, distributed training, serving
autoscaling all already exist independently).
Proposal: Hardware Profiles
Add Hardware Profiles — reusable, named, versioned bundles describing
{accelerator family, resource shape (GPU/CPU/memory), driver/runtime requirements, applicability}
— that any compute-requesting surface can reference by name instead of restating raw resource
flags.
A profile is a named preset, resolved into the platform's existing scheduler-neutral resource
representations at the point of use. It is deliberately not a new resource vocabulary — it
maps onto the resource/placement types the platform already has for HPC job admission and
placement, so using a profile reaches the exact same code paths a hand-typed --gpus 2 reaches
today.
Data model
HardwareProfile — one immutable version of a named profile:
| Field |
Type |
Notes |
name |
string |
slug, e.g. "train-2gpu-a100" |
version |
int |
monotonic per name, assigned on write |
accelerator_family |
string |
nvidia | amd | intel-gaudi | tpu | cpu |
accelerator_model_hint |
string, optional |
advisory, e.g. "A100-80GB" — never trusted without a live capacity check |
gpu_count |
int |
>= 0 |
gpu_fraction |
float |
0 < f <= 1.0, default 1.0 — fractional/MIG-aware |
mig_profile |
string, optional |
e.g. "1g.5gb" |
cpu |
float |
cores |
memory_gb |
float |
|
nodes |
int |
>= 1, default 1 |
driver_tag |
string, optional |
e.g. "cuda-12.4" |
runtime_tag |
string, optional |
e.g. "pytorch-2.4-cu124" |
applicability |
list |
subset of workbench | training | serving | any |
description |
string |
|
A profile name is versioned like a small append-only log: writing a profile never mutates an
existing version, it creates a new one and moves a movable label (default active) to point at
it — so a workbench created last month keeps a record of exactly which version it used, even after
the name has since been updated.
Honest resolution, never fabrication. Resolving a profile against a target cluster/scheduler
must classify the result as one of:
unchecked — no target given (e.g. a local/dev run); the raw ask is returned with no capability
claim.
verified — live discovery of the target confirms the requested GPU count/model.
degraded — the coarse ask (GPU count/CPU/memory) is satisfiable, but a finer claim on the
profile (a specific accelerator model or MIG profile) cannot be confirmed from what the current
discovery probe reports for that target — resolution proceeds on the coarse ask and says exactly
which fields it could not confirm.
unresolvable — the ask exceeds the target's total capacity; resolution is refused outright.
This mirrors the "never silently pretend a capability" convention already used elsewhere in the
platform's HPC/scheduler code (honest fallback on unsupported hardware features, honest capability
reporting on checkpoint/resume) — a hardware profile must never claim a piece of hardware exists
just because an operator once typed its name into a profile.
CLI surface
Folded into the existing hardware-placement command group as a profile subcommand (rather than
a new top-level group), since that group already owns the neutral device-requirement vocabulary a
profile is a named, saved preset of:
exa hardware profile set <name>
--accelerator-family nvidia|amd|intel-gaudi|tpu|cpu [required]
--gpu <int> --gpu-fraction <float> --mig-profile <str>
--cpu <float> --memory-gb <float> --nodes <int>
--accelerator-model-hint <str> --driver-tag <str> --runtime-tag <str>
--applicability <csv> --description <str> --label <str>=active
exa hardware profile list [--applicability workbench|training|serving|any]
exa hardware profile show <name> [--version <int> | --label <str>=active] [--cluster <name>]
exa hardware profile resolve <name> --cluster <name> [--for workbench|training|serving]
exa hardware profile delete <name> [--version <int>] [--yes]
list/show/resolve are read-only; set is a persisting write; delete is a destructive,
confirmation-gated removal.
Phased implementation plan
Phase 1 — Profile registry + CLI (no consumer wiring).
Add the profile record + versioning/label persistence, the pure resolution logic (including the
honest unchecked/verified/degraded/unresolvable classification against live cluster
discovery data), the exa hardware profile ... CLI, and CLI-surface tier classification (which
commands are read-only vs. mutating, for any UI that introspects the CLI). Every mutation is
audited.
Acceptance: creating a profile and resolving it against a known cluster returns the correct
status for each of the four resolution states above (unit-tested, including against a cluster with
no discovery snapshot and a cluster whose capacity is insufficient); versioning round-trips (a
second set creates a new version, an old version stays retrievable by number, and the movable
label points at the newest write); delete-without-version removes every version and its label,
delete-with-version removes only that one.
Phase 2 — Workbench integration.
Let workbench creation accept an optional hardware-profile reference; when given, unset resource
fields on the workbench are filled from the resolved profile (explicit flags on the same call
still win, so a profile is a default, not an override). The workbench record additively gains
columns for which profile (name + version) it was created from, so it stays inspectable after the
profile itself is later updated. Likely touched:
platform/cli/src/examlops/workbenches/__init__.py,
platform/cli/src/examlops/cli/commands/workbench_cmd.py.
Acceptance: creating a workbench with a profile reference produces the profile's CPU/memory
shape; overriding one flag on the same call overrides only that field; a profile whose
applicability excludes workbench is rejected with a clear error, not silently accepted.
Phase 3 — Training & serving integration.
Add an optional hardware-profile flag to the training-run and distributed-training-launch
commands, resolved into the same resource-ask shape those commands already build from raw flags.
Add an optional, additive hardware-profile reference to the per-model serving configuration, so a
deployed model's GPU/CPU/memory shape can also come from a named profile. Likely touched:
platform/cli/src/examlops/cli/commands/pipeline.py,
platform/cli/src/examlops/cli/commands/distributed_cmd.py,
pipelines/pipeline_generator/, serving/ray_serving/.
Acceptance: a distributed-training launch driven by --hardware-profile <name> produces the
identical launch command that specifying --nodes/--gpus-per-node directly would have produced,
proving the profile is sugar over the existing resource path rather than a second execution path;
a model config referencing a profile deploys with the resolved GPU fraction/CPU applied exactly as
today's explicit per-model settings do.
Phase 4 — Dashboard + visibility surface.
A read-only dashboard view for listing/inspecting profiles (admin-gated create/delete), reusing
the same backend code paths the CLI uses rather than a parallel implementation, plus a
hardware-profile picker on the workbench/training/serving creation forms. Surface a profile's
current resolution status (verified/degraded/unresolvable) wherever it is actively in use, so a
degraded or unresolvable profile is visible operationally rather than silently stored.
Acceptance: the dashboard's profile list matches exa hardware profile list --json exactly (no
drift between the two surfaces); a non-admin viewer can list/inspect but not create or delete
profiles.
Why this design
- Reuses existing resource types instead of inventing a parallel vocabulary. A profile
resolves into the platform's existing scheduler-neutral resource-ask and accelerator-placement
types; nothing downstream of resolution needs to learn a new shape.
- Versioned like an append-only log, not an editable row, so "which shape did this workbench
actually get" stays answerable after the named profile is updated — the same pattern this
platform already uses for its prompt/versioning registry.
- Honest about capability, never fabricating a piece of hardware a live discovery probe hasn't
actually reported — consistent with how this platform already treats every other
capability-degradation boundary (scheduler feature support, checkpoint/resume capability,
fractional-GPU fallback).
Alternatives considered
- Leave resource specification ad hoc per surface. Rejected — this is the status quo causing
the drift described above.
- A server-side, admission-time object that mutates/injects a workload's resource spec.
Rejected — this platform's compute substrate is Docker Compose plus bare-metal/HPC scheduler
adapters, not a server-side admission-controller pipeline; a profile here is a client-side named
template resolved before the call, not a server-side mutating hook.
- Extend the existing point-in-time device-placement request type to also carry a name and be
the profile. Rejected — that type is created fresh per placement call and is not meant to be
versioned or referenced ahead of time; conflating the two would make "which version of this
profile did that workbench use" unanswerable.
Problem
Every surface in ExaMLOps that requests compute today specifies its shape with raw, surface-local
flags, and each surface has grown its own vocabulary independently:
exa pipeline run --gpus N --cluster <name|auto>passes a bare GPU count into aResourceAsk(gpus=..., cpus=..., nodes=...).exa pipeline distributed launch <model> --nodes N --gpus-per-node N --strategy fsdp|zero|megatronbuilds a distributed-training launch command from the same underlying shape, with different flag
names.
exa workbench create <name> --project <p> --cpu <cores> --memory-gb <gb>writescpu/memory_gbstraight into the workbench record — with no GPU dimension at all today.
per-model config expresses a GPU/CPU/memory shape declaratively.
None of these compose. An operator who has a resource shape they like — say, "2× A100, 8 CPU
cores, 32GB RAM, CUDA 12.4" — cannot name it once and reuse it across a workbench, a training run,
and a serving deployment. They re-type the equivalent flags at every call site, in a different
flag vocabulary each time, with no record of why a given shape was chosen or that two call sites
were even meant to match. This is straightforward operational drift, and it gets worse as the
platform grows more compute surfaces (HPC fleet placement, distributed training, serving
autoscaling all already exist independently).
Proposal: Hardware Profiles
Add Hardware Profiles — reusable, named, versioned bundles describing
{accelerator family, resource shape (GPU/CPU/memory), driver/runtime requirements, applicability}— that any compute-requesting surface can reference by name instead of restating raw resource
flags.
A profile is a named preset, resolved into the platform's existing scheduler-neutral resource
representations at the point of use. It is deliberately not a new resource vocabulary — it
maps onto the resource/placement types the platform already has for HPC job admission and
placement, so using a profile reaches the exact same code paths a hand-typed
--gpus 2reachestoday.
Data model
HardwareProfile— one immutable version of a named profile:name"train-2gpu-a100"versionaccelerator_familynvidia | amd | intel-gaudi | tpu | cpuaccelerator_model_hint"A100-80GB"— never trusted without a live capacity checkgpu_count>= 0gpu_fraction0 < f <= 1.0, default1.0— fractional/MIG-awaremig_profile"1g.5gb"cpumemory_gbnodes>= 1, default1driver_tag"cuda-12.4"runtime_tag"pytorch-2.4-cu124"applicabilityworkbench | training | serving | anydescriptionA profile name is versioned like a small append-only log: writing a profile never mutates an
existing version, it creates a new one and moves a movable
label(defaultactive) to point atit — so a workbench created last month keeps a record of exactly which version it used, even after
the name has since been updated.
Honest resolution, never fabrication. Resolving a profile against a target cluster/scheduler
must classify the result as one of:
unchecked— no target given (e.g. a local/dev run); the raw ask is returned with no capabilityclaim.
verified— live discovery of the target confirms the requested GPU count/model.degraded— the coarse ask (GPU count/CPU/memory) is satisfiable, but a finer claim on theprofile (a specific accelerator model or MIG profile) cannot be confirmed from what the current
discovery probe reports for that target — resolution proceeds on the coarse ask and says exactly
which fields it could not confirm.
unresolvable— the ask exceeds the target's total capacity; resolution is refused outright.This mirrors the "never silently pretend a capability" convention already used elsewhere in the
platform's HPC/scheduler code (honest fallback on unsupported hardware features, honest capability
reporting on checkpoint/resume) — a hardware profile must never claim a piece of hardware exists
just because an operator once typed its name into a profile.
CLI surface
Folded into the existing hardware-placement command group as a
profilesubcommand (rather thana new top-level group), since that group already owns the neutral device-requirement vocabulary a
profile is a named, saved preset of:
list/show/resolveare read-only;setis a persisting write;deleteis a destructive,confirmation-gated removal.
Phased implementation plan
Phase 1 — Profile registry + CLI (no consumer wiring).
Add the profile record + versioning/label persistence, the pure resolution logic (including the
honest
unchecked/verified/degraded/unresolvableclassification against live clusterdiscovery data), the
exa hardware profile ...CLI, and CLI-surface tier classification (whichcommands are read-only vs. mutating, for any UI that introspects the CLI). Every mutation is
audited.
Acceptance: creating a profile and resolving it against a known cluster returns the correct
status for each of the four resolution states above (unit-tested, including against a cluster with
no discovery snapshot and a cluster whose capacity is insufficient); versioning round-trips (a
second
setcreates a new version, an old version stays retrievable by number, and the movablelabel points at the newest write); delete-without-version removes every version and its label,
delete-with-version removes only that one.
Phase 2 — Workbench integration.
Let workbench creation accept an optional hardware-profile reference; when given, unset resource
fields on the workbench are filled from the resolved profile (explicit flags on the same call
still win, so a profile is a default, not an override). The workbench record additively gains
columns for which profile (name + version) it was created from, so it stays inspectable after the
profile itself is later updated. Likely touched:
platform/cli/src/examlops/workbenches/__init__.py,platform/cli/src/examlops/cli/commands/workbench_cmd.py.Acceptance: creating a workbench with a profile reference produces the profile's CPU/memory
shape; overriding one flag on the same call overrides only that field; a profile whose
applicabilityexcludesworkbenchis rejected with a clear error, not silently accepted.Phase 3 — Training & serving integration.
Add an optional hardware-profile flag to the training-run and distributed-training-launch
commands, resolved into the same resource-ask shape those commands already build from raw flags.
Add an optional, additive hardware-profile reference to the per-model serving configuration, so a
deployed model's GPU/CPU/memory shape can also come from a named profile. Likely touched:
platform/cli/src/examlops/cli/commands/pipeline.py,platform/cli/src/examlops/cli/commands/distributed_cmd.py,pipelines/pipeline_generator/,serving/ray_serving/.Acceptance: a distributed-training launch driven by
--hardware-profile <name>produces theidentical launch command that specifying
--nodes/--gpus-per-nodedirectly would have produced,proving the profile is sugar over the existing resource path rather than a second execution path;
a model config referencing a profile deploys with the resolved GPU fraction/CPU applied exactly as
today's explicit per-model settings do.
Phase 4 — Dashboard + visibility surface.
A read-only dashboard view for listing/inspecting profiles (admin-gated create/delete), reusing
the same backend code paths the CLI uses rather than a parallel implementation, plus a
hardware-profile picker on the workbench/training/serving creation forms. Surface a profile's
current resolution status (verified/degraded/unresolvable) wherever it is actively in use, so a
degraded or unresolvable profile is visible operationally rather than silently stored.
Acceptance: the dashboard's profile list matches
exa hardware profile list --jsonexactly (nodrift between the two surfaces); a non-admin viewer can list/inspect but not create or delete
profiles.
Why this design
resolves into the platform's existing scheduler-neutral resource-ask and accelerator-placement
types; nothing downstream of resolution needs to learn a new shape.
actually get" stays answerable after the named profile is updated — the same pattern this
platform already uses for its prompt/versioning registry.
actually reported — consistent with how this platform already treats every other
capability-degradation boundary (scheduler feature support, checkpoint/resume capability,
fractional-GPU fallback).
Alternatives considered
the drift described above.
Rejected — this platform's compute substrate is Docker Compose plus bare-metal/HPC scheduler
adapters, not a server-side admission-controller pipeline; a profile here is a client-side named
template resolved before the call, not a server-side mutating hook.
the profile. Rejected — that type is created fresh per placement call and is not meant to be
versioned or referenced ahead of time; conflating the two would make "which version of this
profile did that workbench use" unanswerable.