Problem
Exploratory notebooks run in an on-demand, project-scoped workbench are a great place to try out a
model interactively — but today the only way that exploration becomes a versioned, reviewable,
CI-runnable pipeline is to hand-copy the relevant logic into a new file and hand-write it against
the pipeline-as-code DSL (@pipeline / dataset / train / evaluate / promote, compiled to a
validated intermediate representation and run through the existing orchestration + scheduler
stack). Nothing checks the copy is faithful, nothing tells the author which parts of the notebook
they forgot to carry over, and nothing validates the result until they remember to run the
compiler on it.
This proposes a small, deliberately scoped tool to close that gap: cell tagging plus static
extraction — not a full visual pipeline editor (that would be a separate, materially larger
initiative and is explicitly out of scope here).
Proposed design
Cell-tag vocabulary
Standard Jupyter cell metadata (cell.metadata.tags, a plain list of strings — a documented
notebook format feature, not a custom extension) marks which cells matter for extraction. The tag
set mirrors the pipeline DSL's own step vocabulary directly, so there is exactly one vocabulary to
learn:
| Tag |
Meaning |
param |
Declares named literal values (dataset paths, thresholds, hyperparameters) other tagged cells can reference by name — a papermill-style "parameters cell" convention |
dataset |
One dataset the pipeline reads |
train |
The training step |
evaluate |
The evaluation step |
promote |
The promotion/lifecycle step |
skip-export |
Explicit "don't extract this, and don't warn about it" marker |
A tagged cell's source must be a single literal-valued assignment (a dict of plain constants, or a
bare reference to an earlier param) — the extractor parses cell source as data and never
imports, calls, or executes it. This keeps extraction static and side-effect-free: reading a
notebook this way carries no more execution risk than reading a YAML file.
Extraction + validation flow
exa workbench export-pipeline <notebook.ipynb> [--out pipeline.py] [--yaml model.yaml] [--json]:
- Reads the notebook JSON, classifies cells by tag.
- Resolves
param cells into a flat namespace.
- Parses each
dataset/train/evaluate/promote cell's literal dict against that step kind's
real parameter contract.
- Wires the steps into a dependency graph (train's datasets, evaluate's model, promote's
metrics) using the notebook author's own chosen variable names.
- Emits a
@pipeline-decorated Python file using the real DSL helpers.
- Immediately compiles the generated file through the existing pipeline compiler (the same
in-process validation exa pipeline compile performs) — so an extraction that doesn't produce a
valid pipeline fails at export time, not later at run time.
- Reports the compiled pipeline's name/hash/step count, and — critically — every cell it
dropped, so an author who forgot a tag sees it named instead of silently losing content.
Untagged cells are always reported unless explicitly marked skip-export; dropping is never
silent.
Explicit scope statement
This is one-directional: notebook → pipeline file. Editing the generated pipeline file does
not flow back into the notebook, and re-running the export overwrites the file with a fresh
extraction rather than merging. The generated file is a starting point handed to the normal
pipeline-as-code workflow (review it, edit it, version it) — not a live view of the notebook. No
two-way sync is implied or planned.
hpo/custom_python step kinds are intentionally not extractable yet — they exist in the pipeline
IR but have no execution path today, so extracting them would only produce a pipeline that
compiles and is then refused at run time.
CLI surface
- New subcommand under the existing
workbench command group (it operates on a workbench-produced
artifact, not a general-purpose notebook tool).
- Admin-tier: it writes a file and executes the generated pipeline code during its own validation
step, consistent with how the existing pipeline-compile command is tiered.
- Phased:
- Tag vocabulary + static extractor + compile-validation round-trip, writing to an explicit
--out path.
--out defaults inside the owning workbench's own project storage, resolved via
--project/--workbench options.
- A dashboard "export as pipeline" button on a workbench's notebook view, reusing the existing
generic command-console form for admin CLI commands rather than new backend logic.
Acceptance criteria
Modules likely touched
- A new, pure notebook-parsing/DSL-emission module alongside the existing pipeline DSL package
(no changes to the DSL, its intermediate representation, or its validator — the extractor is a
producer for machinery that already exists).
- The
workbench CLI command group (new export-pipeline subcommand).
- The CLI's command-surface/tiering catalog (one new entry).
- Project-storage helpers already used by other workbench/project commands, for Phase 2's default
output path.
- The dashboard's existing generic admin-command console, for Phase 3 (no bespoke backend route
expected).
Non-goals
A full visual, drag-and-drop pipeline/DAG editor is not proposed here. It would need a
persistent graph-editing UI and a different validation loop from a CLI extractor, and is a
separate, materially larger initiative with no clear priority signal yet. This issue is
deliberately scoped to a semi-automated cell-tagging + static-extraction tool only.
Problem
Exploratory notebooks run in an on-demand, project-scoped workbench are a great place to try out a
model interactively — but today the only way that exploration becomes a versioned, reviewable,
CI-runnable pipeline is to hand-copy the relevant logic into a new file and hand-write it against
the pipeline-as-code DSL (
@pipeline/dataset/train/evaluate/promote, compiled to avalidated intermediate representation and run through the existing orchestration + scheduler
stack). Nothing checks the copy is faithful, nothing tells the author which parts of the notebook
they forgot to carry over, and nothing validates the result until they remember to run the
compiler on it.
This proposes a small, deliberately scoped tool to close that gap: cell tagging plus static
extraction — not a full visual pipeline editor (that would be a separate, materially larger
initiative and is explicitly out of scope here).
Proposed design
Cell-tag vocabulary
Standard Jupyter cell metadata (
cell.metadata.tags, a plain list of strings — a documentednotebook format feature, not a custom extension) marks which cells matter for extraction. The tag
set mirrors the pipeline DSL's own step vocabulary directly, so there is exactly one vocabulary to
learn:
paramdatasettrainevaluatepromoteskip-exportA tagged cell's source must be a single literal-valued assignment (a dict of plain constants, or a
bare reference to an earlier
param) — the extractor parses cell source as data and neverimports, calls, or executes it. This keeps extraction static and side-effect-free: reading a
notebook this way carries no more execution risk than reading a YAML file.
Extraction + validation flow
exa workbench export-pipeline <notebook.ipynb> [--out pipeline.py] [--yaml model.yaml] [--json]:paramcells into a flat namespace.dataset/train/evaluate/promotecell's literal dict against that step kind'sreal parameter contract.
metrics) using the notebook author's own chosen variable names.
@pipeline-decorated Python file using the real DSL helpers.in-process validation
exa pipeline compileperforms) — so an extraction that doesn't produce avalid pipeline fails at export time, not later at run time.
dropped, so an author who forgot a tag sees it named instead of silently losing content.
Untagged cells are always reported unless explicitly marked
skip-export; dropping is neversilent.
Explicit scope statement
This is one-directional: notebook → pipeline file. Editing the generated pipeline file does
not flow back into the notebook, and re-running the export overwrites the file with a fresh
extraction rather than merging. The generated file is a starting point handed to the normal
pipeline-as-code workflow (review it, edit it, version it) — not a live view of the notebook. No
two-way sync is implied or planned.
hpo/custom_pythonstep kinds are intentionally not extractable yet — they exist in the pipelineIR but have no execution path today, so extracting them would only produce a pipeline that
compiles and is then refused at run time.
CLI surface
workbenchcommand group (it operates on a workbench-producedartifact, not a general-purpose notebook tool).
step, consistent with how the existing pipeline-compile command is tiered.
--outpath.--outdefaults inside the owning workbench's own project storage, resolved via--project/--workbenchoptions.generic command-console form for admin CLI commands rather than new backend logic.
Acceptance criteria
dataset/train/evaluate/promotecells exports apipeline file that compiles successfully through the existing pipeline compiler.
param-tagged cell's declared values are correctly resolved into the values later taggedcells reference by name.
--jsonmodes); askip-export-tagged cell is dropped without a warning.train/evaluate/promotecell, an unresolved dataset/model/metricsreference, or a tagged cell whose source isn't a literal-valued assignment fails the export
with a clear, cell-specific error — not a silent partial export.
contained within the CLI's sandboxed workspace root, consistent with every other
file-producing command.
and that editing the generated file does not update the source notebook.
Modules likely touched
(no changes to the DSL, its intermediate representation, or its validator — the extractor is a
producer for machinery that already exists).
workbenchCLI command group (newexport-pipelinesubcommand).output path.
expected).
Non-goals
A full visual, drag-and-drop pipeline/DAG editor is not proposed here. It would need a
persistent graph-editing UI and a different validation loop from a CLI extractor, and is a
separate, materially larger initiative with no clear priority signal yet. This issue is
deliberately scoped to a semi-automated cell-tagging + static-extraction tool only.