Skip to content

Notebook-to-Pipeline Authoring: export tagged workbench notebook cells into the pipeline DSL #70

Description

@MSKazemi

Problem

Exploratory notebooks run in an on-demand, project-scoped workbench are a great place to try out a
model interactively — but today the only way that exploration becomes a versioned, reviewable,
CI-runnable pipeline is to hand-copy the relevant logic into a new file and hand-write it against
the pipeline-as-code DSL (@pipeline / dataset / train / evaluate / promote, compiled to a
validated intermediate representation and run through the existing orchestration + scheduler
stack). Nothing checks the copy is faithful, nothing tells the author which parts of the notebook
they forgot to carry over, and nothing validates the result until they remember to run the
compiler on it.

This proposes a small, deliberately scoped tool to close that gap: cell tagging plus static
extraction
— not a full visual pipeline editor (that would be a separate, materially larger
initiative and is explicitly out of scope here).

Proposed design

Cell-tag vocabulary

Standard Jupyter cell metadata (cell.metadata.tags, a plain list of strings — a documented
notebook format feature, not a custom extension) marks which cells matter for extraction. The tag
set mirrors the pipeline DSL's own step vocabulary directly, so there is exactly one vocabulary to
learn:

Tag Meaning
param Declares named literal values (dataset paths, thresholds, hyperparameters) other tagged cells can reference by name — a papermill-style "parameters cell" convention
dataset One dataset the pipeline reads
train The training step
evaluate The evaluation step
promote The promotion/lifecycle step
skip-export Explicit "don't extract this, and don't warn about it" marker

A tagged cell's source must be a single literal-valued assignment (a dict of plain constants, or a
bare reference to an earlier param) — the extractor parses cell source as data and never
imports, calls, or executes it
. This keeps extraction static and side-effect-free: reading a
notebook this way carries no more execution risk than reading a YAML file.

Extraction + validation flow

exa workbench export-pipeline <notebook.ipynb> [--out pipeline.py] [--yaml model.yaml] [--json]:

  1. Reads the notebook JSON, classifies cells by tag.
  2. Resolves param cells into a flat namespace.
  3. Parses each dataset/train/evaluate/promote cell's literal dict against that step kind's
    real parameter contract.
  4. Wires the steps into a dependency graph (train's datasets, evaluate's model, promote's
    metrics) using the notebook author's own chosen variable names.
  5. Emits a @pipeline-decorated Python file using the real DSL helpers.
  6. Immediately compiles the generated file through the existing pipeline compiler (the same
    in-process validation exa pipeline compile performs) — so an extraction that doesn't produce a
    valid pipeline fails at export time, not later at run time.
  7. Reports the compiled pipeline's name/hash/step count, and — critically — every cell it
    dropped
    , so an author who forgot a tag sees it named instead of silently losing content.
    Untagged cells are always reported unless explicitly marked skip-export; dropping is never
    silent.

Explicit scope statement

This is one-directional: notebook → pipeline file. Editing the generated pipeline file does
not flow back into the notebook, and re-running the export overwrites the file with a fresh
extraction rather than merging. The generated file is a starting point handed to the normal
pipeline-as-code workflow (review it, edit it, version it) — not a live view of the notebook. No
two-way sync is implied or planned.

hpo/custom_python step kinds are intentionally not extractable yet — they exist in the pipeline
IR but have no execution path today, so extracting them would only produce a pipeline that
compiles and is then refused at run time.

CLI surface

  • New subcommand under the existing workbench command group (it operates on a workbench-produced
    artifact, not a general-purpose notebook tool).
  • Admin-tier: it writes a file and executes the generated pipeline code during its own validation
    step, consistent with how the existing pipeline-compile command is tiered.
  • Phased:
    1. Tag vocabulary + static extractor + compile-validation round-trip, writing to an explicit
      --out path.
    2. --out defaults inside the owning workbench's own project storage, resolved via
      --project/--workbench options.
    3. A dashboard "export as pipeline" button on a workbench's notebook view, reusing the existing
      generic command-console form for admin CLI commands rather than new backend logic.

Acceptance criteria

  • A notebook with correctly tagged dataset/train/evaluate/promote cells exports a
    pipeline file that compiles successfully through the existing pipeline compiler.
  • A param-tagged cell's declared values are correctly resolved into the values later tagged
    cells reference by name.
  • Every untagged code cell is named in the export report (text and --json modes); a
    skip-export-tagged cell is dropped without a warning.
  • More than one train/evaluate/promote cell, an unresolved dataset/model/metrics
    reference, or a tagged cell whose source isn't a literal-valued assignment fails the export
    with a clear, cell-specific error — not a silent partial export.
  • The command is admin-tier in the CLI's command-surface catalog and its output path is
    contained within the CLI's sandboxed workspace root, consistent with every other
    file-producing command.
  • Documentation states plainly, in the command's own help text, that export is one-directional
    and that editing the generated file does not update the source notebook.

Modules likely touched

  • A new, pure notebook-parsing/DSL-emission module alongside the existing pipeline DSL package
    (no changes to the DSL, its intermediate representation, or its validator — the extractor is a
    producer for machinery that already exists).
  • The workbench CLI command group (new export-pipeline subcommand).
  • The CLI's command-surface/tiering catalog (one new entry).
  • Project-storage helpers already used by other workbench/project commands, for Phase 2's default
    output path.
  • The dashboard's existing generic admin-command console, for Phase 3 (no bespoke backend route
    expected).

Non-goals

A full visual, drag-and-drop pipeline/DAG editor is not proposed here. It would need a
persistent graph-editing UI and a different validation loop from a CLI extractor, and is a
separate, materially larger initiative with no clear priority signal yet. This issue is
deliberately scoped to a semi-automated cell-tagging + static-extraction tool only.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions