Building a minimal but real MLOps loop end to end.
This repo trains a machine-failure classifier on industrial sensor data and wires it into a full train → validate → serve → containerize → publish pipeline, gated end to end by CI. It's deliberately small in scope so the plumbing stays visible — the interesting part isn't the model, it's everything around it.
flowchart LR
subgraph Data
S3[("S3 remote\nvia DVC")]
A[("data/ai4i2020.csv\nAI4I 2020 dataset")]
S3 -- "dvc pull\n(skipped if cache-hit)" --> A
end
subgraph Training
B["train.py --model rf|gbm|xgb"]
B1["Drop leaky columns\n(TWF, HDF, PWF, OSF, RNF)"]
B2["Preprocess\n(OneHotEncode Type)"]
B3["rf: class_weight=balanced\ngbm: sample_weight\nxgb: scale_pos_weight"]
B4["Fixed decision threshold\n(0.30, same across models)"]
end
subgraph Artifacts
C1[("artifacts/model.pkl")]
C2[("artifacts/metrics.json")]
end
subgraph Tracking["Experiment tracking (local, optional)"]
M[("MLflow server\nkind + Postgres")]
end
subgraph Inference
D1["run_model.py\n(CLI predictor)"]
D2["app.py\n(Flask API)"]
end
subgraph Container["Docker"]
E1["Dockerfile\npython:3.12-slim"]
E2["RUN python train.py\n(trains at build time)"]
E3["gunicorn serves app.py\non :5000"]
end
subgraph CI["GitHub Actions CI"]
F["Quality gate\nrecall < 0.5 fails build"]
G["Upload artifacts\n(model.pkl + metrics.json)"]
H["Build & push image\n(main branch only)"]
end
R[("ghcr.io\npredictive-maintenance-mlops")]
A --> B --> B1 --> B2 --> B3 --> B4
B4 --> C1
B4 --> C2
B4 -. "log params/metrics/model\nbest-effort — skipped if unreachable" .-> M
C1 --> D1
C1 --> D2
C2 -. "decision_threshold" .-> D1
C2 -. "decision_threshold" .-> D2
B --> F --> G
F --> H
A -. "COPY . ." .-> E1 --> E2 --> E3
E1 -. "same Dockerfile" .-> H
H --> R
style S3 fill:#e8eaf6,stroke:#5c6bc0
style A fill:#e8eaf6,stroke:#5c6bc0
style C1 fill:#e0f2f1,stroke:#00897b
style C2 fill:#e0f2f1,stroke:#00897b
style D1 fill:#fff3e0,stroke:#fb8c00
style D2 fill:#fff3e0,stroke:#fb8c00
style E1 fill:#ede7f6,stroke:#7e57c2
style E2 fill:#ede7f6,stroke:#7e57c2
style E3 fill:#ede7f6,stroke:#7e57c2
style F fill:#fce4ec,stroke:#d81b60
style G fill:#fce4ec,stroke:#d81b60
style H fill:#fce4ec,stroke:#d81b60
style R fill:#e0f2f1,stroke:#00897b
style M fill:#fff9c4,stroke:#f9a825
train.py is the single source of truth: it produces both the model and the
decision threshold used everywhere downstream, so the CLI and the API never
silently disagree with each other or with the reported metrics.
Iris is 150 rows, perfectly balanced, 4 clean features — it exercises the plumbing but requires zero ML judgment. AI4I 2020 is ~10,000 rows of synthetic industrial sensor readings (temperature, torque, rotational speed, tool wear) with a binary machine-failure label at a ~3.4% positive rate. That forces real decisions instead of a dataset swap in name only:
| Problem | What happens if ignored | How this repo handles it |
|---|---|---|
| Leakage | Five sub-flags (TWF, HDF, PWF, OSF, RNF) directly encode the target — any one being 1 means Machine failure is 1. |
Dropped before training in train.py. |
| Class imbalance | At 3.4% positive, always predicting "no failure" scores ~97% accuracy while catching nothing. | class_weight="balanced"; metrics.json reports precision/recall/F1/ROC-AUC/PR-AUC, with accuracy kept for reference only. |
| Decision threshold | The default 0.5 cutoff gives 91% precision but only 46% recall — missing more than half the real failures. | train.py sweeps thresholds and picks 0.30, balancing to ~75% precision / ~71% recall. Saved in metrics.json, loaded by both run_model.py and app.py. |
| Silent regressions | A retrain could quietly ship a worse model. | CI fails the build if recall on the held-out set drops below 0.5. |
.github/workflows/ci.yaml # pull data -> train -> quality gate -> build & push image to ghcr
data/ai4i2020.csv.dvc # DVC pointer (md5 + size) — the CSV itself is not committed
docs/mlflow-setup.md # connecting train.py to a self-hosted MLflow server
train.py # trains the model (rf/gbm/xgb), writes artifacts/, logs to MLflow
run_model.py # CLI predictor
app.py # Flask API (/health, /predict)
Dockerfile # trains at build time, serves via gunicorn
requirements.txt
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
# dataset is DVC-tracked (not committed) — pull it from the S3 remote first;
# requires AWS credentials with read access to the dvc remote bucket
dvc pull
python train.py
# defaults to --model rf (RandomForest) — this is what CI and the Docker
# image always train. gbm and xgb are available for local comparison:
# python train.py --model gbm
# python train.py --model xgb
# Saved artifacts/model.pkl
# {
# "decision_threshold": 0.3,
# "accuracy": 0.982,
# "precision": 0.75,
# "recall": 0.7059,
# "f1": 0.7273,
# "roc_auc": 0.9585,
# "pr_auc": 0.7745,
# ...
# }
python run_model.py --type L --air-temp 302 --process-temp 311 \
--rpm 1350 --torque 65 --tool-wear 220
# Machine failure: YES (threshold=0.3)
# Failure probability: 0.8567
python app.py
# in another terminal:
curl -X POST http://127.0.0.1:5000/predict \
-H "Content-Type: application/json" \
-d '{"type":"L","air_temp":302,"process_temp":311,"rpm":1350,"torque":65,"tool_wear":220}'| Endpoint | Method | Body | Response |
|---|---|---|---|
/health |
GET | — | {"status": "ok"} |
/predict |
POST | {"type","air_temp","process_temp","rpm","torque","tool_wear"} |
{"machine_failure","failure_probability","decision_threshold"} |
data/ai4i2020.csv is tracked with DVC instead of git —
git only holds data/ai4i2020.csv.dvc, a pointer file with the dataset's
md5 hash and size. The actual bytes live in an S3 bucket
(.dvc/config), fetched on demand with dvc pull.
- Locally: needs AWS credentials (
aws configure/ SSO) with read access to the remote bucket. - In CI: no stored credentials at all — the workflow assumes an IAM
role via OpenID Connect (
aws-actions/configure-aws-credentials+permissions: id-token: write), so there's no long-lived AWS key sitting in GitHub Secrets. - Caching: the pulled CSV is cached with
actions/cache, keyed on the.dvcfile's hash. Unless the dataset actually changes, later CI runs skip the S3 pull entirely instead of re-downloading it on every push and for every Python version in the matrix.
Why not just commit the CSV? At 512KB it barely matters here — this is about practicing the pattern DVC exists for: keep git for code, keep a content-addressed remote for data, and never bloat repo history with binary blobs that grow every time the dataset changes.
train.py --model {rf,gbm,xgb} trains three different approaches to the
same imbalance problem, each logged to MLflow as a separate run — params,
metrics, and the fitted pipeline itself, comparable side by side instead
of eyeballed from printed JSON. See
docs/mlflow-setup.md for connecting train.py
to a self-hosted server (kind + Postgres backend store).
| Model | Imbalance handling | Precision | Recall | F1 | PR-AUC |
|---|---|---|---|---|---|
| RandomForest (default) | class_weight="balanced" |
0.75 | 0.71 | 0.73 | 0.77 |
| GradientBoosting | sample_weight (no native class_weight) |
0.35 | 0.87 | 0.50 | 0.69 |
| XGBoost | scale_pos_weight=28.5 |
0.61 | 0.81 | 0.70 | 0.82 |
No clean winner — RandomForest has the best F1, XGBoost the best PR-AUC.
All three trained at the same fixed 0.30 decision threshold so the
comparison is apples-to-apples: same cutoff, different model. RandomForest
stays the default (what CI and the Docker image train) since F1 is the
more defensible headline number here; --model xgb is the pick if
ranking quality (PR-AUC) matters more than the default threshold's exact
precision/recall tradeoff.
Tracking is best-effort: your kind cluster isn't reachable from GitHub's
runners, so CI trains normally without it — docs/mlflow-setup.md covers
why that's by design, not a gap.
data/ai4i2020.csv must already exist locally before building — it's
DVC-tracked, not committed, and COPY . . only sees what's actually on
disk. Run dvc pull first (same precondition CI has):
dvc pull
docker build -t predictive-maintenance-mlops .
docker run -p 5000:5000 predictive-maintenance-mlops
# in another terminal:
curl -X POST http://127.0.0.1:5000/predict \
-H "Content-Type: application/json" \
-d '{"type":"L","air_temp":302,"process_temp":311,"rpm":1350,"torque":65,"tool_wear":220}'The image trains the model at build time (RUN python train.py) so it
ships ready to serve — no separate artifact-download step. In a real
deployment you'd instead pull a pre-trained artifacts/model.pkl from a
registry (S3, Artifactory) and COPY it in, rather than retraining inside
the image build. The container serves via gunicorn instead of Flask's
dev server (python app.py), which is what production actually runs.
Every push to main builds this same Dockerfile in CI and publishes it to
GitHub Container Registry — no local build required:
docker pull ghcr.io/im-sathiyan-baskaran/predictive-maintenance-mlops:latest
docker run -p 5000:5000 ghcr.io/im-sathiyan-baskaran/predictive-maintenance-mlops:latestImages are tagged latest and with the short commit SHA (sha-xxxxxxx),
so a specific build is always pinnable instead of trusting a moving tag.
.github/workflows/ci.yaml runs on every push
and PR to main.
train_and_save_modeL — matrix across Python 3.10/3.11/3.12:
- Assume the S3 read role via OIDC
- Install dependencies
- Pull
data/ai4i2020.csvfrom the DVC remote (cached — skipped if the dataset hasn't changed since the last run) - Train the model
- Quality gate — asserts
recall >= 0.5onartifacts/metrics.json, failing the job (and therefore the image build downstream) if a retrain quietly regresses - Upload
artifacts/(model + metrics) as a build artifact per Python version
build_and_push_image — runs only on pushes to main (not PRs), after
the training job succeeds:
- Assume the S3 read role via OIDC and pull the dataset (same
cache-or-pull step as the training job —
RUN python train.pyinside the Docker build needs the CSV physically present in the build context, sinceCOPY . .only sees what's actually on disk, not what git tracks) - Log in to
ghcr.iousing the built-inGITHUB_TOKEN— no extra secret to manage - Build the
Dockerfile - Push it as
ghcr.io/im-sathiyan-baskaran/predictive-maintenance-mlops, taggedlatestandsha-<short-sha>
Gating the image push on the training job's matrix (needs:) means a
model that fails the recall gate never gets shipped in a container.
AI4I 2020 Predictive Maintenance Dataset, UCI Machine Learning Repository. S. Matzka, "Explainable Artificial Intelligence for Predictive Maintenance Applications," 2020. Synthetic data designed to reflect real industrial predictive-maintenance conditions.
SwapDone — see Experiment tracking above.RandomForestClassifierforGradientBoostingClassifieror XGBoost and compare PR-AUC.- Promote the winning MLflow run to the Model Registry and have
train.py(or a separateserve.py) load "Production" by name/stage instead of a localartifacts/model.pkl. - Log per-
Type(L/M/H) recall separately — failure dynamics differ by product variant. - Deploy the published
ghcr.ioimage behind ArgoCD instead of a manualdocker run. - Scan the image for vulnerabilities (Trivy) as a CI step before push, and
sign it with
cosignfor supply-chain provenance. - Add a
/metricsendpoint exposing prediction counts and probability distribution for Dynatrace/ELK-style monitoring in production.
feedback and PRs welcome.