Portfolio MLOps project for predicting infrastructure incidents from synthetic infrastructure metrics.
The project goal is to build a small, production-like MLOps platform:
generate synthetic infra dataset
→ train incident prediction model
→ log experiments in MLflow
→ serve model through FastAPI
→ deploy inference service to Kubernetes
→ later add Airflow, JupyterHub, KServe and GPU/CUDA practice
incident-prediction-platform/
├── airflow/
│ └── dags/
│ └── train_incident_model.py
├── artifacts/
├── data/
│ └── generate_dataset.py
├── helm/
│ └── incident-inference/
│ ├── Chart.yaml
│ ├── templates/
│ │ ├── _helpers.tpl
│ │ ├── deployment.yaml
│ │ └── service.yaml
│ └── values.yaml
├── k8s/
│ ├── airflow/
│ │ └── README.md
│ ├── jupyterhub/
│ │ └── README.md
│ ├── kserve/
│ │ └── README.md
│ ├── mlflow/
│ │ └── README.md
│ └── serving/
│ ├── deployment.yaml
│ └── service.yaml
├── README.md
├── serving/
│ ├── app.py
│ ├── Dockerfile
│ ├── model_loader.py
│ └── requirements.txt
└── training/
├── evaluate.py
├── requirements.txt
└── train.py
Minimum local stack:
Python 3.11
Docker
kubectl
kind
Helm
Current validated versions from the local workstation:
Python: 3.11.15
Docker: 27.3.1
kubectl client: v1.35.2
Kustomize: v5.7.1
Helm: v4.1.3
kind: v0.31.0
kind node image: kindest/node:v1.35.0
Important: do not use Python 3.14 as the main project runtime yet. Many ML packages may lag behind the newest Python release. Use Python 3.11 for this project.
Create and activate a virtual environment from the repository root:
cd ~/projects/MLOps/incident-prediction-platform
/opt/homebrew/bin/python3.11 -m venv .venv
source .venv/bin/activate
python --version
python -m pip install --upgrade pip setuptools wheelExpected result:
Python 3.11.x
Install training dependencies:
pip install -r training/requirements.txtIf the project-level dependencies are not split yet, install the initial working set manually:
pip install pandas scikit-learn mlflow fastapi uvicorn pydantic prometheus-client joblibOptionally freeze the local development environment:
pip freeze > requirements-dev.txtkind means Kubernetes in Docker. It runs Kubernetes nodes as local Docker containers.
For this project, kind is used to:
1. Test raw Kubernetes manifests from ./k8s
2. Test the Helm chart from ./helm/incident-inference
3. Deploy the FastAPI inference service locally
4. Validate Deployment, Service, probes, env and resources
5. Later deploy MLflow, Airflow, JupyterHub or KServe in Kubernetes
6. Avoid using real work clusters for experiments
For this project, kind is better than starting with a remote cluster. It gives a safe local Kubernetes environment and fits CI/CD-style validation well.
On macOS with Homebrew:
brew install kind
kind versionCreate the directory for local kind configuration:
mkdir -p kindCreate kind/cluster.yaml:
```bash
cat > kind/cluster.yaml <<'EOF'
kind: Cluster
apiVersion: kind.x-k8s.io/v1alpha4
name: mlops-local
nodes:
- role: control-plane
extraPortMappings:
- containerPort: 30080
hostPort: 8080
protocol: TCP
- containerPort: 30443
hostPort: 8443
protocol: TCP
- role: worker
EOF
```Or create the file manually with this content:
kind: Cluster
apiVersion: kind.x-k8s.io/v1alpha4
name: mlops-local
nodes:
- role: control-plane
extraPortMappings:
- containerPort: 30080
hostPort: 8080
protocol: TCP
- containerPort: 30443
hostPort: 8443
protocol: TCP
- role: workerCreate the cluster:
kind create cluster --config kind/cluster.yamlThe cluster intentionally has two nodes:
control-plane node: Kubernetes control plane
worker node: application workloads
This is still lightweight, but less toy-like than a single-node cluster.
Run this from the repository root with .venv activated:
echo "== Python =="
python3 --version || true
which python3 || true
echo "== Docker =="
docker version --format '{{.Server.Version}}'
docker ps
echo "== kubectl =="
kubectl version --client
kubectl config get-contexts
echo "== Helm =="
helm version
echo "== kind =="
kind version
kind get clustersExpected result after successful bootstrap:
== Python ==
Python 3.11.15
.../incident-prediction-platform/.venv/bin/python3
== Docker ==
27.3.1
== kubectl ==
Client Version: v1.35.2
CURRENT NAME CLUSTER AUTHINFO NAMESPACE
* kind-mlops-local kind-mlops-local kind-mlops-local
== Helm ==
version.BuildInfo{Version:"v4.1.3", ...}
== kind ==
kind v0.31.0 ...
mlops-local
kubectl get nodes -o wide
kubectl get ns
kubectl get pods -AExpected nodes:
NAME STATUS ROLES VERSION
mlops-local-control-plane Ready control-plane v1.35.0
mlops-local-worker Ready <none> v1.35.0
Expected namespaces:
default
kube-node-lease
kube-public
kube-system
local-path-storage
If kubectl version returns this error:
The connection to the server localhost:8080 was refused
it usually means that kubectl is installed, but no active Kubernetes cluster/context is configured.
Check contexts:
kubectl config get-contexts
kubectl config current-contextAfter creating the kind cluster, the current context should be:
kind-mlops-local
First local milestone:
1. Generate synthetic dataset
2. Train a model locally
3. Save model artifact
4. Add MLflow tracking after the first working training script
Run:
source .venv/bin/activate
python data/generate_dataset.py
python training/train.pyDo not start with Airflow, JupyterHub, KServe or CUDA before the basic training flow works. Otherwise the project will turn into YAML practice instead of MLOps practice.
data: synthetic dataset generator.training: model training and evaluation scripts.serving: FastAPI inference service.airflow: training DAG.k8s: raw Kubernetes manifests and component notes.helm: Helm chart for the inference service.artifacts: local generated artifacts, models and temporary outputs.
Result:
Synthetic dataset generated
Training script runs locally
Model artifact saved
Metrics are printed or stored
Result:
MLflow Tracking Server
PostgreSQL backend store
MinIO/S3 artifact store
training/train.py logs params, metrics and model artifacts
Result:
serving/app.py exposes:
GET /health
GET /ready
POST /predict
GET /metrics
Result:
Docker image built
Image loaded into kind
Inference service deployed through Helm
/predict works through port-forward or NodePort
Useful commands for the future serving check:
kubectl get pods
kubectl get svc
kubectl logs deploy/incident-inference
kubectl port-forward svc/incident-inference 8080:80
curl http://localhost:8080/health
curl http://localhost:8080/metrics