Skip to content

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

Incident Prediction Platform

Portfolio MLOps project for predicting infrastructure incidents from synthetic infrastructure metrics.

The project goal is to build a small, production-like MLOps platform:

generate synthetic infra dataset
→ train incident prediction model
→ log experiments in MLflow
→ serve model through FastAPI
→ deploy inference service to Kubernetes
→ later add Airflow, JupyterHub, KServe and GPU/CUDA practice

Project structure

incident-prediction-platform/
├── airflow/
│   └── dags/
│       └── train_incident_model.py
├── artifacts/
├── data/
│   └── generate_dataset.py
├── helm/
│   └── incident-inference/
│       ├── Chart.yaml
│       ├── templates/
│       │   ├── _helpers.tpl
│       │   ├── deployment.yaml
│       │   └── service.yaml
│       └── values.yaml
├── k8s/
│   ├── airflow/
│   │   └── README.md
│   ├── jupyterhub/
│   │   └── README.md
│   ├── kserve/
│   │   └── README.md
│   ├── mlflow/
│   │   └── README.md
│   └── serving/
│       ├── deployment.yaml
│       └── service.yaml
├── README.md
├── serving/
│   ├── app.py
│   ├── Dockerfile
│   ├── model_loader.py
│   └── requirements.txt
└── training/
    ├── evaluate.py
    ├── requirements.txt
    └── train.py

Required local tools

Minimum local stack:

Python 3.11
Docker
kubectl
kind
Helm

Current validated versions from the local workstation:

Python: 3.11.15
Docker: 27.3.1
kubectl client: v1.35.2
Kustomize: v5.7.1
Helm: v4.1.3
kind: v0.31.0
kind node image: kindest/node:v1.35.0

Important: do not use Python 3.14 as the main project runtime yet. Many ML packages may lag behind the newest Python release. Use Python 3.11 for this project.

Python environment

Create and activate a virtual environment from the repository root:

cd ~/projects/MLOps/incident-prediction-platform

/opt/homebrew/bin/python3.11 -m venv .venv
source .venv/bin/activate

python --version
python -m pip install --upgrade pip setuptools wheel

Expected result:

Python 3.11.x

Install training dependencies:

pip install -r training/requirements.txt

If the project-level dependencies are not split yet, install the initial working set manually:

pip install pandas scikit-learn mlflow fastapi uvicorn pydantic prometheus-client joblib

Optionally freeze the local development environment:

pip freeze > requirements-dev.txt

Why kind is used

kind means Kubernetes in Docker. It runs Kubernetes nodes as local Docker containers.

For this project, kind is used to:

1. Test raw Kubernetes manifests from ./k8s
2. Test the Helm chart from ./helm/incident-inference
3. Deploy the FastAPI inference service locally
4. Validate Deployment, Service, probes, env and resources
5. Later deploy MLflow, Airflow, JupyterHub or KServe in Kubernetes
6. Avoid using real work clusters for experiments

For this project, kind is better than starting with a remote cluster. It gives a safe local Kubernetes environment and fits CI/CD-style validation well.

Install kind

On macOS with Homebrew:

brew install kind
kind version

Create a local kind cluster

Create the directory for local kind configuration:

mkdir -p kind

Create kind/cluster.yaml:

```bash
cat > kind/cluster.yaml <<'EOF'
kind: Cluster
apiVersion: kind.x-k8s.io/v1alpha4

name: mlops-local

nodes:
  - role: control-plane
    extraPortMappings:
      - containerPort: 30080
        hostPort: 8080
        protocol: TCP
      - containerPort: 30443
        hostPort: 8443
        protocol: TCP
  - role: worker
EOF
```

Or create the file manually with this content:

kind: Cluster
apiVersion: kind.x-k8s.io/v1alpha4

name: mlops-local

nodes:
  - role: control-plane
    extraPortMappings:
      - containerPort: 30080
        hostPort: 8080
        protocol: TCP
      - containerPort: 30443
        hostPort: 8443
        protocol: TCP
  - role: worker

Create the cluster:

kind create cluster --config kind/cluster.yaml

The cluster intentionally has two nodes:

control-plane node: Kubernetes control plane
worker node: application workloads

This is still lightweight, but less toy-like than a single-node cluster.

Validate local tools

Run this from the repository root with .venv activated:

echo "== Python =="
python3 --version || true
which python3 || true

echo "== Docker =="
docker version --format '{{.Server.Version}}'
docker ps

echo "== kubectl =="
kubectl version --client
kubectl config get-contexts

echo "== Helm =="
helm version

echo "== kind =="
kind version
kind get clusters

Expected result after successful bootstrap:

== Python ==
Python 3.11.15
.../incident-prediction-platform/.venv/bin/python3

== Docker ==
27.3.1

== kubectl ==
Client Version: v1.35.2
CURRENT   NAME               CLUSTER            AUTHINFO           NAMESPACE
*         kind-mlops-local   kind-mlops-local   kind-mlops-local

== Helm ==
version.BuildInfo{Version:"v4.1.3", ...}

== kind ==
kind v0.31.0 ...
mlops-local

Validate the Kubernetes cluster

kubectl get nodes -o wide
kubectl get ns
kubectl get pods -A

Expected nodes:

NAME                        STATUS   ROLES           VERSION
mlops-local-control-plane   Ready    control-plane   v1.35.0
mlops-local-worker          Ready    <none>          v1.35.0

Expected namespaces:

default
kube-node-lease
kube-public
kube-system
local-path-storage

If kubectl version returns this error:

The connection to the server localhost:8080 was refused

it usually means that kubectl is installed, but no active Kubernetes cluster/context is configured.

Check contexts:

kubectl config get-contexts
kubectl config current-context

After creating the kind cluster, the current context should be:

kind-mlops-local

Local ML flow

First local milestone:

1. Generate synthetic dataset
2. Train a model locally
3. Save model artifact
4. Add MLflow tracking after the first working training script

Run:

source .venv/bin/activate

python data/generate_dataset.py
python training/train.py

Do not start with Airflow, JupyterHub, KServe or CUDA before the basic training flow works. Otherwise the project will turn into YAML practice instead of MLOps practice.

Components

  • data: synthetic dataset generator.
  • training: model training and evaluation scripts.
  • serving: FastAPI inference service.
  • airflow: training DAG.
  • k8s: raw Kubernetes manifests and component notes.
  • helm: Helm chart for the inference service.
  • artifacts: local generated artifacts, models and temporary outputs.

Next milestones

Milestone 1: local training

Result:

Synthetic dataset generated
Training script runs locally
Model artifact saved
Metrics are printed or stored

Milestone 2: MLflow

Result:

MLflow Tracking Server
PostgreSQL backend store
MinIO/S3 artifact store
training/train.py logs params, metrics and model artifacts

Milestone 3: FastAPI serving

Result:

serving/app.py exposes:
GET  /health
GET  /ready
POST /predict
GET  /metrics

Milestone 4: Kubernetes deployment

Result:

Docker image built
Image loaded into kind
Inference service deployed through Helm
/predict works through port-forward or NodePort

Useful commands for the future serving check:

kubectl get pods
kubectl get svc
kubectl logs deploy/incident-inference
kubectl port-forward svc/incident-inference 8080:80
curl http://localhost:8080/health
curl http://localhost:8080/metrics

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages