Machine learning pipelines built with Apache Spark (PySpark MLlib), covering classification and regression models. The project runs in a containerized environment with Docker, making it fully reproducible on any machine.
- End-to-end ML workflows using Spark's
PipelineAPI: feature engineering, model training, and evaluation - Classification and regression models trained on distributed data structures
- Reproducible setup with Docker and Docker Compose, with no local Spark installation required
| Path | Description |
|---|---|
notebooks/ |
Jupyter notebooks with the ML pipelines and analysis |
working_dir/ |
Data and working files |
dockerfile |
Container image definition |
docker-compose.yml |
Environment orchestration |
Python · Apache Spark · PySpark MLlib · Docker · Jupyter Notebook
- Clone the repository:
git clone https://github.com/PabloSHerrera/Spark-ML-Pipelines.git
cd Spark-ML-Pipelines- Start the environment:
docker-compose up- Open Jupyter in your browser using the URL shown in the terminal, then run the notebooks in
notebooks/.
Developed as part of the Data Science course at Universidad del Valle de Guatemala.