Skip to content

About

Description: Machine learning pipelines with Apache Spark — classification and regression models

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

7 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Spark ML Pipelines

Machine learning pipelines built with Apache Spark (PySpark MLlib), covering classification and regression models. The project runs in a containerized environment with Docker, making it fully reproducible on any machine.

Highlights

  • End-to-end ML workflows using Spark's Pipeline API: feature engineering, model training, and evaluation
  • Classification and regression models trained on distributed data structures
  • Reproducible setup with Docker and Docker Compose, with no local Spark installation required

Repository Structure

Path Description
notebooks/ Jupyter notebooks with the ML pipelines and analysis
working_dir/ Data and working files
dockerfile Container image definition
docker-compose.yml Environment orchestration

Tech Stack

Python · Apache Spark · PySpark MLlib · Docker · Jupyter Notebook

How to Run

  1. Clone the repository:
   git clone https://github.com/PabloSHerrera/Spark-ML-Pipelines.git
   cd Spark-ML-Pipelines
  1. Start the environment:
   docker-compose up
  1. Open Jupyter in your browser using the URL shown in the terminal, then run the notebooks in notebooks/.

Developed as part of the Data Science course at Universidad del Valle de Guatemala.

About

Description: Machine learning pipelines with Apache Spark — classification and regression models

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages