Large-scale data processing with Apache Spark (PySpark) on Databricks. The project covers the core workflow of distributed data processing: loading raw data, transforming and cleaning it with Spark DataFrames, and running analytical queries at scale.
- Data ingestion from Excel sources into Spark DataFrames
- Data cleaning and transformation using distributed operations
- Aggregations and exploratory analysis with PySpark
- Project configured to run on Databricks (
databricks.yml)
| File | Description |
|---|---|
Lab 8.ipynb |
Main notebook with the Spark processing workflow |
Bases de datos principales PNC.xlsx |
Source dataset |
databricks.yml |
Databricks project configuration |
Python · Apache Spark · PySpark · Databricks · Jupyter Notebook
- Clone the repository:
git clone https://github.com/PabloSHerrera/Spark-Data-Processing.git- Import
Lab 8.ipynbinto a Databricks workspace (or run it locally with PySpark installed). - Upload the dataset and run all cells.
Developed as part of the Data Science course at Universidad del Valle de Guatemala.