Skip to content

Latest commit

 

History

8 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Spark Data Processing

Large-scale data processing with Apache Spark (PySpark) on Databricks. The project covers the core workflow of distributed data processing: loading raw data, transforming and cleaning it with Spark DataFrames, and running analytical queries at scale.

Highlights

  • Data ingestion from Excel sources into Spark DataFrames
  • Data cleaning and transformation using distributed operations
  • Aggregations and exploratory analysis with PySpark
  • Project configured to run on Databricks (databricks.yml)

Repository Structure

File Description
Lab 8.ipynb Main notebook with the Spark processing workflow
Bases de datos principales PNC.xlsx Source dataset
databricks.yml Databricks project configuration

Tech Stack

Python · Apache Spark · PySpark · Databricks · Jupyter Notebook

How to Run

  1. Clone the repository:
   git clone https://github.com/PabloSHerrera/Spark-Data-Processing.git
  1. Import Lab 8.ipynb into a Databricks workspace (or run it locally with PySpark installed).
  2. Upload the dataset and run all cells.

Developed as part of the Data Science course at Universidad del Valle de Guatemala.

About

Introduction to Apache Spark and Databricks for large-scale data processing

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages