Skip to content
View garvvvit's full-sized avatar
🎯
Focusing
🎯
Focusing

Block or report garvvvit

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
garvvvit/README.md

Hi, I'm Garvit Vishwkarma πŸ‘‹

Data Analyst | Aspiring Data Engineer

I am an early-career data professional with hands-on experience in ServiceNow data analytics, SQL-based reporting, ETL workflow development, and data modeling. I am currently focused on building practical Data Engineering projects using Python, SQL, Apache Airflow, MySQL, Docker, Kafka, PySpark, Snowflake, and Databricks.

I enjoy working on real-world datasets, designing data pipelines, creating analytical data models, and transforming raw data into structured, business-ready insights.


πŸš€ What I'm Currently Working On

  • Building end-to-end Data Engineering projects
  • Designing ETL pipelines and data warehouse models
  • Practicing SQL, PySpark, Apache Airflow, and Kafka
  • Learning modern lakehouse workflows using Databricks and Delta Lake
  • Improving GitHub project documentation for job-ready portfolios

πŸ› οΈ Tech Stack

Programming & Analytics:
Python, SQL, Pandas, NumPy

Data Engineering:
ETL Pipelines, Apache Airflow, Kafka, Data Modeling, Star Schema, Data Quality Checks

Databases & Warehouses:
MySQL, PostgreSQL, MongoDB, Snowflake

Big Data & Lakehouse Tools:
PySpark, Databricks, Delta Lake, Docker

Visualization & Reporting:
Power BI, Matplotlib, SQL Reporting

Tools:
Git, GitHub, VS Code, ServiceNow


πŸ“Œ Featured Projects

1. IPL Real-Time Data Engineering Pipeline

End-to-end IPL data engineering pipeline using real Cricsheet ball-by-ball cricket data, Kafka, PySpark, Databricks, Delta Lake, and Spark SQL.

Key Highlights:

  • Ingested real IPL ball-by-ball JSON data from Cricsheet
  • Parsed raw JSON files into structured delivery-level event records
  • Implemented Kafka producer to stream ball-by-ball events into Kafka topics
  • Verified real-time event flow using Kafka console consumer
  • Built Bronze, Silver, and Gold lakehouse layers in Databricks using PySpark
  • Applied data quality checks including null validation, duplicate removal, run validation, and team consistency checks
  • Created Gold analytics tables for team, batter, bowler, venue, and match-level analysis
  • Used Spark SQL to generate final analytical insights

Tech Stack:
Python, Kafka, Docker, PySpark, Databricks, Delta Lake, Spark SQL


2. Retail Data Workflow Automation System

End-to-end Brazilian e-commerce ETL pipeline using Python, Pandas, Apache Airflow, MySQL, Docker, and SQL.

Key Highlights:

  • Processed real Kaggle Olist Brazilian e-commerce data
  • Automated ETL workflow using Apache Airflow
  • Created fact and dimension tables using star schema modeling
  • Loaded transformed data into a MySQL warehouse
  • Added data quality checks and SQL analytical queries
  • Designed the project for data warehouse and reporting use cases

Tech Stack:
Python, Pandas, Apache Airflow, MySQL, Docker, SQL, Data Modeling


πŸ“š Currently Learning

  • Advanced SQL
  • PySpark
  • Apache Airflow
  • Kafka
  • Snowflake
  • Databricks
  • Delta Lake
  • Data Warehouse Design
  • Cloud Data Engineering Concepts

🎯 Career Focus

I am actively building a strong Data Engineering portfolio by working on practical projects that cover:

  • Data ingestion
  • ETL/ELT pipelines
  • Workflow orchestration
  • Data modeling
  • Data quality validation
  • Batch and streaming pipelines
  • Lakehouse architecture
  • SQL-based analytics

πŸ“« Connect With Me

LinkedIn: www.linkedin.com/in/garvvvit
GitHub: github.com/garvvvit
Email: garvitdhiman2002@gmail.com

Pinned Loading

  1. retail-data-workflow-automation-system retail-data-workflow-automation-system Public

    End-to-end Brazilian e-commerce ETL pipeline using Python, Pandas, Apache Airflow, MySQL, Docker, Snowflake, and Databricks concepts.

    Python