Skip to content
View KushAgrawal1's full-sized avatar
🎯
Focusing
🎯
Focusing

Block or report KushAgrawal1

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
KushAgrawal1/README.md

Hi, I'm Kush Agrawal 👋

Data Engineer in training - building production-grade data pipelines and backend services with Python, Kafka, PySpark, Airflow, and AWS.

BSc Computer Science student at the University of West London (graduating 2027), based in London, UK.

📫 LinkedIn · kushagra7067@gmail.com


🔧 What I work with

Languages: Python, SQL, Bash, Java Data Engineering: Kafka, Spark (PySpark, Structured Streaming), Airflow, Cassandra, PostgreSQL, Snowflake Backend: FastAPI, SQLAlchemy (async), REST API design, JWT auth Cloud & DevOps: AWS (Lambda, S3, Kinesis, Glue), Docker, Kubernetes, GitHub Actions CI/CD Testing: pytest, Testcontainers, Great Expectations, TDD


🚀 Featured projects

Real-time medallion ETL over live TfL Underground arrivals - Kafka to Spark Structured Streaming to Cassandra. Streams 4,200+ records per 30-second poll across 11 lines; watermarked deduplication removes 45% duplicates at the silver layer.

Async double-entry ledger REST API - FastAPI, PostgreSQL, Kafka. Deadlock-safe idempotent transfers, JWT auth with RBAC, 93% test coverage across 48 tests, CI/CD via GitHub Actions.

Payment data platform with real-time Kafka ingestion, PySpark batch and stream processing via Medallion architecture, Airflow orchestration, and a PostgreSQL serving layer.

End-to-end ELT pipeline for retail sales data - Airflow TaskFlow API, Docker, AWS S3 and Glue.


🎯 Currently

  • Working towards AWS Cloud Practitioner certification
  • Deepening Spark internals and stream processing patterns
  • Open to data engineering internships and placement opportunities in London

Pinned Loading

  1. Real-Time-Ingestion-Fabric Real-Time-Ingestion-Fabric Public

    Streaming pipeline over live TfL Underground arrivals. Kafka to Spark Structured Streaming to Cassandra with bronze/silver/gold layers. Watermarked deduplication removes ~45% of ingested volume; va…

    Python 2

  2. ledger ledger Public

    Double-entry ledger REST API — FastAPI, PostgreSQL, Kafka, JWT auth, idempotent transfers

    Python

  3. retail_airflow retail_airflow Public

    End-to-end ELT pipeline for retail sales data orchestrated with Apache Airflow (TaskFlow API), containerised with Docker, and integrated with AWS S3 and Glue. Simulates multi-store transactions, tr…

    Python

  4. stock-market-realtime-pipeline-aws stock-market-realtime-pipeline-aws Public

    Real-time stock market analytics pipeline built on AWS — Kinesis, Lambda, S3, DynamoDB, Glue, Athena, and SNS

    Python

  5. payment-data-platform payment-data-platform Public

    Production-grade payment data platform with real-time Kafka ingestion, PySpark batch and stream processing via Medallion architecture (Bronze/Silver/Gold), Airflow orchestration, and a PostgreSQL d…

    Python