Data Engineer in training - building production-grade data pipelines and backend services with Python, Kafka, PySpark, Airflow, and AWS.
BSc Computer Science student at the University of West London (graduating 2027), based in London, UK.
📫 LinkedIn · kushagra7067@gmail.com
Languages: Python, SQL, Bash, Java Data Engineering: Kafka, Spark (PySpark, Structured Streaming), Airflow, Cassandra, PostgreSQL, Snowflake Backend: FastAPI, SQLAlchemy (async), REST API design, JWT auth Cloud & DevOps: AWS (Lambda, S3, Kinesis, Glue), Docker, Kubernetes, GitHub Actions CI/CD Testing: pytest, Testcontainers, Great Expectations, TDD
Real-time medallion ETL over live TfL Underground arrivals - Kafka to Spark Structured Streaming to Cassandra. Streams 4,200+ records per 30-second poll across 11 lines; watermarked deduplication removes 45% duplicates at the silver layer.
Async double-entry ledger REST API - FastAPI, PostgreSQL, Kafka. Deadlock-safe idempotent transfers, JWT auth with RBAC, 93% test coverage across 48 tests, CI/CD via GitHub Actions.
Payment data platform with real-time Kafka ingestion, PySpark batch and stream processing via Medallion architecture, Airflow orchestration, and a PostgreSQL serving layer.
End-to-end ELT pipeline for retail sales data - Airflow TaskFlow API, Docker, AWS S3 and Glue.
- Working towards AWS Cloud Practitioner certification
- Deepening Spark internals and stream processing patterns
- Open to data engineering internships and placement opportunities in London