Data Engineer specializing in cloud-native pipelines and lakehouse architecture. AWS & Huawei certified, building production systems on Databricks, Spark, and Airflow, with hands-on work extending into retrieval infrastructure for AI applications.
I treat data infrastructure as a systems problem: correctness before performance, observability from day one, and hard boundaries between pipeline layers. A pipeline that runs once is not the same as one that runs reliably at scale.
My work spans batch ETL on Databricks, streaming with Kafka and Flink, SQL Server warehouses built on stored-procedure layers, and production RAG APIs with offline evaluation harnesses. The same rigor applies everywhere.
- Medallion Architecture pipelines on Databricks and Delta Lake
- Cloud-native data platforms on AWS: Glue, Redshift, EMR, Athena
- DE for AI workflows: vector indexing, hybrid retrieval, and retrieval evaluation
- ETL orchestration with Apache Airflow
| Certification | Issuer | Valid Until |
|---|---|---|
| AWS Certified Cloud Practitioner (CLF-C02) | Amazon Web Services | May 2029 |
| HCIA Big Data Associate | Huawei Technologies | March 2029 |
Email · LinkedIn · Portfolio · Resume
Building data systems that are reliable by design, maintainable in practice, and useful to the business.

