Heyy,
Welcome to my space!
I'm a Data Engineer with experience building and optimizing distributed data and AI workflows across the finance and medical domains.
-
Data Engineering: Building large-scale ingestion and transformation pipelines using ETL/ELT, CDC, batch, full-load, upsert, and incremental processing, along with workflow orchestration.
-
Data Quality and Governance: Implementing data validation, reconciliation, schema enforcement, automated quality checks, and governance practices to improve pipeline reliability and reduce data discrepancies.
-
AI and ML Workflows: Developing RAG and LLM-integrated applications involving document ingestion, embeddings, semantic retrieval, vector search, and retrieval-augmented generation. Also experienced with planning and deploying agentic workflows on AWS and building ML pipelines supporting analytics.
-
Monitoring and Reliability: Building monitoring and alerting workflows using CloudWatch, logging, lineage analysis, audit logs, and source-level validation to detect failures, investigate root causes, and maintain reliable production data pipelines.
-
Performance Optimization: Optimizing workloads across database, query, distributed job, and pipeline layers through table partitioning, data distribution strategies, SQL and join optimization, reduced data shuffling and scan overhead, and improved parallel execution through effective worker and resource allocation.
I enjoy working on problems involving scalable data systems, distributed processing, data reliability, workflow optimization, and AI-powered applications.
- You can find me hiking, cycling, exploring landscape photography, and trying out new food spots.
- 🤝 Open to collaboration: feel free to drop me a message, and let's create something meaningful together!
Thank you for stopping by!
