I'm an AI/ML Engineer and Data Engineer with a Mathematics & Computing background from RGIPT, currently building production ETL pipelines on AWS at Algo8.AI. My work sits at the intersection of software engineering rigor and applied machine learning β I don't just prototype models in notebooks, I design the pipelines, feature stores, and infrastructure that get them into production.
Across three internships and a current full-time-track role, I've shipped:
- π Anomaly detection systems β Autoencoder-based pipeline hitting 95% precision across 31,505 enterprise access logs
- πΌοΈ Multimodal retrieval β CLIP + BLIP-2 + ChromaDB retrieval-augmented generation, 90% retrieval accuracy
- βοΈ Large-scale data engineering β PySpark feature pipelines and event-driven MLOps on AWS (S3, EC2, Lambda)
- π§© Computer vision β Gabor filter + Multi-Otsu segmentation, CNN benchmarking up to 99.31% validation accuracy
I approach engineering with a product mindset: performance, security, and maintainability aren't afterthoughts β they're part of the spec from day one.
| Domain | Proficiency | Details |
|---|---|---|
| Anomaly Detection | π£π£π£π£π£ | Autoencoders, unsupervised deep learning, threshold tuning β 95% precision on 31K+ logs |
| Multimodal Retrieval / RAG | π£π£π£π£βͺ | CLIP, BLIP-2, ChromaDB vector search, 90% retrieval accuracy |
| Feature Engineering & ETL | π£π£π£π£π£ | PySpark batch pipelines, AWS S3/EC2/Lambda, production data quality |
| Computer Vision | π£π£π£π£βͺ | Gabor filters, Multi-Otsu segmentation, CNN transfer learning (EfficientNetB0, MobileNetV2) |
| LLM Applications | π£π£π£βͺβͺ | GroqCloud (Llama3-70B) inference pipelines, automated report generation |
| MLOps | π£π£π£βͺβͺ | Event-driven pipelines, AWS Lambda triggers, API-based validation via Postman |
π Student Performance Analytics Engine with LLM Feedback Generation
End-to-end pipeline ingesting raw student performance data and auto-generating personalized, subject-wise PDF feedback reports using LLM analysis via GroqCloud β including automated visualizations and batch ZIP export. Built to eliminate manual report writing for educators at scale.
| Aspect | Detail |
|---|---|
| Stack | Python, GroqCloud (Llama3-70B-8192), Hugging Face, PDF generation |
| Scale | Multi-subject, multi-cohort extensible design |
| Performance | Automated batch report generation with ZIP export |
| Security | API-key managed LLM access, no PII persistence |
| Impact | Eliminates manual report authoring for educators |
| Repository | GitHub |
πΌοΈ Multi-Modal Retrieval and Generation System (Vision + Text)
Retrieval-augmented generation pipeline combining vision encoders (CLIP, BLIP-2) with vector similarity search (ChromaDB), optimized for high-throughput embedding retrieval.
| Aspect | Detail |
|---|---|
| Stack | CLIP, BLIP-2, GPT-2, ChromaDB, Hugging Face, Python |
| Scale | Vector-indexed multimodal corpus |
| Performance | 90% retrieval accuracy, 30% throughput improvement over baseline indexing |
| Security | Local embedding store, no external data leakage |
| Impact | Reusable RAG framework for vision + text retrieval tasks |
| Repository | GitHub |
π CNN Architecture Benchmarking β EfficientNetB0 vs MobileNetV2 vs Baseline
Comparative benchmarking of three CNN architectures on CIFAR-10, evaluating accuracy-vs-compute tradeoffs for transfer learning versus training from scratch.
| Aspect | Detail |
|---|---|
| Stack | PyTorch, EfficientNetB0, MobileNetV2, CIFAR-10 |
| Scale | 3 architecture variants, full hyperparameter sweep |
| Performance | 99.31% validation accuracy (EfficientNetB0 transfer learning) |
| Security | N/A β research/benchmarking project |
| Impact | Demonstrated high accuracy at low compute cost for deployment-constrained environments |
| Repository | GitHub |
Feb 2026 β Jun 2026 Β· Remote
Architected production-grade ETL pipelines migrating and transforming data from EC2 to S3, ensuring high data quality and low latency for downstream analytical models.
- Built large-scale feature engineering and batch processing workflows using PySpark
- Enabled event-driven MLOps workflows using AWS Lambda and EC2
- Streamlined API-based data extraction and validation using Postman, reducing pipeline debugging turnaround time
AWS PySpark ETL Lambda MLOps
May 2025 β Aug 2025 Β· Bengaluru
Designed and validated an unsupervised deep learning pipeline for enterprise security anomaly detection.
- Reduced anomaly verification time by 40% by automating analysis of 31,505 access logs
- Achieved 95% precision on security anomaly classification across 1,000+ critical incidents
- Built AccessAI, an end-to-end pipeline flagging abnormal user behavior patterns
Autoencoders Anomaly Detection Deep Learning Python
May 2024 β Jul 2024 Β· Bengaluru
Enhanced structural feature extraction on satellite imagery and contributed to computer vision annotation pipelines.
- Applied and tuned Gabor filter parameters combined with Multi-Otsu segmentation
- Implemented object detection labeling workflows via Makesense.ai for downstream CV models
Computer Vision OpenCV Image Segmentation
Feb 2024 β Mar 2024 Β· Jais, Uttar Pradesh
Designed end-to-end ML research workflows covering preprocessing, feature extraction, training, and validation.
- Strengthened research reproducibility across multiple experimental configurations
- Built foundational competency in statistical modeling and experimental analysis
Statistical Modeling Research Experimental Design
| Recognition | Details |
|---|---|
| π₯ ALLEN SOPAN 2023 | Top 1% among 100,000+ participants nationally |
| π₯ AlgoUniversity | 2nd Runner-Up, Graph Theory Programming Camp |
| π Kode Current Hackathon | Top 10 among 1,000+ competing teams |
| π Lyzr Agentathon | Ranked 38/500+ builders β shortlisted top 50 |
| π₯ Hacktoberfest Hackathon 2025 | Consolation Prize (3rd Place), Bengaluru |
| π» Competitive Programming | 400+ CodeChef, 300+ LeetCode problems solved |
Learning:
- Advanced Multi-Agent Orchestration
- Distributed Systems for ML Infrastructure
- Advanced Retrieval-Augmented Generation Architectures
Building:
- Production-grade MLOps pipelines on AWS
- Multimodal retrieval systems at scale
Exploring:
- Vector database optimization
- Event-driven ML infrastructure
Open_To:
- AI/ML Engineer roles
- Data Science / Data Analyst roles
- Computer Vision Engineer roles