ML Engineer | Training โ Inference โ Production
8+ years building ML systems end-to-end: from training custom models to optimizing inference at the GPU kernel level.
- GPU performance engineering with CUDA and Triton
- LLM training, inference optimization, and evaluation
- Production-oriented RAG and ML systems
| Project | What it demonstrates |
|---|---|
| gpt2-engine | GPT-2 inference with custom Triton kernels, KV caching, quantization, and benchmarks |
| triton-code | Triton kernel experiments and performance analysis |
| cuda-reduction | CUDA reduction strategies benchmarked against CUB |
| tiny-stories-llm | Transformer language-model training from scratch |
| ssl-foundations | Self-supervised representation evaluation with linear probing and k-NN |
| meditations-rag | Agentic RAG with LangGraph, hybrid retrieval, Qdrant, and FastAPI |
๐ซ Open to remote opportunities (contract or full-time) | Flexible on US/EU timezones

