AI/ML Engineer - Lahore, Pakistan
I build LLM systems that make it to production: fine-tuning, RAG, voice agents, evaluation harnesses, and the services around them.
AI Engineer at Symufolk (Oct 2025 - present, Lahore). Building and scaling production-grade AI infrastructure across computer vision, ML systems and intelligent automation pipelines - systems that hold up under real operational load.
Before that: data annotation for autonomous-driving perception models at Motive, and two years at Aspire Analytica building an agentic RAG reconciliation platform (80% faster reconciliation), a FastAPI voice-AI platform for automotive dealerships (+35% bookings, -70% support cost), and backend architecture for five ERP implementations.
- 🔭 Currently working on production AI infrastructure at Symufolk: computer vision, ML systems and automation pipelines
- 🌱 Currently learning multi-agent orchestration with LangGraph, and rigorous LLM evaluation design
- 👯 Looking to collaborate on open-source LLM evaluation harnesses and voice-agent tooling
- 🤔 Looking for help with scaling fine-tuned model inference cost-effectively
- 💬 Ask me about LoRA fine-tuning, RAG retrieval quality, LLM-as-a-judge evaluation, and taking AI prototypes to production
- 📫 How to reach me LinkedIn or taha-ahmad.vercel.app
- 😄 Pronouns he/him
- ⚡ Fun fact I fine-tuned a model to pick cricket batting orders. It disagrees with me about half the time, and it is usually right.
FinEval - LLM Financial Reasoning Benchmark
Benchmarks Gemini 3 against Fin-o1-14B on curated financial reasoning tasks, scored by an LLM-as-judge pipeline. Flask API, batch evaluation scripts, and a Next.js dashboard with per-difficulty breakdowns.
Python - Flask - Next.js - Gemini - LLM-as-a-judge
SkipperAI - Cricket Match Strategy Engine
Fine-tuned Qwen3-1.7B with LoRA/Unsloth to recommend the optimal batting order from live match state, including synthetic training data generation and a FastAPI inference service.
PyTorch - PEFT - Unsloth - FastAPI - Hugging Face
Voice agent with a full STT to LLM to TTS loop over a RAG core. Multi-source ingestion (PDF, DOCX, web links), autonomous tool calling, and WebSocket streaming.
ElevenLabs - Gemini - Flask - WebSockets
Full-stack AI photo editor performing five artistic style transforms on portraits, built on Gemini 2.5 Flash Image.
React - TypeScript - Express - TailwindCSS
LLM & ML - PyTorch, Transformers, PEFT/LoRA, Unsloth, RAG, evaluation harnesses, prompt engineering
Serving - FastAPI, Flask, WebSockets, Docker, Render, Vercel
Full-stack - Next.js, React, TypeScript, Node/Express, Tailwind
Models - Gemini, Qwen, ElevenLabs
AI/ML engineering roles: LLM applications, fine-tuning, RAG systems, agentic workflows. Remote or Lahore-based.