I am Anmol Tripathi, a Quality Data Scientist and Machine Learning Engineer with 7+ years of experience across network operations, analytics, predictive modeling, deep learning, NLP, retrieval-augmented generation, agentic AI, and applied AI.
At Hach, I develop machine-learning and AI solutions for product-quality intelligence, including source-grounded RAG workflows, multi-stage NLP classification, analytics automation, and executive reporting. My public portfolio translates that experience into non-confidential, reproducible projects spanning statistical learning, neural networks, sequence models, Transformers, multimodal AI, retrieval, agentic workflows, and deployment.
Current focus: evidence-grounded RAG and agentic AI, hybrid retrieval and reranking, graph-enhanced reasoning, bounded tool use, transparent evaluation, and production-oriented ML workflows.
|
10 end-to-end projects covering summarization, translation, retrieval, reranking, long-document QA, instruction tuning, VQA, Vision Transformers, CLIP, and RAG, with selected public demos.
|
Applied generation systems including schema-aware Text-to-SQL, streaming-style speech recognition, and source-grounded data-to-text executive reporting.
|
|
13 projects demonstrating statistical reasoning, EDA, feature engineering, forecasting, segmentation, model comparison, explainability, and business communication.
|
Computer-vision projects organized around convolutional architectures, transfer learning, reproducible data pipelines, model evaluation, and visual error analysis.
|
Four evidence-driven systems demonstrating retrieval, graph reasoning, bounded agents, verification, safety, observability, and domain-specific evaluation. The three standalone systems run as complete local engineering implementations; their public Hugging Face Spaces are transparent portfolio demonstrations built from validated outputs rather than unsupported claims of free GPU-backed inference.
|
Evidence-grounded industrial root-cause investigation. Combines BGE-M3 + BM25/RRF retrieval, LangGraph orchestration, local Qwen3 reasoning, bounded analytical tools, independent verification, and Phoenix observability. Frozen test: 93% RCA Top-1 · 100% Top-3 · 100% citation precision · 100% verification pass rate
|
Graph-enhanced repository intelligence and safe coding agent. Maps issues to files and symbols, expands repository relationships, discovers affected tests, and verifies bounded candidate changes inside a network-disabled Docker sandbox. Frozen test: 84.3% File Recall@20 · 75.7% File Recall@10 · 30.9 ms selected-retriever latency
|
|
Temporal financial due-diligence and risk intelligence. Integrates SEC EDGAR filings, XBRL facts, hybrid retrieval, deterministic financial tools, temporal comparison, and a provenance-preserving company-risk graph. Evaluation: 100% textual Recall@10 · 100% XBRL fact selection · 90% blind graph-edge precision · 100% citation precision
|
Public portfolio discovery through grounded retrieval. Searches 220 public AI/ML documents and 3,157 evidence chunks using MiniLM embeddings and hybrid retrieval, then returns source-cited answers with relevance and latency details. Delivery: reproducible Python evaluation · Next.js APIs · GitHub Actions · production Vercel deployment
|
Across these systems: hybrid retrieval · reranking · GraphRAG · agent orchestration · deterministic tools · citation verification · frozen-test evaluation · latency analysis · FastAPI · Docker · CI/CD · responsible-use controls
My portfolio brings together my professional experience, Agentic RAG systems, flagship AI/ML projects, deployed applications, technical case studies, education, credentials, and résumé in one recruiter-friendly experience.
| Project | What it demonstrates | Explore |
|---|---|---|
| Schema-Aware Text-to-SQL | Fine-tunes CodeT5+ 770M with LoRA to generate SQLite from natural-language questions and database schemas, then validates and safely executes approved read-only queries. | Source code · Live application |
| Vision Transformer Browser Classifier | Compares DeiT-tiny with ResNet-18, validates PyTorch-to-ONNX parity, visualizes attention rollout, and performs private WebGPU/WASM inference in the browser. | Source code · Live application |
| Streaming Speech Recognition with Whisper | Transcribes microphone or uploaded audio through a Whisper encoder-decoder workflow with chunking, timestamps, language detection, and robustness-oriented evaluation. | Source code · Live application |
| Data-to-Text Executive Report Generator | Converts structured KPI tables into source-grounded executive narratives while checking numerical claims, exposing source-cell evidence, and blocking unsupported statements. | Source code · Live application |
| Capability | Evidence in this portfolio |
|---|---|
| Applied Data Science | EDA, statistical analysis, feature engineering, forecasting, predictive modeling, and business visualization |
| Classical Machine Learning | Classification, regression, clustering, ensemble modeling, calibration, benchmarking, and error analysis |
| Deep Learning | ANN, CNN, RNN, LSTM, bidirectional LSTM, autoencoder, encoder-decoder, and Transformer architectures |
| NLP, RAG & Agents | Semantic search, hybrid retrieval, reranking, GraphRAG, grounded generation, agent orchestration, tool use, and citation verification |
| Computer Vision & Multimodal AI | CNNs, Vision Transformers, visual question answering, and CLIP image-text retrieval |
| ML & AI Engineering | Reproducible training, local GPU/BF16 inference, FastAPI, Docker, observability, ONNX, automated tests, CI/CD, and deployment |
| Analytics Engineering | Python and SQL automation, data validation, Power BI, Tableau, Excel, KPI reporting, and decision support |
September 2024 - Present · United States
- Develop AI and machine-learning solutions for product-quality intelligence, analytics automation, and operational decision support.
- Designed an internal RAG and LLM-powered agent over structured quality records dating back to 2015 and thousands of technical documents, with source-grounded responses for approximately 50 potential users across R&D and quality teams.
- Developed a calibrated, multi-stage NLP framework for predicting interconnected quality categories using Transformer representations, sparse word- and character-level NLP, structured LightGBM models, ensemble learning, and chronological validation.
- Automate monthly, biweekly, and rolling-period quality analytics with Python, SQL Server, Power BI, and Excel, supporting KPI monitoring, root-cause investigation, data validation, and executive reporting.
Public repositories contain non-confidential portfolio work only. Company data, source code, internal systems, and proprietary methodology are intentionally excluded.
May 2023 - August 2024 · Tucson, Arizona
- Developed predictive workflows from longitudinal wearable-sensor data for approximately 135 research participants.
- Built SQL-to-model pipelines covering preprocessing, temporal feature engineering, sequence generation, model tuning, validation, and prediction-error analysis.
- Compared SVM, LSTM, bidirectional LSTM, CNN, autoencoder, ARIMA, and SARIMA approaches; LSTM-based modeling produced the strongest internal research result and refined the estimated labor-prediction window from approximately 14 days to approximately one day.
- Communicated findings and limitations to interdisciplinary collaborators without presenting research estimates as clinically validated predictions.
Earlier experience - analytics, retail operations, and network engineering
- Student Assistant Manager · University of Arizona BookStores (Dec 2022 - Aug 2023): analyzed approximately 80,000-100,000 monthly sales records, supported inventory planning, built Tableau reporting, and trained 8-10 team members.
- Student Assistant · University of Arizona BookStores (Sep 2022 - Dec 2022): analyzed sales and customer data, validated recurring reports, and supported operational decision-making before promotion within four months.
- Network Operations Center Engineer · Orange Business Services (Apr 2022 - Aug 2022): built an internally evaluated network-failure prediction proof of concept, automated operational analysis, and enhanced KPI dashboards.
- Associate NOC Engineer · Orange Business Services (Dec 2019 - Mar 2022): analyzed more than 10,000 network-performance metrics daily using Python, SQL, and Excel while supporting fault investigation and incident management.
- Graduate Engineering Trainee · Orange Business Services (Jun 2019 - Dec 2019): developed foundations in enterprise networking, monitoring, troubleshooting, data analysis, and operational reporting.
- Frame the decision - define the real problem, user, target, constraints, and meaningful success criteria.
- Establish evidence - validate the data, build baselines, compare candidates, calibrate where needed, and inspect failure modes.
- Engineer for reuse - separate training and inference, preserve metadata, test artifacts, automate validation, and document assumptions.
- Deliver responsibly - ground outputs in evidence, expose limitations, protect sensitive data, and communicate results for technical and business audiences.
| Institution | Program | Academic result |
|---|---|---|
| University of Arizona | Master of Science in Data Science | GPA: 3.889 / 4.000 |
| Texas McCombs School of Business | Post Graduate Program in Data Science and Business Analytics | Overall grade: 4.00 / 4.00 |
| Amity University | Bachelor's Degree in Electronics and Telecommunications | 2015-2019 |
Graduate curriculum
University of Arizona: Ethical Issues in Information; Introduction to Machine Learning; Data Mining and Discovery; Information Research Methods; Data Analysis and Visualization; Artificial Intelligence; Applied NLP; Neural Networks; SQL/NoSQL Databases; Independent Study.
Texas McCombs: Python for Data Science; Statistical Methods for Decision Making; Advanced Statistics; Data Mining; Predictive Modeling; Machine Learning; Time Series Forecasting; Tableau; SQL; Marketing and Retail Analytics; Finance and Risk Analytics; Capstone Project.
- Google Data Analytics Professional Certificate - Google
- Complete A.I. & Machine Learning, Data Science Bootcamp
- Microsoft Excel: Advanced Excel Formulas & Functions
- The Complete SQL Bootcamp: Go from Zero to Hero
- Introduction to Python
View credential records on LinkedIn →
| Path | Best starting point |
|---|---|
| RAG / Agentic AI | ReliabilityOps · RepoAtlas · FilingsGraph · AI Portfolio RAG Assistant |
| Transformers / Multimodal AI | Transformer projects - fine-tuning, retrieval, reranking, ONNX, multimodal AI, evaluation, and deployment |
| NLP / Generative AI | Encoder-decoder projects - Text-to-SQL, speech recognition, and data-to-text systems |
| Data Science / Analytics | Applied DS & ML portfolio - 13 projects spanning analysis, modeling, evaluation, and business communication |
| Computer Vision | CNN projects and Vision Transformer work |
| Sequence Modeling | Simple RNN → LSTM → Bidirectional LSTM |
| Neural Network Foundations | ANN projects - classification, regression, risk, optimization, embeddings, and deployment |
Open to Data Scientist, Machine Learning Engineer, NLP, and Applied AI opportunities.
Visit portfolio · Connect on LinkedIn · Explore all repositories
Primary email · Alternate email
Preview résumé · Download résumé PDF