I'm a Data Scientist with a passion for GenAI, LLM Evaluation & Benchmarking, NLP, Time Series, and GNN. I hold a Master's degree in Statistics and Data Science, with a thesis focused on the optimization of dynamic transportation networks through the application of genetic algorithms. I have experience building end-to-end machine learning pipelines, scalable AI solutions, and evaluation frameworks for AI agents. I'm open to collaborations and opportunities in the field of Data Science. Feel free to reach out to me!
Check out my Medium articles.
- Saylent: The Open-Source AI Visibility Audit: My largest project to date. Buyers no longer Google a category, they ask an assistant — and ask it twice, you get two different answers. Saylent audits what ChatGPT, Claude, Gemini and Perplexity actually tell buyers about a brand, and hands back the evidence behind every line of it. One command,
npx saylent audit example.com, runs a nine-stage pipeline — crawl, brand model, buyer questions, engines, judge, cited pages, site gates, fix plan, score — and writes a single self-contained HTML report. The engineering is in what it refuses to guess: brand mentions are resolved deterministically in code, never by a model, so a hallucination cannot inflate a score; every scored question is sampled more than once so the output is a range rather than an invented number; and answers are judged cross-family (one provider family grading another's output), with the judge required to quote the sentence it labelled. It then leaves the chat window entirely — fetching the pages the engines actually cited to verify presence, and knocking on your own site as six real AI user-agents to catch the gap between whatrobots.txtallows and what your CDN silently blocks. Findings come out as prioritized fixes you track to done, and movement is only ever measured against a frozen question set. It ships as a Node CLI (bring your own keys, local spend caps,--dry-runpricing), a fully self-hosted Next.js + Supabase + Inngest app with Postgres row-level security for multi-tenant isolation, scheduled re-checks, rival comparisons, share links and an operator console (budget cap, kill switch, hot-swappable provider keys, audit log), plus a GitHub Action and an MCP server. Nothing is proxied through a hosted service. Apache-2.0. Docs · Sample report · Live demo. - Project Management System with RAG: An AI-driven project management system using Retrieval-Augmented Generation (RAG) for enhanced efficiency in task execution, team collaboration, and decision-making. Read the article on Medium.
- VisualInsight: An app leveraging Google Generative AI (Gemini) to analyze images. This Streamlit-based web application allows users to upload images for analysis, storing both the original images and results in Amazon S3. Read the article on Medium.
- APDTFlow: A Modular Forecasting Framework for Time Series Data: APDTFlow is a modern and extensible forecasting framework for time series data that leverages advanced techniques including neural ordinary differential equations (Neural ODEs), transformer-based components, and probabilistic modeling. Its modular design allows researchers and practitioners to experiment with multiple forecasting models and easily extend the framework for new methods. With over 44,000 downloads, APDTFlow is actively used by the time series forecasting community. Read the article on Medium.
- Toolscore: LLM Tool-Calling Evaluation Framework: A deterministic, zero-API-cost Python testing framework that evaluates how well Large Language Models perform when making tool calls. It provides composite scoring across invocation accuracy, argument matching, sequence accuracy, and redundancy detection, with auto-detection support for OpenAI, Anthropic, Google Gemini, LangChain, and MCP. With over 11,000 downloads, it includes a full CLI (eval, compare, regression), GitHub Action integration, interactive debugging, multi-format reporting (HTML/JSON/CSV/Markdown), cost tracking, and a Pytest plugin for CI/CD pipelines. Read the article on Medium.
- PromptBeacon: AI Brand-Visibility Measurement Engine: An open-source Python engine that measures whether AI assistants recommend your brand — across six providers in a single run, on web-grounded answers with real citations rather than model memory alone. It reports Share of Voice against your competitors, a stability score with confidence intervals (a single LLM answer is not a measurement), and glass-box funnel visibility that shows exactly where you drop out. Installs and runs keyless (
promptbeacon demo "Nike"), and ships a pytest plugin and GitHub Action so a build can fail when AI stops recommending you. With over 4,000 downloads. - FlowPrompt: Prompts as Code, Measured: A Python framework that turns prompts into typed classes with Pydantic-validated outputs — and then proves which version is actually better.
compare()runs an A/B test across prompt variants and returns accuracy, latency, p-values and effect sizes; full experiments add production traffic splitting, sticky user assignment and multi-armed bandits. Provider-agnostic, with a pytest plugin for prompt regression tests in CI, OpenTelemetry tracing and multimodal support. With over 3,000 downloads. - Bot Chat DeepSeek:end-to-end pipeline of fine-tuning a Large Language Model (LLM) on AWS. This repository demonstrates how to prepare a custom dataset, fine-tune a language model, deploy it on AWS SageMaker, and interact with it via a Flask-based API.
- Football Scout RAG: An AI-powered football scouting agent designed to scrape data from Transfermarkt, providing advanced player statistics and comparisons.
- LLM-Code-Review 🤖: An automated code review tool powered by Large Language Models (LLMs), integrated as a GitHub Action. It provides real-time code analysis on every pull request, offering detailed feedback on security risks, performance bottlenecks, and code quality—right in your pull requests. No more waiting for reviews or missing critical issues.







