Skip to content
View yotambraun's full-sized avatar

Block or report yotambraun

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
yotambraun/README.md

Hi, I'm Yotam Braun

I'm a Data Scientist with a passion for GenAI, LLM Evaluation & Benchmarking, NLP, Time Series, and GNN. I hold a Master's degree in Statistics and Data Science, with a thesis focused on the optimization of dynamic transportation networks through the application of genetic algorithms. I have experience building end-to-end machine learning pipelines, scalable AI solutions, and evaluation frameworks for AI agents. I'm open to collaborations and opportunities in the field of Data Science. Feel free to reach out to me!

Check out my Medium articles.



Tech Stack


  • Languages:

    Python R SQL


  • Frameworks & Tools:

    PyTorch TensorFlow mlflow ClearML Dataiku Hugging Face Apache Airflow Apache Kafka ClickHouse Docker LangChain LangGraph Transformers AWS Azure Databricks PySpark Dask Git GitHub Actions YAML FastAPI PostgreSQL SQL-Server Nixtla Darts Yolo OPTUNA scikit-learn Scipy Terrafom


🚀 Cool Projects

  • Saylent: The Open-Source AI Visibility Audit: My largest project to date. Buyers no longer Google a category, they ask an assistant — and ask it twice, you get two different answers. Saylent audits what ChatGPT, Claude, Gemini and Perplexity actually tell buyers about a brand, and hands back the evidence behind every line of it. One command, npx saylent audit example.com, runs a nine-stage pipeline — crawl, brand model, buyer questions, engines, judge, cited pages, site gates, fix plan, score — and writes a single self-contained HTML report. The engineering is in what it refuses to guess: brand mentions are resolved deterministically in code, never by a model, so a hallucination cannot inflate a score; every scored question is sampled more than once so the output is a range rather than an invented number; and answers are judged cross-family (one provider family grading another's output), with the judge required to quote the sentence it labelled. It then leaves the chat window entirely — fetching the pages the engines actually cited to verify presence, and knocking on your own site as six real AI user-agents to catch the gap between what robots.txt allows and what your CDN silently blocks. Findings come out as prioritized fixes you track to done, and movement is only ever measured against a frozen question set. It ships as a Node CLI (bring your own keys, local spend caps, --dry-run pricing), a fully self-hosted Next.js + Supabase + Inngest app with Postgres row-level security for multi-tenant isolation, scheduled re-checks, rival comparisons, share links and an operator console (budget cap, kill switch, hot-swappable provider keys, audit log), plus a GitHub Action and an MCP server. Nothing is proxied through a hosted service. Apache-2.0. Docs · Sample report · Live demo.
  • Project Management System with RAG: An AI-driven project management system using Retrieval-Augmented Generation (RAG) for enhanced efficiency in task execution, team collaboration, and decision-making. Read the article on Medium.
  • VisualInsight: An app leveraging Google Generative AI (Gemini) to analyze images. This Streamlit-based web application allows users to upload images for analysis, storing both the original images and results in Amazon S3. Read the article on Medium.
  • APDTFlow: A Modular Forecasting Framework for Time Series Data: APDTFlow is a modern and extensible forecasting framework for time series data that leverages advanced techniques including neural ordinary differential equations (Neural ODEs), transformer-based components, and probabilistic modeling. Its modular design allows researchers and practitioners to experiment with multiple forecasting models and easily extend the framework for new methods. With over 44,000 downloads, APDTFlow is actively used by the time series forecasting community. Read the article on Medium.
  • Toolscore: LLM Tool-Calling Evaluation Framework: A deterministic, zero-API-cost Python testing framework that evaluates how well Large Language Models perform when making tool calls. It provides composite scoring across invocation accuracy, argument matching, sequence accuracy, and redundancy detection, with auto-detection support for OpenAI, Anthropic, Google Gemini, LangChain, and MCP. With over 11,000 downloads, it includes a full CLI (eval, compare, regression), GitHub Action integration, interactive debugging, multi-format reporting (HTML/JSON/CSV/Markdown), cost tracking, and a Pytest plugin for CI/CD pipelines. Read the article on Medium.
  • PromptBeacon: AI Brand-Visibility Measurement Engine: An open-source Python engine that measures whether AI assistants recommend your brand — across six providers in a single run, on web-grounded answers with real citations rather than model memory alone. It reports Share of Voice against your competitors, a stability score with confidence intervals (a single LLM answer is not a measurement), and glass-box funnel visibility that shows exactly where you drop out. Installs and runs keyless (promptbeacon demo "Nike"), and ships a pytest plugin and GitHub Action so a build can fail when AI stops recommending you. With over 4,000 downloads.
  • FlowPrompt: Prompts as Code, Measured: A Python framework that turns prompts into typed classes with Pydantic-validated outputs — and then proves which version is actually better. compare() runs an A/B test across prompt variants and returns accuracy, latency, p-values and effect sizes; full experiments add production traffic splitting, sticky user assignment and multi-armed bandits. Provider-agnostic, with a pytest plugin for prompt regression tests in CI, OpenTelemetry tracing and multimodal support. With over 3,000 downloads.
  • Bot Chat DeepSeek:end-to-end pipeline of fine-tuning a Large Language Model (LLM) on AWS. This repository demonstrates how to prepare a custom dataset, fine-tune a language model, deploy it on AWS SageMaker, and interact with it via a Flask-based API.
  • Football Scout RAG: An AI-powered football scouting agent designed to scrape data from Transfermarkt, providing advanced player statistics and comparisons.
  • LLM-Code-Review 🤖: An automated code review tool powered by Large Language Models (LLMs), integrated as a GitHub Action. It provides real-time code analysis on every pull request, offering detailed feedback on security risks, performance bottlenecks, and code quality—right in your pull requests. No more waiting for reviews or missing critical issues.

Let's Connect!





Pinned Loading

  1. Project_Management_System_with_RAG Project_Management_System_with_RAG Public

    Python 10 2

  2. APDTFlow APDTFlow Public

    APDTFlow is a modern and extensible forecasting framework for time series data that leverages advanced techniques including neural ordinary differential equations (Neural ODEs), transformer-based c…

    Python 45 5

  3. bot_chat_deepseek bot_chat_deepseek Public

    Python 1

  4. Toolscore Toolscore Public

    Python framework for evaluating LLM tool-calling behavior with comprehensive metrics on accuracy, efficiency, and correctness

    Python 5

  5. saylent saylent Public

    The open-source audit of what AI assistants say about your brand, with the receipts. CLI, self-hostable app, GitHub Action.

    TypeScript

  6. flowprompt flowprompt Public

    Type-safe prompt management with automatic optimization for LLMs. DSPy-style optimization, A/B testing, multimodal support, and more.

    Python 6