LLM & RAG evaluation testing framework✨
-
Updated
Sep 1, 2026 - Python
LLM & RAG evaluation testing framework✨
A dashboard for monitoring and analyzing AI application performance.
Lightweight eval framework for LLMs & AI apps combining deterministic scoring, LLM-as-judge, and regression testing.
A structured, reusable framework for the defensive quality assurance and evaluation of AI systems — LLMs, chatbots, RAG pipelines, and agents. 21 test categories, evaluation scorecards, and an AI-specific bug taxonomy.
QA sandbox sportsbook demonstrating API, SQL, and E2E testing with Playwright, a token-authed test harness, and AI-assisted workflows.
AI framework for dynamic calibration of ESC/TCS automotive systems, focused on Functional Safety (ISO 26262).
Automated AI Quality Assurance test suite for evaluating LLM legal contract summaries using DeepEval, Pytest, and local Ollama (qwen2.5:7b).
Enterprise-style AI Quality Evaluation Framework for testing Generative AI/LLM applications using automated evals, hallucination detection, prompt regression testing, latency validation, and CI/CD quality gates.
To associate your repository with the ai-quality-assurance topic, visit your repo's landing page and select "manage topics."