SDET focused on test automation and AI/LLM evaluation β building the tooling that catches bugs and bad model behavior before your users do.
π§° Mobile/Web E2E Β· UI Β· API testing
π Mexico Β· π₯¨ Working from home
I build evaluation, observability, and red-teaming pipelines for LLM-based agents β mostly around a fictional "Trailhead Travel" support bot I use as a testbed for these tools.
-
π§ͺ trailhead-travel-agent-eval LLM-agent quality suite (single-turn Q&A, multi-turn chat, RAG) using DeepEval β correctness, safety, and hallucination metrics scored by a local judge model.
-
π trailhead-travel-rag-eval RAG-pipeline evaluation using Ragas β faithfulness, context precision/recall, and answer relevancy, with noise-aware CI thresholds.
-
π trailhead-travel-observability Tracing and CI regression gating for an LLM agent using LangSmith β versioned eval datasets, a GitHub Actions prompt-regression gate, and drift monitoring.
-
π‘οΈ trailhead-travel-red-team Automated red-teaming and safety layering for a RAG agent β Promptfoo attack generation, Guardrails AI validation, Llama Guard classification, before/after break-rate comparison.
π Currently building out UI/E2E test automation projects with Selenium and Playwright β stay tuned.
Test Automation
AI/LLM Evaluation