Production-ready enterprise RAG and structured data extraction pipeline using LangGraph, Instructor, and Pydantic with automated CI evaluation metrics.
-
Updated
Aug 28, 2026 - Python
Production-ready enterprise RAG and structured data extraction pipeline using LangGraph, Instructor, and Pydantic with automated CI evaluation metrics.
benchmarking jailbreak-dataset fine-tuning-tools llm-evaluation official high-priority benchmarking task evaluation checkpoint shared memory sandbox API endpoints structured data schemas collaborative environments
Replication-first study of sociolinguistic effects on LLM epistemic judgments, with behavioral validity gates before mechanistic interpretation.
Defense-in-depth safety study on Mistral-7B: how much of an LLM's safety lives in the weights vs. external guardrails, and how fragile it is under fine-tuning. Defensive research.
To associate your repository with the llm-evaluations topic, visit your repo's landing page and select "manage topics."