A modern, production-grade Java service combining Retrieval-Augmented Generation (RAG) and an LLM-as-a-Judge Evaluation Engine to interpret, classify, and index complex multi-turn IT support and customer service email chains.
Operations and support organizations process thousands of multi-turn email threads. Auditing whether customer issues are truly resolved, bypassed with temporary workarounds, or abandoned due to unresponsive participants is slow, error-prone, and unscalable when done manually.
This repository provides an automated, model-agnostic solution:
- LLM-as-a-Judge Resolution Engine: Performs turn-by-turn chronological transcript analysis, decomposes threads into granular customer issues, and outputs strongly typed verdicts (
RESOLVED,WORKAROUND,UNRESOLVED,UNKNOWN) backed by exact cited evidence and sentiment analysis. - Semantic RAG Ingestion Pipeline: Ingests email chains and arbitrary text, splits them into semantic chunks, generates vector embeddings using local ONNX transformer models, and indexes them into PGVector (PostgreSQL) with rich evaluation metadata for fast similarity retrieval.
- Pluggable Multi-Model Architecture: Built on Spring AI 2.0 with dynamic model resolution, allowing seamless switching between cloud LLMs (Google Gemini, OpenAI, Claude) and local inference engines (Ollama).
flowchart TD
subgraph Ingestion["1. Document & Email Ingestion"]
EML[".eml Email Threads / Text Payloads"] --> Tika["Apache Tika Reader"]
Tika --> Parser["EmailParser\n(Header & Chronological Turn Extraction)"]
end
subgraph LLMJudge["2. LLM-as-a-Judge Service"]
Parser --> JudgeService["EmailResolutionJudgeService"]
Resolver["ChatModelResolver\n(Dynamic Model & Temperature Override)"] --> JudgeService
JudgeService --> LLM["LLM Provider\n(Google Gemini / Local ONNX / Ollama)"]
LLM --> Analysis["Structured EmailLlmAnalysis\n- Confidence Score & Rationale\n- Decomposed Issues\n- Status & Cited Evidence"]
end
subgraph Storage["3. Vector Ingestion & RAG"]
Analysis --> IngestService["EmailIngestionService"]
IngestService --> Embeddings["ONNX Local Embeddings\n(384 dimensions)"]
Embeddings --> PGVector[("PGVector Store\n(PostgreSQL + HNSW Index)")]
end
subgraph Query["4. Search & API Surface"]
Client["REST API Consumers / QA Teams"] --> Controller["ApiController / JudgeApiController"]
Controller --> DocQuery["DocumentQueryService\n(Cosine Similarity Search)"]
DocQuery --> PGVector
end
- Granular Issue Decomposition: Identifies each distinct problem within a single multi-turn email thread rather than treating the entire conversation as a monolithic ticket.
- Strict Resolution Taxonomy:
RESOLVED: Root cause diagnosed, permanent fix applied, customer confirmation verified.WORKAROUND: Temporary mitigation, bypass, or rollback applied; underlying defect remains open.UNRESOLVED: Issue failing, blocked on dependencies/permissions, or abandoned.UNKNOWN: Insufficient or ambiguous context to classify.
- Evidence & Chain-of-Thought Rationale: Produces step-by-step reasoning alongside exact quoted sentences from conversation turns for full human auditability.
- Local ONNX Embeddings: Runs 384-dimensional embedding models locally using ONNX runtime without incurring external embedding API costs or latency.
- Dynamic Model Resolution: Runtime override of target LLM model name and temperature via
JudgeOptionswithout service restart. - Comprehensive Benchmark Fixtures: Includes over 300 ground-truth email chain fixtures (
dataset/*.eml) spanning all resolution categories.
When an email thread is evaluated, the engine returns a structured EmailLlmAnalysis:
{
"confidenceScore": 0.95,
"rationale": "Turn 1 reported VPN connection drops. Turn 2 suggested MTU configuration changes. Turn 3 confirmed successful connection with no further drops.",
"issues": [
{
"status": "RESOLVED",
"issue": "VPN connection drops after 15 minutes of idle time",
"keyEvidence": [
"Applying the MTU 1420 fix resolved all connection drops completely.",
"Tested for 3 hours with stable connection."
],
"rootCauseSummary": "Packet fragmentation due to default MTU size mismatch on gateway",
"finalCustomerSentiment": "SATISFIED",
"resolutionStepsTaken": "Updated client network adapter MTU setting to 1420."
}
]
}- Java 26 JDK or newer
- Maven 3.9+
- PostgreSQL with the
pgvectorextension enabled (for vector store persistence) - Google Gemini API Key (or another configured Spring AI provider)
Set your Gemini API key in your environment:
# Linux / macOS
export GEMINI_API_KEY="your-gemini-api-key"
# Windows (PowerShell)
$env:GEMINI_API_KEY="your-gemini-api-key"
# Windows (Command Prompt)
set GEMINI_API_KEY=your-gemini-api-keyEnsure PostgreSQL is running and update src/main/resources/application.yml if necessary:
spring:
ai:
google:
genai:
api-key: ${GEMINI_API_KEY:}
chat:
options:
model: gemini-3.5-flash
embedding.transformer.enabled: true
embedding.transformer.cache.directory: ./onnx-models
vectorstore:
pgvector:
initialize-schema: true
index-type: HNSW
distance-type: COSINE_DISTANCE
dimensions: 384
table-name: vector_store
datasource:
url: jdbc:postgresql://localhost:5432/ragdb
username: ${VECTOR_DB_USR:}
password: ${VECTOR_DB_PWD:}# Compile and package
mvn clean package
# Run the Spring Boot application
mvn spring-boot:runThe application will start on http://localhost:8080.
Interactive Swagger documentation is available at http://localhost:8080/swagger-ui.html when the application is running.
Evaluates an uploaded .eml email chain and returns structured analysis without saving to vector database.
- URL:
POST /api/judge/evaluate-file - Content-Type:
multipart/form-data - Parameters:
input(File, required): The.emlemail file.modelName(Query param, optional): Specific LLM model identifier to use (e.g.gemini-3.5-flash).temperature(Query param, optional): Sampling temperature (e.g.0.0for deterministic judge output).
curl -X POST "http://localhost:8080/api/judge/evaluate-file?temperature=0.0" \
-F "input=@dataset/EMAIL_CHAIN_001_RESOLVED.eml"Evaluates an email chain and indexes each decomposed issue, root cause, and metadata into PGVector.
- URL:
POST /api/ingest-email - Content-Type:
multipart/form-data - Parameters:
input(File, required):.emlemail file.
curl -X POST "http://localhost:8080/api/ingest-email" \
-F "input=@dataset/EMAIL_CHAIN_001_RESOLVED.eml"Chunks and embeds arbitrary text documents into the vector store.
- URL:
POST /api/ingest-string - Content-Type:
application/json - Body:
{
"content": "Kubernetes pod evicted due to disk pressure on node worker-04.",
"source": "incident-reports",
"description": "Node storage alert",
"topics": ["infrastructure", "kubernetes", "disk-pressure"]
}curl -X POST "http://localhost:8080/api/ingest-string" \
-H "Content-Type: application/json" \
-d '{"content": "Kubernetes pod evicted due to disk pressure...", "source": "ops", "description": "Alert", "topics": ["k8s"]}'Searches the vector store for documents matching the semantic meaning of the query.
- URL:
GET /api/ask?q={queryText}
curl "http://localhost:8080/api/ask?q=VPN+connection+drops"Run the test suite using Maven:
mvn testUnit and slice tests run against an in-memory H2 database and mock configurations, ensuring zero external cloud dependencies during CI/CD test runs.
demo-ragapp/
βββ dataset/ # 300+ labeled .eml ground-truth email fixtures
β βββ EMAIL_CHAIN_001_RESOLVED.eml
β βββ EMAIL_CHAIN_002_WORKAROUND.eml
β βββ EMAIL_CHAIN_008_UNRESOLVED.eml
β βββ EMAIL_CHAIN_012_ABANDONED.eml
β βββ vector-database-snapshot.csv # Ground truth benchmark reference
βββ src/
β βββ main/
β β βββ java/org/example/
β β β βββ DemoApplication.java # Spring Boot application entry point
β β β βββ controller/
β β β β βββ ApiController.java # Ingestion & similarity search endpoints
β β β β βββ JudgeApiController.java # LLM judge evaluation endpoint
β β β βββ ingestion/
β β β β βββ EmailIngestionService.java # Evaluates & indexes email issues
β β β β βββ StringIngestionService.java # Splits & indexes text snippets
β β β βββ model/
β β β β βββ EmailLlmAnalysis.java # Evaluation verdict record
β β β β βββ EmailLlmAnalysisIssue.java # Decomposed issue record
β β β β βββ EmailMessage.java # Parsed email message record
β β β β βββ EmailTurn.java # Chronological turn representation
β β β β βββ IngestionReq.java # Text ingestion DTO
β β β β βββ JudgeOptions.java # Dynamic model execution options
β β β βββ service/
β β β βββ ChatModelResolver.java # Pluggable model resolver interface
β β β βββ DefaultChatModelResolver.java # ChatClient & multi-model provider
β β β βββ DocumentQueryService.java # Vector store similarity search
β β β βββ EmailParser.java # Email header & turn separator
β β β βββ EmailResolutionJudgeService.java # Judge interface
β β β βββ EmailResolutionJudgeServiceImpl.java # Prompt rubric & LLM evaluation
β β βββ resources/
β β βββ application.yml # Database, PGVector, and Spring AI configuration
β βββ test/ # Unit and integration test suites
βββ SPEC.md # System specification & architectural decisions
βββ pom.xml # Maven project dependencies and build plugins
βββ README.md # Project overview and documentation
This project is licensed under the Apache 2.0 License.