Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

RAG Support Agent

A hybrid-search RAG system that answers questions from a business's own documents — PDF, DOCX, CSV, or plain text — with source citations, and refuses to guess when the answer isn't in the docs.

Domain-agnostic by design: the same engine runs a SaaS product-support bot and an internal HR-policy bot from two different document folders, with zero code changes. Point it at any business's docs and it works.

Architecture

The problem this solves

Support teams answer the same questions over and over — "how do I reset my password," "what's the refund policy," "how much parental leave do I get" — because the answers live in scattered PDFs and wikis nobody has time to search. A generic RAG chatbot is the single most-requested "AI agent" job type on Upwork, across every industry.

What makes this different from a tutorial-level RAG demo

Tutorial RAG demo This system
Retrieval Vector search only Hybrid: BM25 keyword search + vector search, fused with Reciprocal Rank Fusion
Result ordering Whatever the vector search returns Reranked with a second, independent relevance signal
Document formats Clean markdown files PDF, DOCX, CSV, TXT — what real client docs actually look like
Chunking Fixed character splitting Structure-aware — splits on headings first, falls back to size-based
Confidence "Trust me, it retrieves well" Eval harness with real IR metrics: Recall@3, MRR, source precision
Runs without an API key Usually not Yes — fully offline embeddings, LLM optional

Results

Eval harness (eval/run_eval.py) against 10 labeled Q&A pairs:

Metric Score
Recall@3 100% (10/10)
MRR 0.73
Source precision 45%

13/13 unit and integration tests passing. The eval caught two real retrieval bugs during development — a BM25 zero-IDF edge case and a reranker weighting issue — both documented with the fix in docs/case-study.md, along with what that 45% precision number actually reveals about the retrieval/precision tradeoff.

Architecture

Documents (PDF/DOCX/CSV/TXT)
        │
        ▼
Parser → structure-aware chunker
        │
        ▼
LocalLSAEmbedder (TF-IDF + SVD)
        │
   ┌────┴────┐
   ▼         ▼
ChromaDB    BM25+
(vector)   (keyword)
   │         │
   └────┬────┘
        ▼
Hybrid fusion (Reciprocal Rank Fusion)
        │
        ▼
Reranker (term-coverage cross-check)
        │
        ▼
Grounded answer + source citations

The same RAGPipeline object (app/pipeline.py) powers the live API, the eval harness, and every test — one code path, not a separate "demo version."

Tech stack

  • FastAPI — /ask endpoint plus a live /eval/results endpoint that reruns the eval and returns scores as JSON
  • ChromaDB — real embedded vector database (not an in-memory hack)
  • BM25+ (rank-bm25) — keyword retrieval, chosen over classic BM25 specifically to avoid a zero-IDF edge case on small knowledge bases (see case study)
  • scikit-learn TF-IDF + TruncatedSVD — local, zero-cost semantic embeddings (LSA). Written behind a swappable interface — see "on embeddings" below
  • pypdf, python-docx, pandas — real document format parsing
  • Anthropic Claude (optional) — upgrades answers from extractive (top chunk returned directly) to natural-language generation when ANTHROPIC_API_KEY is set

On embeddings

This repo uses TF-IDF + SVD instead of neural embeddings (sentence-transformers) because it was built in a sandboxed environment without internet access to Hugging Face or disk space for model weights — not because it's the "right" production choice. It's a real semantic embedding technique (LSA), not keyword search in disguise, and it's genuinely useful for keeping the whole repo clonable with zero setup cost. But every embedding class implements the same BaseEmbedder interface, so swapping in real transformer embeddings is a contained ~15-line change — the exact shape of that change is written out in app/embedding/embedder.py. Full reasoning in the case study.

Running it locally

git clone <this-repo>
cd rag-support-agent
pip install -r requirements.txt

cp .env.example .env   # optional: add ANTHROPIC_API_KEY for generated (not just extractive) answers
uvicorn app.main:app --reload

Open http://127.0.0.1:8000 for the chat UI, or query directly:

curl -X POST http://127.0.0.1:8000/ask \
  -H "Content-Type: application/json" \
  -d '{"question": "What are the API rate limits on the free plan?", "knowledge_base": "saas_support"}'

Swap knowledge_base to hr_policies to see the exact same engine answer from a completely different document set.

Run tests and eval:

pytest tests/ -v
python -m eval.run_eval

Or with Docker:

docker build -t rag-support-agent .
docker run -p 8000:8000 rag-support-agent

Adding your own documents

Drop PDF/DOCX/CSV/TXT files into a new folder under sample_docs/, then query with knowledge_base set to that folder name. No code changes needed — this is the actual point of the project.

What I deliberately didn't build

  • No auth or multi-tenant document isolation — see case study for reasoning
  • No streaming responses — straightforward FastAPI addition, orthogonal to what this project demonstrates
  • No file-watching re-ingestion — documents are indexed once at query time; live re-indexing on file change is real but separate production work

Full reasoning, including the two real bugs this project's tests caught and what the eval numbers actually reveal, is in docs/case-study.md.

Project structure

app/
├── ingestion/       parsers (PDF/DOCX/CSV/TXT) + structure-aware chunker
├── embedding/        swappable embedder interface, LocalLSAEmbedder default
├── retrieval/        ChromaDB vector store, BM25+ retriever, hybrid fusion
├── rerank/           term-coverage reranker
├── answer.py          grounded answer generation, LLM-optional
├── pipeline.py        wires the full flow together — single entrypoint
└── main.py            FastAPI app

tests/                13 unit + integration tests
eval/                 scenarios.jsonl + scoring harness
docs/                 architecture diagram + case study writeup
sample_docs/          two demo knowledge bases (SaaS support, HR policies)
frontend/             single-page chat UI

Second in a portfolio series demonstrating production AI agent architecture. First project: hotel-ai-booking-agent — a tool-calling booking agent with code-enforced guardrails.

About

Hybrid-search RAG agent that answers questions from a business's own docs (PDF/DOCX/CSV/TXT) with source citations. BM25 + vector search + reranking, eval-scored (100% recall@3), domain-agnostic. Built with FastAPI, ChromaDB, and a swappable embedding layer.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages