A small, dependency-light retrieval-augmented generation pipeline over the Bhagavad-gītā. Scrape the verses into a structured corpus, embed them into a local vector store, and ask natural-language questions that are answered with verse-cited context by any local OpenAI-compatible LLM.
It's a clean, readable example of the whole RAG loop — scrape → chunk → embed → retrieve → prompt → cite — with no framework magic. Three short scripts, ~400 lines total.
scripts/scrape.py → corpus/bg.jsonl (one structured record per verse)
scripts/embed.py → chroma_db/ (ChromaDB index, all-MiniLM-L6-v2)
scripts/query.py → cited answer (retrieve top-k → local LLM)
- One verse = one document. The embedding text concatenates reference + translation + purport, so a question retrieves the verse and its commentary. Sanskrit transliteration and synonyms are kept as metadata for citation but not embedded.
- Strict grounding. The system prompt forces the model to answer only from
retrieved verses, cite every claim (
BG 2.47), quote translations verbatim, and say "not enough context" rather than speculate. - Bring your own LLM.
query.pytalks to any OpenAI-compatible/v1/chat/completionsendpoint — llama.cpp's server, vLLM, Ollama's/v1, etc. Nothing is hardcoded to a vendor.
pip install requests beautifulsoup4 chromadb
# 1. Build the corpus (polite scraper, caches raw HTML, ~0.5s/verse)
python scripts/scrape.py
# 2. Embed into a local ChromaDB index
python scripts/embed.py
# 3. Ask — point at your local model first
export GITA_BRAIN_URL="http://localhost:8080/v1/chat/completions"
export GITA_MODEL="your-model-name"
python scripts/query.py "what does the Gita say about karma yoga?"
# retrieval only, no LLM:
python scripts/query.py --raw --k 8 "how is the soul described?"| Env var | Default | Meaning |
|---|---|---|
GITA_BRAIN_URL |
http://localhost:8080/v1/chat/completions |
OpenAI-compatible chat endpoint |
GITA_MODEL |
local-model |
model name passed to the endpoint |
This repo ships code only. The corpus (corpus/, raw/) and the built
index (chroma_db/) are not included — the Bhagavad-gītā As It Is (1972
Macmillan edition) translation and purports are copyrighted by the Bhaktivedanta
Book Trust. scrape.py builds your own local copy for personal study use; please
respect the BBT's rights and don't redistribute the generated corpus.
MIT — see LICENSE. Applies to the code in this repository, not to any source text you generate with it.