Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 

Repository files navigation

gita-rag

A small, dependency-light retrieval-augmented generation pipeline over the Bhagavad-gītā. Scrape the verses into a structured corpus, embed them into a local vector store, and ask natural-language questions that are answered with verse-cited context by any local OpenAI-compatible LLM.

It's a clean, readable example of the whole RAG loop — scrape → chunk → embed → retrieve → prompt → cite — with no framework magic. Three short scripts, ~400 lines total.

scripts/scrape.py   →  corpus/bg.jsonl      (one structured record per verse)
scripts/embed.py    →  chroma_db/           (ChromaDB index, all-MiniLM-L6-v2)
scripts/query.py    →  cited answer         (retrieve top-k → local LLM)

Why it's built this way

  • One verse = one document. The embedding text concatenates reference + translation + purport, so a question retrieves the verse and its commentary. Sanskrit transliteration and synonyms are kept as metadata for citation but not embedded.
  • Strict grounding. The system prompt forces the model to answer only from retrieved verses, cite every claim (BG 2.47), quote translations verbatim, and say "not enough context" rather than speculate.
  • Bring your own LLM. query.py talks to any OpenAI-compatible /v1/chat/completions endpoint — llama.cpp's server, vLLM, Ollama's /v1, etc. Nothing is hardcoded to a vendor.

Quick start

pip install requests beautifulsoup4 chromadb

# 1. Build the corpus (polite scraper, caches raw HTML, ~0.5s/verse)
python scripts/scrape.py

# 2. Embed into a local ChromaDB index
python scripts/embed.py

# 3. Ask — point at your local model first
export GITA_BRAIN_URL="http://localhost:8080/v1/chat/completions"
export GITA_MODEL="your-model-name"
python scripts/query.py "what does the Gita say about karma yoga?"

# retrieval only, no LLM:
python scripts/query.py --raw --k 8 "how is the soul described?"
Env var Default Meaning
GITA_BRAIN_URL http://localhost:8080/v1/chat/completions OpenAI-compatible chat endpoint
GITA_MODEL local-model model name passed to the endpoint

A note on the source text

This repo ships code only. The corpus (corpus/, raw/) and the built index (chroma_db/) are not included — the Bhagavad-gītā As It Is (1972 Macmillan edition) translation and purports are copyrighted by the Bhaktivedanta Book Trust. scrape.py builds your own local copy for personal study use; please respect the BBT's rights and don't redistribute the generated corpus.

License

MIT — see LICENSE. Applies to the code in this repository, not to any source text you generate with it.

About

Verse-cited retrieval-augmented generation over the Bhagavad-gita (scrape → embed → query). Code-only; bring your own corpus + local LLM.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages