Ask questions about FDA drug labels in plain English and get grounded, cited answers — never invented ones. PV-Insight combines semantic retrieval over official openFDA label text with a Neo4j knowledge graph enriched with real FAERS adverse-event report counts.
Built entirely on free and open-source tools and free cloud tiers, and validated with a risk-based OQ/PQ test suite aligned to FDA Computer Software Assurance (CSA) and GAMP-5 principles.
⚠️ Intended use: informational, research, and demonstration purposes only. Not medical advice, and not a validated system of record for regulatory decision-making. Always verify against the official label.
Question: "What are the most frequently reported adverse events for tramadol, and what does its label warn about?"
Answer:
The most frequently reported adverse events for tramadol are: Dependence (n=7823), Overdose (n=3889), Vomiting (n=3327)… The label warns about the risk of addiction, abuse, and misuse, as well as life-threatening respiratory depression… [1, 2, 4, 5]
Sources: [1] TRAMADOL HYDROCHLORIDE – adverse_reactions · [2] warnings_and_cautions · [4] boxed_warning
Two retrieval paths are fused into one answer:
- Semantic search over label text → the exact passages, cited by number.
- Knowledge graph → drug class, related drugs, and FAERS report counts.
┌──────────── openFDA API ────────────┐
│ /drug/label.json /drug/event.json │
└──────────┬──────────────┬────────────┘
labels │ │ FAERS counts
▼ ▼
normalize + dedup enrich_faers
│ │
data/labels.jsonl │ (immutable staging snapshot)
│ │
┌─────────────┴───┐ ▼
▼ ▼ ┌──────────────┐
HF embeddings (metadata) │ Neo4j AuraDB │ Drug ─ AdverseEvent
│ │ graph │ Manufacturer/Class
▼ └──────┬───────┘
Chroma (local) ──export──▶ data/vectors.npz │
│ │
└──────┬───────┘
▼
GraphRAG orchestrator
│
HF Inference LLM
│
Streamlit UI / CLI
| Layer | Technology | Cost |
|---|---|---|
| Source data | openFDA (labels + FAERS) | Free, no key required |
| Embeddings | HF Inference API — all-MiniLM-L6-v2 (384-dim) |
Free tier |
| Vector search | Chroma (local) → .npz snapshot (deployed) |
Free |
| Graph | Neo4j AuraDB Free | Free tier |
| Generation | HF Inference — Llama 3.1 8B Instruct (auto-fallback) | Free tier |
| UI / hosting | Streamlit + Community Cloud | Free tier |
Why two vector backends? Locally, Chroma is writable — needed for ingestion.
Cloud hosts have ephemeral filesystems, so the deployed app loads an immutable,
pre-built .npz snapshot: no database process, ~4 MB, instant cold start, and
exact (not approximate) search. An OQ test asserts both backends return identical
rankings and distances.
Prerequisites: Python 3.11+ and a free Hugging Face token. Neo4j is optional — without it you lose graph context but still get cited answers.
git clone https://github.com/<your-username>/pv-insight.git
cd pv-insight
python -m venv .venv
source .venv/bin/activate # Windows: .\.venv\Scripts\Activate.ps1
pip install -r requirements-dev.txt # dev install (includes Chroma + pytest)
cp .env.example .env # Windows: Copy-Item .env.example .envEdit .env and set at minimum:
HF_API_TOKEN=hf_your_token_hereOptional, for the knowledge graph (AuraDB Free):
NEO4J_URI=neo4j+s://xxxxxxxx.databases.neo4j.io
NEO4J_USERNAME=neo4j
NEO4J_PASSWORD=your-passwordVerify your setup, then build the dataset:
python -m pv_insight.readiness # all components should report OK
python -m pv_insight.ingest --max-records 50 # fetch labels from openFDA
python -m pv_insight.load # embed + load graph & vectors
python -m pv_insight.enrich_faers # add FAERS adverse events (needs Neo4j)Run it:
streamlit run streamlit_app.py # web app
python -m pv_insight.ask "What warnings apply to naproxen?" # CLI| Command | Purpose |
|---|---|
python -m pv_insight.readiness |
Health-check every service; UTC audit record. Exit 0 = ready |
python -m pv_insight.ingest [--search Q] [--max-records N] |
Fetch + normalize openFDA labels → data/labels.jsonl |
python -m pv_insight.load [--limit N] [--skip-graph] |
Embed chunks → vector store; MERGE drugs → graph |
python -m pv_insight.enrich_faers [--top-events N] |
Add HAS_ADVERSE_EVENT edges with FAERS counts |
python -m pv_insight.export_snapshot |
Build data/vectors.npz for deployment |
python -m pv_insight.search "text" [-k N] |
Raw semantic search (no LLM) |
python -m pv_insight.ask "question" [--no-graph] |
Full GraphRAG answer with citations |
pytest |
Full OQ + PQ validation suite |
All loaders are idempotent — re-running converges instead of duplicating.
Set via .env locally, or as Streamlit secrets / environment variables when deployed.
See .env.example for the full annotated list.
| Variable | Default | Notes |
|---|---|---|
HF_API_TOKEN |
— | Required. Embeddings + generation |
HF_EMBEDDING_MODEL |
sentence-transformers/all-MiniLM-L6-v2 |
384-dim |
HF_LLM_MODEL |
meta-llama/Llama-3.1-8B-Instruct |
Auto-falls-back if unavailable |
NEO4J_URI / NEO4J_USERNAME / NEO4J_PASSWORD |
— | Optional; enables graph context |
VECTOR_BACKEND |
chroma |
chroma (local) · snapshot (deployed) · pinecone |
EMBEDDING_BACKEND |
hf |
hf (API) or local (sentence-transformers) |
OPENFDA_API_KEY |
— | Optional; raises quota 1k → 120k/day |
APP_PASSCODE |
— | If set, the web UI requires this passcode |
RAG_TOP_K / LLM_TEMPERATURE |
5 / 0.0 |
Temperature 0 maximizes reproducibility |
Graph (Neo4j)
(:Drug {set_id, generic_name, brand_names, record_id, effective_time})
-[:MANUFACTURED_BY]-> (:Manufacturer {name})
-[:IN_CLASS]-> (:PharmClass {name})
-[:CONTAINS]-> (:Substance {name})
-[:HAS_ADVERSE_EVENT {report_count}]-> (:AdverseEvent {name}) // FAERSExplore it in the Neo4j console:
MATCH (d:Drug)-[r:HAS_ADVERSE_EVENT]->(a) RETURN d, r, a LIMIT 50Vectors — one per label-section chunk, id {record_id}:{section}:{index},
carrying generic_name, section, source_id, and retrieved_at for lineage.
Records — each DrugLabelRecord embeds a Provenance block (source endpoint,
exact query, openFDA id, UTC timestamp, SHA-256 of the raw payload).
Deploy free on Streamlit Community Cloud with passcode-gated access — see docs/DEPLOYMENT.md for the click-by-click guide.
The system is designed for auditability, not just accuracy:
- Grounding — the LLM may answer only from retrieved passages and must refuse when they are insufficient; every claim carries a citation.
- Lineage — provenance on every record;
source_id+retrieved_aton every vector. - Reproducibility — deterministic ids, idempotent
MERGE,temperature=0. - Qualification — 20 automated tests: OQ (deterministic machinery) and PQ (criteria-based acceptance for non-deterministic AI behaviour).
pytest -m "not integration" # OQ — offline, fast
pytest -m integration # PQ — live servicesFull protocol and traceability matrix: docs/VALIDATION.md.
- Corpus size — ships with ~50 drug labels; increase with
--max-records. - Neo4j AuraDB Free auto-pauses after ~3 days idle. Resume it at console.neo4j.io; graph context degrades gracefully meanwhile.
- HF free-tier model churn — models get retired;
llm.pyauto-falls-back through a candidate list, but the default may need occasional refreshing. - Adverse events are label-derived text plus FAERS counts — FAERS reports are voluntary and do not establish causation, nor are counts incidence rates.
- Answer completeness depends on retrieval depth (
-k/ the UI slider).
See CONTRIBUTING.md. Please run pytest before opening a PR.
MIT. openFDA data is public domain; you are responsible for complying with the openFDA terms of service.