Skip to content

Repository files navigation

💊 PV-Insight — Pharmacovigilance GraphRAG Assistant

Ask questions about FDA drug labels in plain English and get grounded, cited answers — never invented ones. PV-Insight combines semantic retrieval over official openFDA label text with a Neo4j knowledge graph enriched with real FAERS adverse-event report counts.

Built entirely on free and open-source tools and free cloud tiers, and validated with a risk-based OQ/PQ test suite aligned to FDA Computer Software Assurance (CSA) and GAMP-5 principles.

⚠️ Intended use: informational, research, and demonstration purposes only. Not medical advice, and not a validated system of record for regulatory decision-making. Always verify against the official label.


What it does

Question: "What are the most frequently reported adverse events for tramadol, and what does its label warn about?"

Answer:

The most frequently reported adverse events for tramadol are: Dependence (n=7823), Overdose (n=3889), Vomiting (n=3327)… The label warns about the risk of addiction, abuse, and misuse, as well as life-threatening respiratory depression… [1, 2, 4, 5]

Sources: [1] TRAMADOL HYDROCHLORIDE – adverse_reactions · [2] warnings_and_cautions · [4] boxed_warning

Two retrieval paths are fused into one answer:

  1. Semantic search over label text → the exact passages, cited by number.
  2. Knowledge graph → drug class, related drugs, and FAERS report counts.

Architecture

                 ┌──────────── openFDA API ────────────┐
                 │  /drug/label.json   /drug/event.json │
                 └──────────┬──────────────┬────────────┘
                     labels │              │ FAERS counts
                            ▼              ▼
                    normalize + dedup   enrich_faers
                            │              │
                  data/labels.jsonl        │   (immutable staging snapshot)
                            │              │
              ┌─────────────┴───┐          ▼
              ▼                 ▼    ┌──────────────┐
     HF embeddings        (metadata) │ Neo4j AuraDB │  Drug ─ AdverseEvent
              │                      │    graph     │  Manufacturer/Class
              ▼                      └──────┬───────┘
   Chroma (local) ──export──▶ data/vectors.npz      │
                                     │              │
                                     └──────┬───────┘
                                            ▼
                                   GraphRAG orchestrator
                                            │
                                    HF Inference LLM
                                            │
                              Streamlit UI  /  CLI
Layer Technology Cost
Source data openFDA (labels + FAERS) Free, no key required
Embeddings HF Inference API — all-MiniLM-L6-v2 (384-dim) Free tier
Vector search Chroma (local) → .npz snapshot (deployed) Free
Graph Neo4j AuraDB Free Free tier
Generation HF Inference — Llama 3.1 8B Instruct (auto-fallback) Free tier
UI / hosting Streamlit + Community Cloud Free tier

Why two vector backends? Locally, Chroma is writable — needed for ingestion. Cloud hosts have ephemeral filesystems, so the deployed app loads an immutable, pre-built .npz snapshot: no database process, ~4 MB, instant cold start, and exact (not approximate) search. An OQ test asserts both backends return identical rankings and distances.


Quick start

Prerequisites: Python 3.11+ and a free Hugging Face token. Neo4j is optional — without it you lose graph context but still get cited answers.

git clone https://github.com/<your-username>/pv-insight.git
cd pv-insight

python -m venv .venv
source .venv/bin/activate          # Windows: .\.venv\Scripts\Activate.ps1

pip install -r requirements-dev.txt   # dev install (includes Chroma + pytest)
cp .env.example .env                  # Windows: Copy-Item .env.example .env

Edit .env and set at minimum:

HF_API_TOKEN=hf_your_token_here

Optional, for the knowledge graph (AuraDB Free):

NEO4J_URI=neo4j+s://xxxxxxxx.databases.neo4j.io
NEO4J_USERNAME=neo4j
NEO4J_PASSWORD=your-password

Verify your setup, then build the dataset:

python -m pv_insight.readiness        # all components should report OK
python -m pv_insight.ingest --max-records 50   # fetch labels from openFDA
python -m pv_insight.load                      # embed + load graph & vectors
python -m pv_insight.enrich_faers              # add FAERS adverse events (needs Neo4j)

Run it:

streamlit run streamlit_app.py                            # web app
python -m pv_insight.ask "What warnings apply to naproxen?"  # CLI

Command reference

Command Purpose
python -m pv_insight.readiness Health-check every service; UTC audit record. Exit 0 = ready
python -m pv_insight.ingest [--search Q] [--max-records N] Fetch + normalize openFDA labels → data/labels.jsonl
python -m pv_insight.load [--limit N] [--skip-graph] Embed chunks → vector store; MERGE drugs → graph
python -m pv_insight.enrich_faers [--top-events N] Add HAS_ADVERSE_EVENT edges with FAERS counts
python -m pv_insight.export_snapshot Build data/vectors.npz for deployment
python -m pv_insight.search "text" [-k N] Raw semantic search (no LLM)
python -m pv_insight.ask "question" [--no-graph] Full GraphRAG answer with citations
pytest Full OQ + PQ validation suite

All loaders are idempotent — re-running converges instead of duplicating.

Configuration

Set via .env locally, or as Streamlit secrets / environment variables when deployed. See .env.example for the full annotated list.

Variable Default Notes
HF_API_TOKEN — Required. Embeddings + generation
HF_EMBEDDING_MODEL sentence-transformers/all-MiniLM-L6-v2 384-dim
HF_LLM_MODEL meta-llama/Llama-3.1-8B-Instruct Auto-falls-back if unavailable
NEO4J_URI / NEO4J_USERNAME / NEO4J_PASSWORD — Optional; enables graph context
VECTOR_BACKEND chroma chroma (local) · snapshot (deployed) · pinecone
EMBEDDING_BACKEND hf hf (API) or local (sentence-transformers)
OPENFDA_API_KEY — Optional; raises quota 1k → 120k/day
APP_PASSCODE — If set, the web UI requires this passcode
RAG_TOP_K / LLM_TEMPERATURE 5 / 0.0 Temperature 0 maximizes reproducibility

Data model

Graph (Neo4j)

(:Drug {set_id, generic_name, brand_names, record_id, effective_time})
  -[:MANUFACTURED_BY]->  (:Manufacturer {name})
  -[:IN_CLASS]->         (:PharmClass {name})
  -[:CONTAINS]->         (:Substance {name})
  -[:HAS_ADVERSE_EVENT {report_count}]-> (:AdverseEvent {name})   // FAERS

Explore it in the Neo4j console:

MATCH (d:Drug)-[r:HAS_ADVERSE_EVENT]->(a) RETURN d, r, a LIMIT 50

Vectors — one per label-section chunk, id {record_id}:{section}:{index}, carrying generic_name, section, source_id, and retrieved_at for lineage.

Records — each DrugLabelRecord embeds a Provenance block (source endpoint, exact query, openFDA id, UTC timestamp, SHA-256 of the raw payload).

Deployment

Deploy free on Streamlit Community Cloud with passcode-gated access — see docs/DEPLOYMENT.md for the click-by-click guide.

Validation & compliance

The system is designed for auditability, not just accuracy:

  • Grounding — the LLM may answer only from retrieved passages and must refuse when they are insufficient; every claim carries a citation.
  • Lineage — provenance on every record; source_id + retrieved_at on every vector.
  • Reproducibility — deterministic ids, idempotent MERGE, temperature=0.
  • Qualification — 20 automated tests: OQ (deterministic machinery) and PQ (criteria-based acceptance for non-deterministic AI behaviour).
pytest -m "not integration"   # OQ — offline, fast
pytest -m integration         # PQ — live services

Full protocol and traceability matrix: docs/VALIDATION.md.

Known limitations

  • Corpus size — ships with ~50 drug labels; increase with --max-records.
  • Neo4j AuraDB Free auto-pauses after ~3 days idle. Resume it at console.neo4j.io; graph context degrades gracefully meanwhile.
  • HF free-tier model churn — models get retired; llm.py auto-falls-back through a candidate list, but the default may need occasional refreshing.
  • Adverse events are label-derived text plus FAERS counts — FAERS reports are voluntary and do not establish causation, nor are counts incidence rates.
  • Answer completeness depends on retrieval depth (-k / the UI slider).

Contributing

See CONTRIBUTING.md. Please run pytest before opening a PR.

License

MIT. openFDA data is public domain; you are responsible for complying with the openFDA terms of service.

About

Pharmacovigilance GraphRAG assistant: cited, grounded Q&A over FDA drug labels using Neo4j + FAERS adverse-event data. Free-tier stack with risk-based OQ/PQ validation.

Topics

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages