Skip to content

Repository files navigation

Data Integrity Case File

CI CodeQL Status Python FastAPI React SQLite Ollama Compliance Docker

ALCOA+ · Data Integrity · Investigation · CAPA Readiness · Audit Evidence · Local AI Triage

Portfolio-safe prototype — synthetic data only

Screenshots · Quick start · Case study · Investigation playbook · AI Assistant · Security


Data boundary: All case files and evidence notes are fictional. This is an educational workspace and must not be used for actual regulatory filings, batch releases, or official QA investigations.


Portfolio preview

Investigation summary Synthetic case library
Synthetic data-integrity investigation metrics Synthetic data-integrity case library

See the case study for the business problem, users, decisions, evidence, and production boundary.

What This Is

An investigative workspace modeling how quality and IT compliance professionals structure data integrity investigations under ALCOA+ and 21 CFR Part 11:

  1. Intake — Capture the signal (audit finding, system discrepancy, user access anomaly)
  2. ALCOA+ gap analysis — Evaluate affected attributes: Attributable, Legible, Contemporaneous, Original, Accurate, Complete, Consistent, Enduring, Available
  3. Evidence log — Record audit trail reviews, technical metadata, access logs, screenshots, interview notes
  4. Root cause & CAPA formulation — Structure corrective and preventive action items
  5. AI-assisted triage — Local LLM (Ollama + llama3.2:3b) suggests which ALCOA+ attributes to prioritize — human review required for every suggestion
  6. Workflow lifecycle — Track case stages from intake through QA closure

Designed to demonstrate fluency in FDA, MHRA, and PIC/S data integrity frameworks through clean software architecture.

Stack: Python 3.11 · FastAPI 0.141 · Pydantic v2 · SQLite (WAL) · React 19 · Docker Compose · Ollama (local AI) · GitHub Actions


Security (local demo)

This is not production IAM. Controls are intentional for a portfolio prototype:

Control Behavior
API key Header X-API-Key required on all routes except /health, /docs, /redoc, /openapi.json
Default key dev-api-key-change-me — override with env API_KEY
Rate limit 120 req/min per client (default); AI suggest endpoint 10 req/min
Security headers X-Frame-Options, X-Content-Type-Options, CSP, Referrer-Policy
CORS Allowlist: localhost:5173 / 127.0.0.1:5173
Audit actor Mutations still require X-Actor for the application audit log
Ollama Published only on 127.0.0.1:11434 on the host
Container API image runs as non-root user appuser
# Examples
curl http://localhost:8000/health
curl -H "X-API-Key: dev-api-key-change-me" http://localhost:8000/summary
curl -H "X-API-Key: $API_KEY" -H "X-Actor: A.Reyes" \
  -H "Content-Type: application/json" \
  -d '{"title":"Demo case","system":"LIMS-01","signal_type":"audit_finding","opened_by":"A.Reyes"}' \
  http://localhost:8000/cases

Env vars (API / Compose): API_KEY, RATE_LIMIT_DEFAULT, RATE_LIMIT_WINDOW_SECONDS, RATE_LIMIT_AI, RATE_LIMIT_AI_WINDOW_SECONDS.


Quick Start

Option A — Docker Compose (recommended)

# 1. Clone
git clone https://github.com/alianisreyesr/data-integrity-case-file.git
cd data-integrity-case-file

# 2. Optional: set a non-default API key
export API_KEY=dev-api-key-change-me

# 3. Start all services (API + frontend + Ollama)
docker compose up --build

# 4. Verify (health is public; other routes need the key)
curl http://localhost:8000/health
curl -H "X-API-Key: $API_KEY" http://localhost:8000/ai/status

Ollama note: Model pull runs via the ollama-init one-shot service. Data persists in the ollama_data volume.

Option B — Local Python (no Docker)

# 1. Prerequisites: Python 3.11+, Ollama installed (https://ollama.com)
ollama pull llama3.2:3b

# 2. Install dependencies
python -m venv .venv && source .venv/bin/activate
pip install -r requirements-dev.txt

# 3. Seed synthetic data
python data/seed.py

# 4. Run (optional: export API_KEY=...)
uvicorn app.main:app --reload
# API: http://localhost:8000
# Docs: http://localhost:8000/docs

API Endpoints

All endpoints below except GET /health require X-API-Key. Write operations also require X-Actor.

Core Workflow

Method Path Description
GET /health Service health + data boundary reminder (public)
GET /summary Case counts, open gaps, CAPA stats
GET /cases List all cases (filter by ?status=)
POST /cases Open a new DI case
GET /cases/{id} Get case detail
GET /cases/{id}/alcoa-gaps List ALCOA+ gap assessments
POST /cases/{id}/alcoa-gaps Record a gap finding
GET /cases/{id}/evidence List evidence entries
POST /cases/{id}/evidence Add an evidence record
GET /cases/{id}/capas List CAPAs for a case
POST /cases/{id}/capas Create a CAPA item
GET /audit-log Full audit log (filter by ?case_id=)

AI-Assisted Triage (Human Review Required)

Method Path Description
GET /ai/status Ollama service + model availability check
POST /cases/{id}/ai-suggest-gaps Generate ALCOA+ gap suggestions (local LLM; stricter rate limit)
GET /cases/{id}/ai-suggestions List all AI suggestions for a case
POST /ai-suggestions/{id}/review Accept / reject / modify a suggestion (one-time; accepting writes ALCOA+ gaps)

Interactive docs: http://localhost:8000/docs (Swagger UI)


AI Layer

The local AI assistant uses llama3.2:3b via Ollama running entirely on your machine — no data leaves your environment.

flowchart TB
  A["User opens a case"] --> B["POST /ai-suggest-gaps"]
  B --> C["ai.py calls Ollama /api/chat"]
  C --> D["Model returns structured JSON suggestions"]
  D --> E["Hash response with SHA-256 and store it"]
  E --> F["Qualified human accepts, rejects, or modifies the suggestion set"]
  F --> G["Accept: write one alcoa_gaps row per attribute + audit_log entry.\nReject/modify: audit_log entry only, no gap written."]
Loading

The AI never writes directly to case records. Every suggestion requires explicit human action before any gap is recorded, a suggestion can only be reviewed once, and the SHA-256 is recomputed and checked against the stored response on every read — not just recorded once and forgotten — so a mismatch is surfaced (integrity_verified: false) instead of silently trusted.

See docs/AI_ASSISTANT.md for the full design rationale.


Project Structure

flowchart TB
  R["data-integrity-case-file"]
  R --> A["app — FastAPI, security, database, AI, and review modules"]
  A --> AA["main.py and router.py — application and REST endpoints"]
  A --> AB["security.py and database.py — controls and SQLite persistence"]
  A --> AC["ai modules — local inference, readiness, and human review"]
  R --> D["data — synthetic case seeder"]
  R --> T["tests — API, security, AI, contract, and status tests"]
  R --> O["docs — AI design, references, roadmap, and frontend audit"]
  R --> F["frontend — React investigation board"]
  R --> P["Docker and Python dependency files"]
Loading

Roadmap

Phase Milestone Status
Phase 0 Documentation & ALCOA+ regulatory map ✅ Complete
Phase 1 Case, Finding, Evidence & CAPA domain models + AI layer ✅ Complete
Phase 2 FastAPI endpoints & synthetic case library ✅ Complete
Phase 2.1 API key, rate limits, security headers, non-root Docker ✅ Complete
Phase 3 Investigation board & detail reviewer UI (React 19) ✅ Complete
Phase 4 Extended test suite, CodeQL scanning & hardened Docker delivery ✅ Complete

Regulated Portfolio Ecosystem

Project Domain Focus Status
Quality Deviation Risk Monitor Deviation prioritization & explainable risk scoring ✅ Active · 112 tests
CSV Evidence Tracker Requirements traceability, IQ/OQ/PQ test execution, audit trail ✅ Active · 44 tests
GxP Change Control Controlled change lifecycle & approvals ✅ Active · 68 tests
CSA Assurance Planner Risk-based software assurance planning, FDA CSA alignment ✅ Active
GxP Batch Data Pipeline Batch manufacturing pipeline — DuckDB · dbt · quality gates ✅ Active · 12 tests

Running Tests

# Unit + contract + security tests (no Ollama required — Ollama calls are mocked)
export API_KEY=test-api-key   # tests also set this internally
pytest tests/ -v

# With coverage
pytest tests/ --cov=app --cov-report=term-missing

AI tests mock the Ollama HTTP call so the full test suite runs offline.


Built by Alianis Reyes-Reyes · LinkedIn · Portfolio

Information Systems @ UPRM · Eli Lilly Tech@Lilly Alumni

About

Portfolio-safe data integrity investigation workspace — ALCOA+ gap analysis, evidence, CAPA readiness. Synthetic data only. Python · FastAPI · React.

Topics

Resources

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages