An intelligent document question-answering system built with RAG (Retrieval-Augmented Generation) architecture. Upload documents, ask questions, and get AI-powered answers with source citations. Built with FastAPI, React, ChromaDB, and OpenAI.
- Document Processing: Support for PDF, DOCX, Markdown, and TXT files
- Intelligent Chunking: Automatic text chunking with overlap for better context
- Vector Search: Semantic search using embeddings and ChromaDB
- LLM Integration: Support for OpenAI and local LLM models
- Streaming Responses: Real-time streaming of AI responses
- Source Citations: Track which documents and chunks were used for answers
- Modern UI: Beautiful, responsive React frontend
- RESTful API: FastAPI backend with comprehensive endpoints
- Docker Support: Easy deployment with Docker Compose
- Architecture
- Tech Stack
- Prerequisites
- Installation
- Configuration
- Usage
- API Documentation
- Deployment
- Enterprise
- Security
- Development
- Future Improvements
The system follows a RAG (Retrieval-Augmented Generation) architecture:
โโโโโโโโโโโโโโโ
โ Frontend โ React Application
โ (React) โ - Document Upload
โโโโโโโโฌโโโโโโโ - Chat Interface
โ - Document Management
โ HTTP
โผ
โโโโโโโโโโโโโโโ
โ Backend โ FastAPI Application
โ (FastAPI) โ - REST API
โโโโโโโโฌโโโโโโโ - Document Processing
โ
โโโโโโโโโโโ
โ โ
โผ โผ
โโโโโโโโโโโโ โโโโโโโโโโโโโโโโ
โ Document โ โ Embedding โ
โProcessor โ โ Model โ
โโโโโโฌโโโโโโ โโโโโโโโฌโโโโโโโโ
โ โ
โ โผ
โ โโโโโโโโโโโโโโโโ
โ โ Vector Store โ
โ โ (ChromaDB) โ
โ โโโโโโโโฌโโโโโโโโ
โ โ
โโโโโโโโโโโโโโโโโโ
โ
โผ
โโโโโโโโโโโโโโโโ
โ LLM Service โ
โ (OpenAI/Local)โ
โโโโโโโโโโโโโโโโ
- Document Upload: User uploads a document through the frontend
- Text Extraction: Backend extracts text based on file type (PDF, DOCX, etc.)
- Chunking: Text is split into overlapping chunks for better context
- Embedding: Each chunk is converted to a vector using sentence transformers
- Storage: Embeddings and metadata are stored in ChromaDB vector database
- Query Processing: User question is embedded and used to search similar chunks
- Context Retrieval: Top-K most relevant chunks are retrieved
- Answer Generation: LLM generates answer based on retrieved context
- Response: Answer with source citations is returned to the user
- FastAPI: Modern Python web framework
- sentence-transformers: Embedding generation
- ChromaDB: Vector database for similarity search
- PyPDF2: PDF text extraction
- python-docx: DOCX text extraction
- OpenAI API: LLM integration (optional: local models)
- React: UI framework
- Vite: Build tool and dev server
- Axios: HTTP client
- Modern CSS: Responsive design with gradients
- Docker: Containerization
- Docker Compose: Multi-container orchestration
- Python 3.11+
- Node.js 18+
- Docker and Docker Compose (optional, for containerized deployment)
- OpenAI API key (if using OpenAI LLM)
- Clone the repository:
git clone <repository-url>
cd portproj- Create a
.envfile:
cp .env.example .env
# Edit .env and add your OPENAI_API_KEY- Start the services:
docker-compose up -d- Access the application:
- Frontend: http://localhost:3000
- Backend API (proxied): http://localhost:3000/api
- API Docs: http://localhost:3000/api/docs
- Navigate to backend directory:
cd backend- Create virtual environment:
python -m venv venv
source venv/bin/activate # On Windows: venv\Scripts\activate- Install dependencies:
pip install -r requirements.txt- Set environment variables:
export OPENAI_API_KEY=your_api_key_here
export LLM_PROVIDER=openai
export LLM_MODEL=gpt-3.5-turbo
export EMBEDDING_MODEL=all-MiniLM-L6-v2- Run the server:
uvicorn app.main:app --reload- Navigate to frontend directory:
cd frontend- Install dependencies:
npm install- Start development server:
npm run dev- Access the application at http://localhost:3000
| Variable | Description | Default |
|---|---|---|
OPENAI_API_KEY |
OpenAI API key for LLM | Required |
ENV |
Runtime environment name (development, production, etc.) |
development locally; production in compose |
ENABLE_API_DOCS |
Expose /docs, /redoc, /openapi.json |
false in production by default |
LLM_PROVIDER |
LLM provider: openai or local |
openai |
LLM_MODEL |
Model name (e.g., gpt-3.5-turbo) |
gpt-3.5-turbo |
EMBEDDING_MODEL |
Embedding model name | all-MiniLM-L6-v2 |
VECTOR_STORE_PATH |
Path to ChromaDB storage | ./chroma_db |
VITE_API_URL |
Backend API URL for frontend | /api (recommended behind nginx) |
API_KEY |
Optional API key required for protected backend endpoints | Empty (disabled) |
API_KEYS |
Optional comma-separated additional API keys | Empty |
API_KEY_HEADER_NAME |
Header used for API key auth | X-API-Key |
MAX_UPLOAD_BYTES |
Max upload size in bytes | 10485760 (10MB) |
CORS_ALLOWED_ORIGINS |
Comma-separated frontend origins | http://localhost:3000 |
RATE_LIMIT_ENABLED |
Enable API rate limiting | true |
RATE_LIMIT_DEFAULT |
Default limiter | 200/minute |
RATE_LIMIT_UPLOAD |
Upload endpoint limit | 20/hour |
RATE_LIMIT_ASK |
Ask endpoints limit | 120/hour |
RATE_LIMIT_READ |
Read endpoints limit | 600/hour |
AUTH_MODE |
anonymous, api_key, jwt, jwt_or_api_key |
jwt_or_api_key |
JWT_JWKS_URL |
OIDC JWKS URL (Auth0/Azure AD/Okta/Keycloak) | Empty |
JWT_ISSUER / JWT_AUDIENCE |
Optional JWT validation | Empty |
STATS_ALLOWED_ROLES |
Comma-separated roles allowed to call GET /stats |
Empty (any authenticated caller) |
ENABLE_METRICS |
Expose Prometheus metrics at GET /metrics |
true |
For SSO/OIDC, tenant claims, RBAC on stats, audit JSON logs, and Prometheus metrics, see ENTERPRISE.md.
- PDF (
.pdf) - Microsoft Word (
.docx,.doc) - Markdown (
.md,.markdown) - Plain Text (
.txt)
- Navigate to the "Documents" tab
- Drag and drop a file or click to select
- Wait for processing (document is chunked and embedded)
- Document appears in the list
- Navigate to the "Chat" tab
- Type your question in the input box
- Click "Send" or press Enter
- View the answer with source citations
- Check documents in the Documents tab to limit search scope
- Uncheck to search across all documents
- Click "Stream" button for real-time token streaming
- Useful for longer responses
POST /api/documents/upload- Upload a documentGET /api/documents- List all documentsGET /api/documents/{document_id}- Get document detailsDELETE /api/documents/{document_id}- Delete a document
POST /api/ask- Ask a question (returns full answer)POST /api/ask/stream- Ask a question (streaming response)
GET /api/stats- Get system statisticsGET /api/health- Health check
Note: The stable HTTP contract is /api/v1/*. The UI and nginx dev/proxy setup call /api/*, which is rewritten to /api/v1/*.
curl -X POST "http://localhost:3000/api/documents/upload" \
-H "Content-Type: multipart/form-data" \
-F "file=@document.pdf"curl -X POST "http://localhost:3000/api/ask" \
-H "Content-Type: application/json" \
-H "X-API-Key: <your_api_key_if_enabled>" \
-d '{
"question": "What is this document about?",
"top_k": 5
}'curl "http://localhost:3000/api/stats"API documentation is available at http://localhost:3000/api/docs when ENABLE_API_DOCS=true (disabled by default in the provided compose file).
- Build Docker images:
docker-compose build- Set production environment variables:
# Edit docker-compose.yml or use .env fileRecommended production settings
- Set
API_KEY(and optionallyAPI_KEYS) so the API is not anonymously accessible. - Set
ENV=productionand keepENABLE_API_DOCS=falseunless you explicitly want Swagger exposed. - Keep
CORS_ALLOWED_ORIGINSlimited to your real frontend origin(s). - Prefer same-origin API calls (
VITE_API_URL=/api) and avoid baking secrets into the frontend bundle.
- Run in production mode:
docker-compose up -d- Deploy backend to AWS ECS or EC2
- Use RDS or S3 for vector store persistence
- Deploy frontend to S3 + CloudFront
- Deploy backend to Cloud Run
- Use Cloud Storage for vector store
- Deploy frontend to Firebase Hosting
- Deploy backend to Azure Container Instances
- Use Azure Blob Storage for vector store
- Deploy frontend to Azure Static Web Apps
- Vector Store: Consider using managed vector databases (Pinecone, Weaviate) for production
- LLM: Use API keys from environment variables, never commit them
- Scaling: Use load balancers and multiple backend instances
- Monitoring: Add logging and monitoring (e.g., Prometheus, Grafana)
portproj/
โโโ backend/
โ โโโ app/
โ โ โโโ api/ # API routes
โ โ โโโ models/ # ML models
โ โ โโโ services/ # Business logic
โ โ โโโ utils/ # Utilities
โ โโโ requirements.txt
โ โโโ Dockerfile
โโโ frontend/
โ โโโ src/
โ โ โโโ components/ # React components
โ โ โโโ pages/ # Page components
โ โ โโโ services/ # API clients
โ โโโ package.json
โ โโโ Dockerfile
โโโ ml/
โ โโโ models/ # Model files
โ โโโ embeddings/ # Embedding scripts
โ โโโ notebooks/ # Jupyter notebooks
โโโ docker-compose.yml
โโโ README.md
# Backend tests (when implemented)
cd backend
pytest
# Frontend tests (when implemented)
cd frontend
npm test- Backend: Follow PEP 8, use type hints
- Frontend: Follow ESLint rules, use functional components
- Add user authentication and multi-user support (API key tenant isolation is implemented; user accounts are not)
- Implement document versioning
- Add support for more file types (Excel, PowerPoint)
- Improve chunking strategies (semantic chunking)
- Add conversation history persistence
- Implement document search/filtering
- Support for multiple LLM providers (Anthropic, Cohere)
- Fine-tuned embedding models
- Advanced RAG techniques (re-ranking, query expansion)
- Export conversations to PDF/Markdown
- Batch document processing
- API rate limiting and API-key authentication
- Multi-modal support (images, tables in documents)
- Real-time collaboration features
- Advanced analytics and insights
- Custom model fine-tuning interface
- Integration with cloud storage (S3, Google Drive)
- Mobile app support
- Embedding Generation: ~100-200ms per chunk (CPU)
- Vector Search: ~10-50ms for top-5 results
- LLM Response: ~1-3s (OpenAI GPT-3.5-turbo)
- Document Processing: ~1-5s per document (depends on size)
- Fork the repository
- Create a feature branch (
git checkout -b feature/amazing-feature) - Commit your changes (
git commit -m 'Add amazing feature') - Push to the branch (
git push origin feature/amazing-feature) - Open a Pull Request
RAG, Retrieval-Augmented Generation, Document Q&A, Document Question Answering, AI Document Assistant, LLM Integration, Vector Database, ChromaDB, FastAPI, React, OpenAI, Embeddings, Semantic Search, Document Processing, PDF Q&A, Document Chatbot, RAG Implementation, AI-Powered Search, Document Intelligence, Natural Language Processing, NLP, Machine Learning, ML Portfolio Project
This project is open source and available under the MIT License.
- Built with FastAPI
- Embeddings powered by sentence-transformers
- Vector database: ChromaDB
- UI framework: React
Note: This is a portfolio project demonstrating RAG architecture, full-stack development, and ML integration. For production use, consider additional security, scalability, and monitoring features.