A powerful RAG (Retrieval-Augmented Generation) application that combines large language models with document retrieval capabilities for intelligent report generation and conversational AI.
This project implements a FastAPI-based service that integrates:
- Large Language Models (LLMs) - LLaMA 3.1 models (70B parameter versions)
- Vector Database - FAISS for efficient similarity search
- Text Embeddings - HuggingFace embeddings for semantic understanding
- LangChain - For orchestrating RAG workflows
The system can handle both conversational chat and structured report generation tasks, automatically retrieving relevant context from a knowledge base to enhance response quality.
- Intelligent Chat API - Context-aware conversational interface
- Structured Report Generation - Automatic generation of three-part reports (What, Why, How)
- Streaming Responses - Real-time token-by-token response streaming
- RAG Integration - Retrieves relevant documents to augment LLM responses
- Chinese Language Optimized - Specialized support for Chinese text processing
- CORS Support - Ready for web application integration
- Request Validation - Pydantic models for type safety
- Comprehensive Logging - Detailed logging for monitoring and debugging
- Error Handling - Robust exception handling and recovery
- Unicode Normalization - Proper handling of Chinese characters
- FastAPI Server - RESTful API endpoints for chat and status
- LLM Engine - LLaMA 3.1 model with 4-bit/8-bit quantization
- Vector Store - FAISS database for document embeddings
- RAG Chain - Custom LangChain implementation for report generation
- Document Processor - Text splitting and embedding generation
User Request → API Endpoint → Task Detection
↓
┌───────────────┴────────────────┐
↓ ↓
Report Generation Regular Chat
↓ ↓
Topic Extraction Prompt Construction
↓ ↓
Vector Retrieval LLM Generation
↓ ↓
Context Augmentation Token Streaming
↓ ↓
Structured Generation → Response → User
- Python 3.8+
- CUDA-compatible GPU (recommended for model inference)
- 32GB+ RAM (for 70B models)
pip install fastapi uvicorn
pip install torch transformers
pip install langchain langchain-community
pip install faiss-cpu # or faiss-gpu for GPU support
pip install sentence-transformers
pip install pydantic- Clone the repository
git clone https://github.com/forestByTheSeashore/rag_llm.git
cd rag_llm- Configure model paths
Edit the model paths in rag-demo.py or rag-demo-V2.py:
MODEL_PATH = "/path/to/your/llama-model"
EMBEDDING_MODEL_PATH = "/path/to/embedding-model"
DATA_PATH = "/path/to/your/data.json"- Prepare your data
Create a JSON file with the following structure:
[
{
"article_id": "1",
"title": "Article Title",
"content": "Article content...",
"source": "Source name",
"publishtime": "2024-01-01"
}
]Using rag-demo.py (LLaMA 3.1 4-bit quantized):
python rag-demo.pyUsing rag-demo-V2.py (LLaMA 3.1 8-bit Chinese optimized):
python rag-demo-V2.pyThe API server will start on http://0.0.0.0:8000
GET /Response:
{
"status": "API is running",
"model": "llama3.1-4bit"
}POST /api/chatRequest Body:
{
"model": "llama3.1-4bit",
"messages": [
{
"role": "user",
"content": "Your question here"
}
]
}Response (Streaming):
{
"model": "llama3.1-4bit",
"created_at": "2024-01-01T12:00:00",
"message": {
"role": "assistant",
"content": "Response content..."
},
"done": false
}import requests
url = "http://localhost:8000/api/chat"
data = {
"model": "llama3.1-4bit",
"messages": [
{"role": "user", "content": "What is artificial intelligence?"}
]
}
response = requests.post(url, json=data, stream=True)
for line in response.iter_lines():
if line:
print(line.decode('utf-8'))data = {
"model": "llama3.1-4bit",
"messages": [
{"role": "user", "content": "生成一份关于人工智能的报告,1500字"}
]
}The system automatically detects report generation tasks and creates structured reports with three sections:
- 是什么 (What is it) - Definition and characteristics
- 为什么 (Why) - Importance and impact
- 怎么做 (How) - Implementation strategies
- Uses LLaMA 3.1 70B model with 4-bit quantization
- Temperature: 0.7
- Optimized for balanced response quality and speed
- Uses LLaMA 3.1 70B Chinese-optimized model
- 8-bit quantization for better quality
- Temperature: 0.4, Top-p: 0.9
- Enhanced Unicode handling for Chinese text
- More conservative sampling parameters
| Parameter | rag-demo.py | rag-demo-V2.py |
|---|---|---|
| Quantization | 4-bit | 8-bit |
| Temperature | 0.7 | 0.4 |
| Top-p | - | 0.9 |
| Max Tokens | 1500 | 1500 |
| Do Sample | Yes | Yes |
- Chunk Size: 1000 characters
- Chunk Overlap: 0
- Retrieval K: 3 documents
- Report Sections: 3 (What, Why, How)
- Word Distribution: [0.2, 0.4, 0.4]
The system automatically detects report keywords:
- 报告 (report)
- 总结 (summary)
- 概述 (overview)
- 分析 (analysis)
For report generation:
- Extracts topic and word count from user message
- Retrieves relevant documents for each section
- Filters sentences containing the topic
- Constructs section-specific prompts with context
- Generates structured content token-by-token
Automatic cleaning of:
- Duplicate assistant tags
- Special characters
- Invalid Unicode sequences
- Non-printable characters
- Memory: 70B models require significant GPU memory
- Quantization: 4-bit uses ~35GB, 8-bit uses ~70GB
- Inference Speed: Depends on hardware; expect 5-10 tokens/sec on consumer GPUs
- FAISS Indexing: One-time cost at startup; fast retrieval thereafter
-
Out of Memory
- Use 4-bit quantization
- Reduce batch size
- Enable CPU offloading with
device_map="auto"
-
Slow Generation
- Check GPU utilization
- Reduce max_new_tokens
- Consider smaller models
-
Unicode Errors
- Use rag-demo-V2.py for better Chinese support
- Check JSON file encoding (must be UTF-8)
-
Model Loading Fails
- Verify MODEL_PATH is correct
- Ensure sufficient disk space
- Check CUDA compatibility
Contributions are welcome! Please feel free to submit pull requests or open issues for bugs and feature requests.
This project is provided as-is for educational and research purposes.
- LLaMA 3.1 by Meta AI
- LangChain framework
- FAISS by Facebook Research
- HuggingFace Transformers and Embeddings
- FastAPI framework
For questions or issues, please open an issue on the GitHub repository.