Skip to content
forestByTheSeashorePublic

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

1 watching

Forks

Latest commit

 

History

5 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 

Repository files navigation

RAG LLM - Retrieval-Augmented Generation with Large Language Models

A powerful RAG (Retrieval-Augmented Generation) application that combines large language models with document retrieval capabilities for intelligent report generation and conversational AI.

Overview

This project implements a FastAPI-based service that integrates:

  • Large Language Models (LLMs) - LLaMA 3.1 models (70B parameter versions)
  • Vector Database - FAISS for efficient similarity search
  • Text Embeddings - HuggingFace embeddings for semantic understanding
  • LangChain - For orchestrating RAG workflows

The system can handle both conversational chat and structured report generation tasks, automatically retrieving relevant context from a knowledge base to enhance response quality.

Features

Core Capabilities

  • Intelligent Chat API - Context-aware conversational interface
  • Structured Report Generation - Automatic generation of three-part reports (What, Why, How)
  • Streaming Responses - Real-time token-by-token response streaming
  • RAG Integration - Retrieves relevant documents to augment LLM responses
  • Chinese Language Optimized - Specialized support for Chinese text processing

Technical Features

  • CORS Support - Ready for web application integration
  • Request Validation - Pydantic models for type safety
  • Comprehensive Logging - Detailed logging for monitoring and debugging
  • Error Handling - Robust exception handling and recovery
  • Unicode Normalization - Proper handling of Chinese characters

Architecture

Components

  1. FastAPI Server - RESTful API endpoints for chat and status
  2. LLM Engine - LLaMA 3.1 model with 4-bit/8-bit quantization
  3. Vector Store - FAISS database for document embeddings
  4. RAG Chain - Custom LangChain implementation for report generation
  5. Document Processor - Text splitting and embedding generation

Data Flow

User Request → API Endpoint → Task Detection
                                    ↓
                    ┌───────────────┴────────────────┐
                    ↓                                ↓
            Report Generation                Regular Chat
                    ↓                                ↓
          Topic Extraction                  Prompt Construction
                    ↓                                ↓
          Vector Retrieval                     LLM Generation
                    ↓                                ↓
          Context Augmentation                 Token Streaming
                    ↓                                ↓
          Structured Generation → Response → User

Installation

Prerequisites

  • Python 3.8+
  • CUDA-compatible GPU (recommended for model inference)
  • 32GB+ RAM (for 70B models)

Dependencies

pip install fastapi uvicorn
pip install torch transformers
pip install langchain langchain-community
pip install faiss-cpu  # or faiss-gpu for GPU support
pip install sentence-transformers
pip install pydantic

Setup

  1. Clone the repository
git clone https://github.com/forestByTheSeashore/rag_llm.git
cd rag_llm
  1. Configure model paths

Edit the model paths in rag-demo.py or rag-demo-V2.py:

MODEL_PATH = "/path/to/your/llama-model"
EMBEDDING_MODEL_PATH = "/path/to/embedding-model"
DATA_PATH = "/path/to/your/data.json"
  1. Prepare your data

Create a JSON file with the following structure:

[
  {
    "article_id": "1",
    "title": "Article Title",
    "content": "Article content...",
    "source": "Source name",
    "publishtime": "2024-01-01"
  }
]

Usage

Starting the Server

Using rag-demo.py (LLaMA 3.1 4-bit quantized):

python rag-demo.py

Using rag-demo-V2.py (LLaMA 3.1 8-bit Chinese optimized):

python rag-demo-V2.py

The API server will start on http://0.0.0.0:8000

API Endpoints

Health Check

GET /

Response:

{
  "status": "API is running",
  "model": "llama3.1-4bit"
}

Chat API

POST /api/chat

Request Body:

{
  "model": "llama3.1-4bit",
  "messages": [
    {
      "role": "user",
      "content": "Your question here"
    }
  ]
}

Response (Streaming):

{
  "model": "llama3.1-4bit",
  "created_at": "2024-01-01T12:00:00",
  "message": {
    "role": "assistant",
    "content": "Response content..."
  },
  "done": false
}

Examples

Regular Chat

import requests

url = "http://localhost:8000/api/chat"
data = {
    "model": "llama3.1-4bit",
    "messages": [
        {"role": "user", "content": "What is artificial intelligence?"}
    ]
}

response = requests.post(url, json=data, stream=True)
for line in response.iter_lines():
    if line:
        print(line.decode('utf-8'))

Report Generation (Chinese)

data = {
    "model": "llama3.1-4bit",
    "messages": [
        {"role": "user", "content": "生成一份关于人工智能的报告,1500字"}
    ]
}

The system automatically detects report generation tasks and creates structured reports with three sections:

  • 是什么 (What is it) - Definition and characteristics
  • 为什么 (Why) - Importance and impact
  • 怎么做 (How) - Implementation strategies

File Descriptions

rag-demo.py

  • Uses LLaMA 3.1 70B model with 4-bit quantization
  • Temperature: 0.7
  • Optimized for balanced response quality and speed

rag-demo-V2.py

  • Uses LLaMA 3.1 70B Chinese-optimized model
  • 8-bit quantization for better quality
  • Temperature: 0.4, Top-p: 0.9
  • Enhanced Unicode handling for Chinese text
  • More conservative sampling parameters

Configuration

Model Parameters

Parameter rag-demo.py rag-demo-V2.py
Quantization 4-bit 8-bit
Temperature 0.7 0.4
Top-p - 0.9
Max Tokens 1500 1500
Do Sample Yes Yes

RAG Configuration

  • Chunk Size: 1000 characters
  • Chunk Overlap: 0
  • Retrieval K: 3 documents
  • Report Sections: 3 (What, Why, How)
  • Word Distribution: [0.2, 0.4, 0.4]

Advanced Features

Custom Report Generation

The system automatically detects report keywords:

  • 报告 (report)
  • 总结 (summary)
  • 概述 (overview)
  • 分析 (analysis)

Context Augmentation

For report generation:

  1. Extracts topic and word count from user message
  2. Retrieves relevant documents for each section
  3. Filters sentences containing the topic
  4. Constructs section-specific prompts with context
  5. Generates structured content token-by-token

Message Cleaning

Automatic cleaning of:

  • Duplicate assistant tags
  • Special characters
  • Invalid Unicode sequences
  • Non-printable characters

Performance Considerations

  • Memory: 70B models require significant GPU memory
  • Quantization: 4-bit uses ~35GB, 8-bit uses ~70GB
  • Inference Speed: Depends on hardware; expect 5-10 tokens/sec on consumer GPUs
  • FAISS Indexing: One-time cost at startup; fast retrieval thereafter

Troubleshooting

Common Issues

  1. Out of Memory

    • Use 4-bit quantization
    • Reduce batch size
    • Enable CPU offloading with device_map="auto"
  2. Slow Generation

    • Check GPU utilization
    • Reduce max_new_tokens
    • Consider smaller models
  3. Unicode Errors

    • Use rag-demo-V2.py for better Chinese support
    • Check JSON file encoding (must be UTF-8)
  4. Model Loading Fails

    • Verify MODEL_PATH is correct
    • Ensure sufficient disk space
    • Check CUDA compatibility

Contributing

Contributions are welcome! Please feel free to submit pull requests or open issues for bugs and feature requests.

License

This project is provided as-is for educational and research purposes.

Acknowledgments

  • LLaMA 3.1 by Meta AI
  • LangChain framework
  • FAISS by Facebook Research
  • HuggingFace Transformers and Embeddings
  • FastAPI framework

Contact

For questions or issues, please open an issue on the GitHub repository.

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages