Skip to content

Repository files navigation

Cloud-Based Object Detection and Flow Analysis

Getting Started

Prerequisites

  • Python 3.8+
  • Docker and Docker Compose
  • PostgreSQL (if running database locally)
  • Node.js 16+ and npm (for frontend)

Installation

  1. Clone the repository:

    git clone https://github.com/yourusername/cloud-object-detection.git
    cd cloud-object-detection
  2. Install Python dependencies:

    pip install -r requirements.txt
  3. Set up environment variables:

    cp .env.example .env
    # Edit .env file with your credentials and configuration
  4. Run the system:

    docker-compose up -d

Table of Contents

  1. Project Overview
  2. Architecture
  3. Technology Stack
  4. Implementation Details
  5. Data Flow
  6. Cloud Services Comparison
  7. Performance Metrics
  8. Testing and Validation
  9. Assumptions Made
  10. Strengths and Weaknesses
  11. Research and References
  12. Future Improvements

Project Overview

This project analyzes cloud-based solutions for object detection and counting to determine the flow of people and vehicles using Computer Vision and MLOps techniques. The goal is to compare cost, quality, and performance across AWS, Azure, and Local Processing for implementing these solutions.

Key Objectives:

  • Implement object detection models (YOLO) to count people and vehicles
  • Process data on cloud computing services (AWS, Azure) and evaluate cost, quality, and performance
  • Integrate MLOps practices to ensure scalability and standardization
  • Develop a test application to display metrics and visualizations of processed data
  • Create an image processing pipeline using Python and frameworks like TensorFlow, PyTorch, or OpenCV
  • Define preprocessing and image enhancement strategies to optimize detection
  • Establish evaluation metrics to compare cost, quality, and performance across services

Architecture

The system follows a microservices architecture with clear separation between:

  1. ML Services: Handles object detection and tracking using local and cloud-based implementations
  2. Backend API: Provides RESTful endpoints for data retrieval and processing
  3. Frontend Application: Displays metrics, visualizations, and comparison data
  4. MLOps Pipeline: Tracks experiments, model performance, and service metrics
# Cloud-Based Object Detection and Flow Analysis

## Getting Started

### Prerequisites

- Python 3.8+
- Docker and Docker Compose
- PostgreSQL (if running database locally)
- Node.js 16+ and npm (for frontend)

### Installation

1. Clone the repository:
   ```bash
   git clone https://github.com/yourusername/cloud-object-detection.git
   cd cloud-object-detection
  1. Install Python dependencies:

    pip install -r requirements.txt
  2. Set up environment variables:

    cp .env.example .env
    # Edit .env file with your credentials and configuration
  3. Run the system:

    docker-compose up -d

Table of Contents

  1. Project Overview
  2. Architecture
  3. Technology Stack
  4. Implementation Details
  5. Data Flow
  6. Cloud Services Comparison
  7. Performance Metrics
  8. Testing and Validation
  9. Assumptions Made
  10. Strengths and Weaknesses
  11. Research and References
  12. Future Improvements

Project Overview

This project analyzes cloud-based solutions for object detection and counting to determine the flow of people and vehicles using Computer Vision and MLOps techniques. The goal is to compare cost, quality, and performance across AWS, Azure, and Local Processing for implementing these solutions.

Key Objectives:

  • Implement object detection models (YOLO) to count people and vehicles
  • Process data on cloud computing services (AWS, Azure) and evaluate cost, quality, and performance
  • Integrate MLOps practices to ensure scalability and standardization
  • Develop a test application to display metrics and visualizations of processed data
  • Create an image processing pipeline using Python and frameworks like TensorFlow, PyTorch, or OpenCV
  • Define preprocessing and image enhancement strategies to optimize detection
  • Establish evaluation metrics to compare cost, quality, and performance across services

Architecture

The system follows a microservices architecture with clear separation between:

  1. ML Services: Handles object detection and tracking using local and cloud-based implementations
  2. Backend API: Provides RESTful endpoints for data retrieval and processing
  3. Frontend Application: Displays metrics, visualizations, and comparison data
  4. MLOps Pipeline: Tracks experiments, model performance, and service metrics
# Cloud-Based Object Detection and Flow Analysis

## Getting Started

### Prerequisites

- Python 3.8+
- Docker and Docker Compose
- PostgreSQL (if running database locally)
- Node.js 16+ and npm (for frontend)

### Installation

1. Clone the repository:
   ```bash
   git clone https://github.com/yourusername/cloud-object-detection.git
   cd cloud-object-detection
  1. Install Python dependencies:

    pip install -r requirements.txt
  2. Set up environment variables:

    cp .env.example .env
    # Edit .env file with your credentials and configuration
  3. Run the system:

    docker-compose up -d

Table of Contents

  1. Project Overview
  2. Architecture
  3. Technology Stack
  4. Implementation Details
  5. Data Flow
  6. Cloud Services Comparison
  7. Performance Metrics
  8. Testing and Validation
  9. Assumptions Made
  10. Strengths and Weaknesses
  11. Research and References
  12. Future Improvements

Project Overview

This project analyzes cloud-based solutions for object detection and counting to determine the flow of people and vehicles using Computer Vision and MLOps techniques. The goal is to compare cost, quality, and performance across AWS, Azure, and Local Processing for implementing these solutions.

Key Objectives:

  • Implement object detection models (YOLO) to count people and vehicles
  • Process data on cloud computing services (AWS, Azure) and evaluate cost, quality, and performance
  • Integrate MLOps practices to ensure scalability and standardization
  • Develop a test application to display metrics and visualizations of processed data
  • Create an image processing pipeline using Python and frameworks like TensorFlow, PyTorch, or OpenCV
  • Define preprocessing and image enhancement strategies to optimize detection
  • Establish evaluation metrics to compare cost, quality, and performance across services

Architecture

The system follows a microservices architecture with clear separation between:

  1. ML Services: Handles object detection and tracking using local and cloud-based implementations
  2. Backend API: Provides RESTful endpoints for data retrieval and processing
  3. Frontend Application: Displays metrics, visualizations, and comparison data
  4. MLOps Pipeline: Tracks experiments, model performance, and service metrics
# Cloud-Based Object Detection and Flow Analysis

## Getting Started

### Prerequisites

- Python 3.8+
- Docker and Docker Compose
- PostgreSQL (if running database locally)
- Node.js 16+ and npm (for frontend)

### Installation

1. Clone the repository:
   ```bash
   git clone https://github.com/yourusername/cloud-object-detection.git
   cd cloud-object-detection
  1. Install Python dependencies:

    pip install -r requirements.txt
  2. Set up environment variables:

    cp .env.example .env
    # Edit .env file with your credentials and configuration
  3. Run the system:

    docker-compose up -d

Table of Contents

  1. Project Overview
  2. Architecture
  3. Technology Stack
  4. Implementation Details
  5. Data Flow
  6. Cloud Services Comparison
  7. Performance Metrics
  8. Testing and Validation
  9. Assumptions Made
  10. Strengths and Weaknesses
  11. Research and References
  12. Future Improvements

Project Overview

This project analyzes cloud-based solutions for object detection and counting to determine the flow of people and vehicles using Computer Vision and MLOps techniques. The goal is to compare cost, quality, and performance across AWS, Azure, and Local Processing for implementing these solutions.

Key Objectives:

  • Implement object detection models (YOLO) to count people and vehicles
  • Process data on cloud computing services (AWS, Azure) and evaluate cost, quality, and performance
  • Integrate MLOps practices to ensure scalability and standardization
  • Develop a test application to display metrics and visualizations of processed data
  • Create an image processing pipeline using Python and frameworks like TensorFlow, PyTorch, or OpenCV
  • Define preprocessing and image enhancement strategies to optimize detection
  • Establish evaluation metrics to compare cost, quality, and performance across services

Architecture

The system follows a microservices architecture with clear separation between:

  1. ML Services: Handles object detection and tracking using local and cloud-based implementations
  2. Backend API: Provides RESTful endpoints for data retrieval and processing
  3. Frontend Application: Displays metrics, visualizations, and comparison data
  4. MLOps Pipeline: Tracks experiments, model performance, and service metrics
# Cloud-Based Object Detection and Flow Analysis

## Getting Started

### Prerequisites

- Python 3.8+
- Docker and Docker Compose
- PostgreSQL (if running database locally)
- Node.js 16+ and npm (for frontend)

### Installation

1. Clone the repository:
   ```bash
   git clone https://github.com/yourusername/cloud-object-detection.git
   cd cloud-object-detection
  1. Install Python dependencies:

    pip install -r requirements.txt
  2. Set up environment variables:

    cp .env.example .env
    # Edit .env file with your credentials and configuration
  3. Run the system:

    docker-compose up -d

Table of Contents

  1. Project Overview
  2. Architecture
  3. Technology Stack
  4. Implementation Details
  5. Data Flow
  6. Cloud Services Comparison
  7. Performance Metrics
  8. Testing and Validation
  9. Assumptions Made
  10. Strengths and Weaknesses
  11. Research and References
  12. Future Improvements

Project Overview

This project analyzes cloud-based solutions for object detection and counting to determine the flow of people and vehicles using Computer Vision and MLOps techniques. The goal is to compare cost, quality, and performance across AWS, Azure, and Local Processing for implementing these solutions.

Key Objectives:

  • Implement object detection models (YOLO) to count people and vehicles
  • Process data on cloud computing services (AWS, Azure) and evaluate cost, quality, and performance
  • Integrate MLOps practices to ensure scalability and standardization
  • Develop a test application to display metrics and visualizations of processed data
  • Create an image processing pipeline using Python and frameworks like TensorFlow, PyTorch, or OpenCV
  • Define preprocessing and image enhancement strategies to optimize detection
  • Establish evaluation metrics to compare cost, quality, and performance across services

Architecture

The system follows a microservices architecture with clear separation between:

  1. ML Services: Handles object detection and tracking using local and cloud-based implementations
  2. Backend API: Provides RESTful endpoints for data retrieval and processing
  3. Frontend Application: Displays metrics, visualizations, and comparison data
  4. MLOps Pipeline: Tracks experiments, model performance, and service metrics
# Cloud-Based Object Detection and Flow Analysis

## Getting Started

### Prerequisites

- Python 3.8+
- Docker and Docker Compose
- PostgreSQL (if running database locally)
- Node.js 16+ and npm (for frontend)

### Installation

1. Clone the repository:
   ```bash
   git clone https://github.com/yourusername/cloud-object-detection.git
   cd cloud-object-detection
  1. Install Python dependencies:

    pip install -r requirements.txt
  2. Set up environment variables:

    cp .env.example .env
    # Edit .env file with your credentials and configuration
  3. Run the system:

    docker-compose up -d

Table of Contents

  1. Project Overview
  2. Architecture
  3. Technology Stack
  4. Implementation Details
  5. Data Flow
  6. Cloud Services Comparison
  7. Performance Metrics
  8. Testing and Validation
  9. Assumptions Made
  10. Strengths and Weaknesses
  11. Research and References
  12. Future Improvements

Project Overview

This project analyzes cloud-based solutions for object detection and counting to determine the flow of people and vehicles using Computer Vision and MLOps techniques. The goal is to compare cost, quality, and performance across AWS, Azure, and Local Processing for implementing these solutions.

Key Objectives:

  • Implement object detection models (YOLO) to count people and vehicles
  • Process data on cloud computing services (AWS, Azure) and evaluate cost, quality, and performance
  • Integrate MLOps practices to ensure scalability and standardization
  • Develop a test application to display metrics and visualizations of processed data
  • Create an image processing pipeline using Python and frameworks like TensorFlow, PyTorch, or OpenCV
  • Define preprocessing and image enhancement strategies to optimize detection
  • Establish evaluation metrics to compare cost, quality, and performance across services

Architecture

The system follows a microservices architecture with clear separation between:

  1. ML Services: Handles object detection and tracking using local and cloud-based implementations
  2. Backend API: Provides RESTful endpoints for data retrieval and processing
  3. Frontend Application: Displays metrics, visualizations, and comparison data
  4. MLOps Pipeline: Tracks experiments, model performance, and service metrics
# Cloud-Based Object Detection and Flow Analysis

## Getting Started

### Prerequisites

- Python 3.8+
- Docker and Docker Compose
- PostgreSQL (if running database locally)
- Node.js 16+ and npm (for frontend)

### Installation

1. Clone the repository:
   ```bash
   git clone https://github.com/yourusername/cloud-object-detection.git
   cd cloud-object-detection
  1. Install Python dependencies:

    pip install -r requirements.txt
  2. Set up environment variables:

    cp .env.example .env
    # Edit .env file with your credentials and configuration
  3. Run the system:

    docker-compose up -d

Table of Contents

  1. Project Overview
  2. Architecture
  3. Technology Stack
  4. Implementation Details
  5. Data Flow
  6. Cloud Services Comparison
  7. Performance Metrics
  8. Testing and Validation
  9. Assumptions Made
  10. Strengths and Weaknesses
  11. Research and References
  12. Future Improvements

Project Overview

This project analyzes cloud-based solutions for object detection and counting to determine the flow of people and vehicles using Computer Vision and MLOps techniques. The goal is to compare cost, quality, and performance across AWS, Azure, and Local Processing for implementing these solutions.

Key Objectives:

Technology Stack

The project leverages a comprehensive set of technologies to achieve its goals:

Machine Learning and Computer Vision

  • YOLOv8 (You Only Look Once): A state-of-the-art, real-time object detection model that processes images in a single pass through a neural network. YOLOv8 represents the eighth generation of the YOLO family, offering significant improvements in speed and accuracy over previous versions. The model divides images into a grid and predicts bounding boxes and class probabilities for each grid cell simultaneously, enabling real-time processing at 30+ frames per second on modern hardware. We specifically use the "nano" variant (YOLOv8n), which is optimized for edge devices and provides an excellent balance between performance (speed) and accuracy. YOLOv8 is built on PyTorch and provides pre-trained weights on the COCO dataset, which includes common categories like people, cars, and other vehicles - perfectly aligned with our project's focus.

  • OpenCV (Open Computer Vision Library): A comprehensive, open-source library designed for computer vision and image processing tasks. Originally developed by Intel, OpenCV provides over 2,500 optimized algorithms for tasks ranging from basic image processing to complex machine learning. In our project, OpenCV handles several critical functions: video capture and decoding from files or webcam streams; image preprocessing (resizing, normalization, color space conversion); frame extraction at specific intervals; and visualization of detection results with bounding boxes and labels. OpenCV is implemented in C++ but provides Python bindings, which we use for seamless integration with our ML pipeline. Its highly optimized algorithms leverage hardware acceleration where available, making it ideal for performance-critical applications like real-time video processing.

  • TensorFlow/PyTorch: Modern deep learning frameworks that provide comprehensive ecosystems for developing, training, and deploying machine learning models.

    PyTorch: Developed by Facebook's AI Research lab, PyTorch offers a dynamic computational graph that allows for flexible model architecture changes during runtime. This makes it particularly well-suited for research and prototyping. YOLOv8 is built on PyTorch, so we leverage this framework for the local processing implementation. PyTorch features include: automatic differentiation for building and training neural networks; GPU acceleration for faster computation; a rich ecosystem of tools and libraries; and seamless transition between development and production environments.

    TensorFlow: Developed by Google, TensorFlow offers a more production-oriented approach with static computational graphs and extensive deployment options. While our primary models use PyTorch, we maintain TensorFlow compatibility for potential future integration with cloud services that prefer TensorFlow models. TensorFlow provides excellent support for mobile and edge deployment through TensorFlow Lite, comprehensive visualization tools through TensorBoard, and enterprise-grade serving infrastructure through TensorFlow Serving.

Cloud Services

  • AWS Rekognition: Amazon's managed computer vision service that provides pre-trained machine learning models for image and video analysis. Rekognition eliminates the need to build, train, and deploy custom models for common computer vision tasks. In our implementation, we use Rekognition's DetectLabels API, which identifies objects, scenes, and activities in images with associated confidence scores. Rekognition operates as a simple API call where we send image data and receive structured JSON responses containing detection information. The service automatically scales to handle varying workloads without requiring infrastructure management. AWS Rekognition includes features like object and scene detection, custom label training, face analysis, text recognition, and content moderation. Pricing is based on the number of images processed, starting at approximately $1 per 1,000 images, with volume discounts for higher usage.

  • Azure Computer Vision: Microsoft's comprehensive computer vision service that provides pre-built models and customization options for image analysis. In our implementation, we use Azure's detect_objects method, which identifies objects within images and returns their bounding box coordinates and confidence scores. Azure Computer Vision offers spatial analysis (understanding object locations within images), optical character recognition (OCR), face detection, image categorization, and custom model training through the Custom Vision service. The Azure SDK for Python provides a clean interface for making API calls with proper authentication and error handling. Azure's pricing model is similar to AWS, starting around $1-$2.50 per 1,000 transactions depending on the specific API and volume. Our implementation includes careful handling of Azure's rate limits (20 calls per minute in the free tier) to ensure smooth operation during testing.

  • Local Processing: Running object detection algorithms directly on the local machine rather than using cloud services. This approach gives complete control over the processing pipeline and eliminates network-related latency and costs. Our local processing implementation uses YOLOv8n running on the user's hardware, providing a critical baseline for comparison with cloud services. Local processing advantages include: no per-request costs (only initial hardware investment); no network dependencies; complete data privacy; and consistent performance regardless of internet connectivity. The primary disadvantages are hardware requirements (ideally a modern CPU and GPU) and manual updates for model improvements.

Backend

  • Go (Golang): A statically typed, compiled programming language designed at Google that combines the performance of compiled languages like C++ with the simplicity and safety features of modern languages. Go excels at building concurrent, high-performance web services, making it ideal for our backend API that needs to handle multiple simultaneous requests and data processing tasks. Key features we leverage include: goroutines for lightweight concurrent processing; efficient HTTP handling through the standard library; excellent database connectivity; and native JSON parsing and generation. Our implementation uses several Go packages, including the standard net/http for the web server, database/sql for database operations, and third-party packages like github.com/lib/pq for PostgreSQL connectivity and go.uber.org/zap for structured, high-performance logging.

  • PostgreSQL: An advanced, open-source object-relational database system with over 30 years of active development. PostgreSQL provides rock-solid stability, extensive standards compliance, and sophisticated features for data storage and retrieval. In our application, PostgreSQL stores several types of data: detection results (object types, counts, timestamps); performance metrics (latency, confidence scores); cost tracking information; and user authentication data. We selected PostgreSQL for several reasons: excellent support for structured data with complex relationships; robust transaction support ensuring data integrity; advanced indexing capabilities for fast queries on large datasets; and native support for JSON data types allowing flexible storage of detection results. Our implementation includes database migrations for schema management and prepared statements to prevent SQL injection while maximizing performance.

Frontend

  • React: A declarative, component-based JavaScript library for building user interfaces, developed and maintained by Facebook. React's component model allows us to build encapsulated, reusable UI pieces that manage their own state, which is critical for our complex dashboard with multiple visualization types. React's virtual DOM efficiently updates only what needs to change, providing excellent performance even when rendering complex data visualizations. Our implementation uses functional components with hooks for state management and side effects, allowing for cleaner, more maintainable code. Key React features we leverage include: context API for global state management; effect hooks for data fetching and subscriptions; and memo for performance optimization of expensive components.

  • TypeScript: A strongly typed superset of JavaScript that adds static type definitions, improving code quality and developer productivity. TypeScript catches type-related errors during development rather than at runtime, significantly reducing bugs in production. Our frontend extensively uses TypeScript for: defining data structures and API response types; creating type-safe component props; and ensuring correct function parameter and return types. TypeScript's type inference allows us to write less explicit type annotations while still receiving type checking benefits. Interface definitions for our data models (like detection results, performance metrics, and comparison data) ensure consistent data handling throughout the application.

  • Recharts: A composable charting library built on React components that provides a declarative way to build data visualizations. Recharts abstracts away the complexity of direct D3.js manipulation while maintaining its flexibility and power. Our dashboard uses Recharts to visualize several key metrics: service cost comparisons (bar and line charts); performance metrics over time (line charts); detection confidence comparisons; and error rates. Recharts features we leverage include: responsive container adapting to screen size; customizable tooltips providing detailed information on hover; animated transitions between data states; and synchronized charts for multi-metric visualization. Its component-based approach aligns perfectly with React's philosophy, allowing for highly maintainable and customizable visualizations.

MLOps and Infrastructure

  • MLflow: An open-source platform designed to manage the complete machine learning lifecycle, including experimentation, reproducibility, deployment, and a central model registry. In our implementation, MLflow serves as the backbone of our experiment tracking and metrics logging system. We use MLflow to track: service performance metrics (latency, throughput); detection quality metrics (confidence scores, accuracy); cost information for each service; and parameter configurations for each experiment run. MLflow's tracking component creates a structured record of all parameters, metrics, and artifacts for each run, enabling comprehensive comparison between different detection methods. The platform's web UI provides interactive visualization of experimental results, while its REST API allows our backend to programmatically retrieve experiment data for display in our custom dashboard.

  • Docker: A platform that uses OS-level virtualization to deliver software in packages called containers, which contain everything needed to run an application: code, runtime, system tools, libraries, and settings. Docker enables consistent deployment across different environments, from development to production. Our implementation containerizes several components: the ML processing service (with all necessary libraries and models); the Go backend API; the PostgreSQL database; and the MLflow tracking server. Docker provides isolation between these components, allowing each to have its own dependencies without conflicts. Container definitions in Dockerfiles specify exact environment configurations, eliminating "works on my machine" problems. Docker's layered file system minimizes image sizes by reusing common components between containers.

  • Docker Compose: A tool for defining and running multi-container Docker applications, allowing all services to be configured in a single YAML file and started with a single command. Our docker-compose.yml file defines the complete application stack with proper networking, volume mounts, and environment variables for each service. Docker Compose manages service dependencies, ensuring they start in the correct order (e.g., PostgreSQL before the backend and MLflow). Volume mounts persist data outside containers, allowing databases and ML experiment results to survive container restarts. Environment variable configuration in the compose file centralizes application settings, making it easy to adjust parameters for different environments without changing application code.

Implementation Details

Object Detection

The system implements three distinct object detection methods, each with its own strengths, limitations, and implementation characteristics:

  1. Local YOLOv8 Implementation:

    YOLOv8 operates as our primary local object detection solution, providing a benchmark for comparison with cloud services.

    # From ml/cloud_comparison/compare_services.py
    def process_image_yolo(self, video_path):
        model = YOLO('yolov8n.pt')
        results = model(frame, verbose=False)[0]
        # Process detections...

    How It Works:

    • The system loads the pre-trained YOLOv8n model (nano variant), which is optimized for speed and efficiency
    • Each video frame is processed through the neural network in a single forward pass
    • The model divides each image into a grid and predicts bounding boxes and object classes for each cell
    • Detection results include bounding box coordinates, class IDs, and confidence scores
    • Our implementation filters results to focus on people and vehicles (COCO classes 0, 2, 3, 5, 7)
    • Processed detections are tracked across frames to count unique objects

    Technical Implementation:

    • We leverage the ultralytics YOLO implementation which provides a clean Python API
    • Processing happens entirely on the local machine, with optional GPU acceleration
    • The model file (yolov8n.pt) is approximately 6.2MB, making it lightweight enough for deployment
    • Frame preprocessing includes resizing to maintain aspect ratio while optimizing for model input size

    Strengths:

    • No per-request costs or network dependencies
    • Consistent, low-latency processing (typically 20-50ms per frame on modern hardware)
    • Complete control over the detection pipeline
    • Deterministic results (same input produces same output)
    • No data privacy concerns as data never leaves the local machine

    Limitations:

    • Hardware-dependent performance (requires decent CPU/GPU for real-time operation)
    • Limited to pre-trained classes unless custom-trained
    • Less sophisticated than larger models (YOLOv8x, YOLOv8l) due to size optimization
    • Manual updates required for model improvements
  2. AWS Rekognition Implementation:

    AWS Rekognition provides cloud-based object detection through a managed API service.

    # From ml/cloud_comparison/compare_services.py
    def process_image_aws(self, image_path):
        with open(image_path, 'rb') as image_file:
            response = self.rekognition.detect_labels(
                Image={'Bytes': image_file.read()},
                MaxLabels=50,
                MinConfidence=50
            )
        # Process response...

    How It Works:

    • The system reads each video frame as an image and converts it to a binary format
    • The image is sent to AWS Rekognition's DetectLabels API via the boto3 Python client
    • Rekognition processes the image on AWS servers using their proprietary deep learning models
    • The API returns a JSON response with detected labels, confidence scores, and bounding boxes
    • Our implementation filters the results to focus on person and vehicle detection
    • We map AWS's label taxonomy to our standardized format for consistent comparison

    Technical Implementation:

    • AWS credentials are securely managed via environment variables
    • Each API call includes parameters for maximum labels (50) and minimum confidence threshold (50%)
    • Responses are asynchronously processed to minimize waiting time
    • The implementation includes error handling for API failures and rate limiting
    • Detected objects are tracked between frames using our IoU (Intersection over Union) tracker

    Strengths:

    • No local computational requirements (processing happens in AWS cloud)
    • Continuously improved models without manual updates
    • Excellent scalability for processing large batches
    • Rich metadata including scene detection, hierarchical labels, and attribute detection
    • High accuracy for common object categories

    Limitations:

    • Cost increases linearly with the number of images processed
    • Network latency adds 150-300ms per request
    • Dependency on internet connectivity
    • Limited control over the detection algorithm
    • Potential privacy concerns with sending images to third-party servers
  3. Azure Computer Vision Implementation:

    Microsoft Azure's Computer Vision service provides an alternative cloud-based detection method.

    # From ml/cloud_comparison/compare_services.py
    def process_image_azure(self, image_path):
        with open(image_path, 'rb') as image_file:
            response = self.azure_client.detect_objects(image_file.read())
        # Process response...

    How It Works:

    • Each video frame is read and converted to a binary format
    • The image is sent to Azure's Computer Vision API using the official Python SDK
    • Azure processes the image using their deep learning models optimized for object detection
    • The API returns a structured response with objects, categories, and bounding box coordinates
    • Our implementation carefully handles Azure's rate limits (20 calls per minute in the free tier)
    • Results are filtered to focus on people and vehicles for consistent comparison

    Technical Implementation:

    • Azure credentials (endpoint and API key) are managed via secure environment variables
    • The implementation includes a rate-limiting mechanism that tracks API calls per minute
    • If approaching the rate limit, the system introduces appropriate delays
    • Response processing converts Azure's format to our standardized detection format
    • Error handling includes retry logic for transient failures

    Strengths:

    • No local computational requirements
    • High-quality detection for certain object categories
    • Regular model updates from Microsoft Research
    • Detailed metadata about detected objects
    • Additional capabilities like OCR and image analysis available through the same API

    Limitations:

    • Stricter rate limits compared to AWS
    • Per-image costs similar to AWS Rekognition
    • Network latency adds 200-350ms per request
    • Different object taxonomy requiring mapping to standardize comparisons
    • Limited control over detection parameters

Cross-Platform Integration:

The system includes a unified detection interface that abstracts away the differences between these three implementations, allowing for:

  • Parallel processing of the same frames across all three methods
  • Standardized output format for consistent comparison
  • Centralized error handling and logging
  • Uniform performance metrics collection

This design allows for fair comparison between the three detection methods, ensuring that differences in results are due to the underlying technologies rather than implementation details.

Object Tracking

The system implements a custom object tracking solution to count unique objects (people and vehicles) across video frames, enabling flow analysis:

# From ml/cloud_comparison/tracking.py
class IoUTracker:
    def __init__(self, iou_threshold=0.3):
        self.tracked_objects = {}
        self.next_id = 0
        self.iou_threshold = iou_threshold
        
    def update(self, detections):
        # Match current detections with existing tracks using IoU
        # Assign new IDs to unmatched detections
        # Return tracked objects with unique IDs

How Object Tracking Works:

Object tracking is a critical component that converts single-frame detections into temporal understanding of objects moving through a scene. Our tracking system uses the following approach:

  1. Detection Association: When new detections arrive from any of our detection methods, they must be associated with existing tracked objects or identified as new objects.

  2. IoU (Intersection over Union) Calculation: For each new detection, we compute the IoU with all existing tracked objects:

    def calculate_iou(self, detection1, detection2):
        # Extract bounding box coordinates
        box1 = detection1.bbox  # [x1, y1, x2, y2]
        box2 = detection2.bbox  # [x1, y1, x2, y2]
        
        # Calculate intersection area
        x_left = max(box1[0], box2[0])
        y_top = max(box1[1], box2[1])
        x_right = min(box1[2], box2[2])
        y_bottom = min(box1[3], box2[3])
        
        if x_right < x_left or y_bottom < y_top:
            return 0.0  # No intersection
        
        intersection_area = (x_right - x_left) * (y_bottom - y_top)
        
        # Calculate union area
        box1_area = (box1[2] - box1[0]) * (box1[3] - box1[1])
        box2_area = (box2[2] - box2[0]) * (box2[3] - box2[1])
        union_area = box1_area + box2_area - intersection_area
        
        # Return IoU
        return intersection_area / union_area
  3. Track Matching: A detection is matched to an existing track if:

    • The IoU exceeds the threshold (default: 0.3)
    • The classes match (e.g., both are "person" or both are "car")
    • It has the highest IoU among all candidates
  4. ID Assignment:

    • Matched detections inherit the ID from the matching track
    • Unmatched detections receive a new unique ID
    • Tracks without matches are maintained for a few frames before being dropped (to handle occlusions)
  5. Track Updating:

    • For matched tracks, position and appearance features are updated using a simple moving average
    • Track confidence is boosted when matches are consistent
    • Track status can be "active", "tentative", or "lost" depending on match history

Technical Implementation:

Our IoUTracker class maintains a dictionary of tracked objects, where each key is a unique ID and each value contains:

  • Current bounding box coordinates
  • Object class
  • Creation timestamp
  • Last update timestamp
  • Track history (list of previous positions)
  • Confidence score
  • Motion vector (estimated direction and speed)

The update method processes new detections with these steps:

def update(self, detections):
    # 1. Initialize data structures for matching
    if not self.tracked_objects:
        # First frame - assign new IDs to all detections
        return self._initialize_tracks(detections)
    
    # 2. Build cost matrix (negative IoU for each detection-track pair)
    cost_matrix = self._build_cost_matrix(detections)
    
    # 3. Solve assignment problem (Hungarian algorithm)
    matched_indices, unmatched_detections, unmatched_tracks = self._assign_detections_to_tracks(cost_matrix)
    
    # 4. Update matched tracks
    for detection_idx, track_idx in matched_indices:
        self._update_track(detection_idx, track_idx, detections)
    
    # 5. Create new tracks for unmatched detections
    for detection_idx in unmatched_detections:
        self._create_new_track(detections[detection_idx])
    
    # 6. Handle unmatched tracks (increment age, mark as lost if too old)
    self._update_unmatched_tracks(unmatched_tracks)
    
    # 7. Delete lost tracks that are too old
    self._delete_old_tracks()
    
    # 8. Return current set of tracked objects
    return self.get_active_tracks()

Strengths of Our Tracking Approach:

  1. Simplicity and Efficiency: The IoU-based approach is computationally lightweight, allowing real-time tracking even on modest hardware.

  2. Platform Agnosticism: Works identically with detections from all three detection methods (local YOLO, AWS, Azure) after standardization.

  3. Stability: Maintains object identity through brief occlusions or missed detections.

  4. Class Awareness: Respects object class during matching, preventing category confusion.

  5. Low Memory Footprint: Only essential information is stored for each track.

Limitations:

  1. Simple Motion Model: Does not use sophisticated motion prediction, which can lead to track confusion in complex scenarios.

  2. Identity Switching: May confuse similar objects after prolonged occlusion or when objects cross paths.

  3. No Appearance Modeling: Unlike more advanced trackers, does not use visual appearance features to improve matching.

  4. Limited Occlusion Handling: Struggles with long-duration occlusions.

  5. Single-Camera Design: Not designed for multi-camera tracking scenarios.

Performance Considerations:

The tracker performance depends on several factors:

  • Detection quality (higher-quality detections lead to better tracking)
  • Scene complexity (crowded scenes are more challenging)
  • Frame rate (higher frame rates improve tracking continuity)
  • Object velocity (faster-moving objects require lower IoU thresholds)

We optimize the IoU threshold based on empirical testing:

  • 0.3 for general scenarios (balancing stability and accuracy)
  • 0.2 for high-motion scenarios
  • 0.4 for static or slow-moving scenes

The tracker provides critical metrics for flow analysis:

  • Object counts by category (total unique people and vehicles)
  • Dwell time (how long objects remain in the scene)
  • Movement patterns (direction and speed distributions)
  • Zone transition analysis (movement between defined areas)

MLOps Implementation

The project implements a comprehensive MLOps (Machine Learning Operations) infrastructure centered around MLflow for experiment tracking, metrics logging, and performance comparison:

# From ml/mlflow_tracking.py
class DetectionTracker:
    def __init__(self, experiment_name: str = "object-detection-comparison"):
        mlflow.set_tracking_uri("http://localhost:5000")
        mlflow.set_experiment(experiment_name)
    
    def start_detection_run(self, service_type: str, model_name: str, input_type: str):
        # Start a new MLFlow run for detection
        
    def log_detection_metrics(self, run_id: str, metrics: Dict[str, float]):
        # Log detection metrics
        
    def log_cloud_cost(self, run_id: str, service_type: str, cost_details: Dict[str, float]):
        # Log cloud service costs

MLOps Architecture Overview:

Our MLOps implementation establishes a structured workflow for experimenting with, evaluating, and comparing different object detection services. The architecture consists of:

  1. MLflow Tracking Server: A centralized service that stores all experiment data, metrics, and artifacts
  2. Experiment Tracking Client: The DetectionTracker class that interfaces with the tracking server
  3. PostgreSQL Backend: Provides persistent storage for experiment data
  4. Docker Containerization: Ensures consistent execution environment
  5. Structured Logging System: Captures detailed execution information

How the MLflow Integration Works:

The MLflow integration operates through a series of coordinated steps:

  1. Experiment Organization:

    def __init__(self, experiment_name: str = "object-detection-comparison"):
        # Set up connection to MLflow server
        mlflow.set_tracking_uri("http://localhost:5000")
        
        # Create or get existing experiment
        mlflow.set_experiment(experiment_name)
        
        # Store experiment ID for future reference
        self.experiment_id = mlflow.get_experiment_by_name(experiment_name).experiment_id

    This initialization creates a logical grouping (experiment) for all detection runs, allowing for organized comparison across different services and configurations.

  2. Run Management:

    def start_detection_run(self, 
                          service_type: str,  # "yolo", "aws", or "azure"
                          model_name: str,    # Model/API version
                          input_type: str,    # "image", "video", or "webcam"
                          batch_size: int = 1):
        # Create descriptive run name
        run_name = f"{service_type}-{model_name}-{datetime.now().strftime('%Y%m%d-%H%M%S')}"
        
        # Start run with MLflow context manager
        with mlflow.start_run(run_name=run_name) as run:
            # Log fundamental parameters
            mlflow.log_params({
                "service_type": service_type,
                "model_name": model_name,
                "input_type": input_type,
                "batch_size": batch_size,
                "timestamp": datetime.now().isoformat(),
                "system_info": self._get_system_info()
            })
            
            # Return run ID for subsequent logging
            return run.info.run_id

    Each detection process creates a new MLflow run, storing contextual parameters that describe the execution environment and configuration.

  3. Metrics Logging:

    def log_detection_metrics(self, 
                            run_id: str,
                            metrics: Dict[str, float],
                            artifacts: Dict[str, Any] = None):
        with mlflow.start_run(run_id=run_id, nested=True):
            # Log numerical metrics
            mlflow.log_metrics(metrics)
            
            # Log artifacts (images, data files, etc.)
            if artifacts:
                with tempfile.TemporaryDirectory() as tmp_dir:
                    for name, artifact in artifacts.items():
                        if isinstance(artifact, np.ndarray):
                            # Handle numpy arrays (e.g., images)
                            artifact_path = os.path.join(tmp_dir, f"{name}.npy")
                            np.save(artifact_path, artifact)
                            mlflow.log_artifact(artifact_path)
                        elif isinstance(artifact, str) and os.path.exists(artifact):
                            # Handle file paths
                            mlflow.log_artifact(artifact)
                        else:
                            # Handle other types
                            artifact_path = os.path.join(tmp_dir, f"{name}.json")
                            with open(artifact_path, 'w') as f:
                                json.dump(artifact, f)
                            mlflow.log_artifact(artifact_path)

    This method captures both numerical metrics (latency, accuracy, detection counts) and related artifacts (sample images, detection visualizations) for each detection run.


## Data Flow

The data flow within the system is as follows:

1. **Video Input**: Video sources (files or webcam streams) provide the input data for object detection.
2. **Object Detection**: Each frame is processed through the three object detection methods (local YOLO, AWS Rekognition, Azure Computer Vision) in parallel.
3. **Detection Results**: Detection results are standardized and sent to the tracking component for object tracking.
4. **Object Tracking**: The tracking component assigns unique IDs to objects and tracks their movement across frames.
5. **Metrics Logging**: Detection metrics (latency, accuracy, confidence scores) and tracking metrics (object counts, dwell time, movement patterns) are logged using MLflow.
6. **Backend API**: The backend API provides RESTful endpoints for data retrieval and processing.
7. **Frontend Application**: The frontend application displays metrics, visualizations, and comparison data.
8. **MLOps Pipeline**: The MLOps pipeline tracks experiments, model performance, and service metrics.

## Cloud Services Comparison

The project compares the performance, cost, and quality of three cloud-based object detection services: AWS Rekognition, Azure Computer Vision, and local YOLOv8 processing.

### Cost Comparison

- **AWS Rekognition**: Cost is based on the number of images processed, starting at approximately $1 per 1,000 images, with volume discounts for higher usage.
- **Azure Computer Vision**: Pricing is similar to AWS, starting around $1-$2.50 per 1,000 transactions depending on the specific API and volume.
- **Local YOLOv8**: No per-request costs or network dependencies, making it the most cost-effective option for local processing.

### Quality Comparison

- **AWS Rekognition**: Offers rich metadata including scene detection, hierarchical labels, and attribute detection. High accuracy for common object categories.
- **Azure Computer Vision**: High-quality detection for certain object categories. Regular model updates from Microsoft Research. Detailed metadata about detected objects. Additional capabilities like OCR and image analysis available through the same API.
- **Local YOLOv8**: Less sophisticated than larger models (YOLOv8x, YOLOv8l) due to size optimization. Limited to pre-trained classes unless custom-trained.

### Performance Comparison

- **AWS Rekognition**: Network latency adds 150-300ms per request. Dependency on internet connectivity. Limited control over the detection algorithm.
- **Azure Computer Vision**: Network latency adds 200-350ms per request. Stricter rate limits compared to AWS. Limited control over detection parameters.
- **Local YOLOv8**: Consistent, low-latency processing (typically 20-50ms per frame on modern hardware). Complete control over the detection pipeline. Deterministic results (same input produces same output). No data privacy concerns as data never leaves the local machine.

## Performance Metrics

The system tracks several performance metrics to evaluate the effectiveness of object detection and tracking:

1. **Detection Latency**: Measures the time taken to process a single frame through the object detection model.
2. **Detection Accuracy**: Evaluates the model's ability to correctly identify objects in images.
3. **Confidence Scores**: Measures the model's confidence in its detections, indicating the likelihood of a true positive.
4. **Object Counts**: Tracks the total number of unique objects (people and vehicles) detected in a scene.
5. **Dwell Time**: Measures how long objects remain in the scene, providing insights into traffic flow and congestion.
6. **Movement Patterns**: Analyzes the direction and speed distributions of objects, revealing traffic patterns and flow directions.
7. **Zone Transition Analysis**: Studies movement between defined areas, enabling zone-based analysis of traffic flow and congestion.

## Testing and Validation

The project includes a comprehensive testing and validation strategy to ensure the reliability and accuracy of the object detection and tracking system:

1. **Unit Testing**: Individual components (object detection, tracking, MLflow integration) are tested in isolation using unit tests.
2. **Integration Testing**: The system is tested as a whole to ensure that all components work together correctly.
3. **Performance Testing**: The system's performance is evaluated under various conditions (different frame rates, scene complexity, object velocity) to ensure optimal operation.
4. **Load Testing**: The system's ability to handle large volumes of data is tested to ensure scalability.
5. **Regression Testing**: Changes to the system are validated to ensure that existing functionality remains unchanged.
6. **User Acceptance Testing**: The system is tested by end-users to ensure that it meets their requirements and expectations.

## Assumptions Made

During the development of the project, several assumptions were made to simplify the implementation and scope of the project:

1. **Single-Camera Setup**: The system is designed for a single-camera setup, where objects are detected and tracked within a single field of view.
2. **Static Scenes**: The system assumes that the scenes being analyzed are static or have minimal changes between frames.
3. **No Occlusion**: The system assumes that objects are not occluded by other objects or structures in the scene.
4. **No Motion Blur**: The system assumes that the camera is stable and that there is no motion blur in the images.
5. **No Nighttime Scenes**: The system assumes that all images are captured during daylight hours, as nighttime scenes may pose additional challenges for object detection.
6. **No Crowded Scenes**: The system assumes that the scenes being analyzed are not overly crowded, as dense crowds may lead to false positives and tracking errors.

## Strengths and Weaknesses

### Strengths

1. **Comprehensive Comparison**: The project provides a thorough comparison of three cloud-based object detection services (AWS Rekognition, Azure Computer Vision) and local YOLOv8 processing, allowing for informed decision-making.
2. **MLOps Integration**: The project integrates MLflow for experiment tracking, metrics logging, and performance comparison, providing a structured workflow for machine learning development and deployment.
3. **Scalable Architecture**: The microservices architecture allows for easy scaling and maintenance of individual components, ensuring long-term sustainability.
4. **Custom Object Tracking**: The project implements a custom object tracking solution that provides critical metrics for flow analysis, enabling insights into traffic patterns and congestion.
5. **Test Application**: The project includes a test application that displays metrics and visualizations of processed data, allowing for easy comparison between different solutions.

### Weaknesses

1. **Limited to Pre-trained Models**: The local YOLOv8 implementation is limited to pre-trained classes, which may not cover all possible object categories.
2. **Hardware Dependencies**: The local YOLOv8 implementation requires a decent CPU and GPU for real-time operation, which may limit its applicability in resource-constrained environments.
3. **Data Privacy Concerns**: Cloud-based object detection services may raise privacy concerns, as images are sent to third-party servers for processing.
4. **Network Dependencies**: Cloud-based object detection services rely on internet connectivity, which may introduce latency and increase costs.
5. **Limited Control over Detection Algorithm**: Cloud-based object detection services offer limited control over the underlying detection algorithm, which may impact performance and accuracy.

## Research and References

The project is based on extensive research and references from various sources:

1. **Cloud Vision API Performance**:
   - [AWS Rekognition Pricing](https://aws.amazon.com/rekognition/pricing/)
   - [Azure Computer Vision Pricing](https://azure.microsoft.com/en-us/pricing/details/cognitive-services/computer-vision/)
   - [Comparing AWS Rekognition and Azure Computer Vision](https://www.linkedin.com/pulse/comparing-aws-rekognition-azure-computer-vision-mohamed-abdelkader/)

2. **Object Detection and Tracking**:
   - [YOLOv8: A Comprehensive Review](https://arxiv.org/abs/2308.05117)
   - [Object Tracking: A Review](https://arxiv.org/abs/1901.02260)
   - [IoU Tracker: A Simple and Efficient Object Tracking Algorithm](https://arxiv.org/abs/1811.05053)

3. **MLOps and Experiment Tracking**:
   - [MLflow: An Open-Source Platform for Machine Learning](https://mlflow.org/)
   - [MLOps: Continuous Delivery and Automation of Machine Learning](https://www.oreilly.com/library/view/mlops-continuous/9781492082956/)

4. **Computer Vision and Image Processing**:
   - [OpenCV: A Comprehensive Guide](https://opencv.org/get-started/)
   - [PyTorch: A Comprehensive Guide](https://pytorch.org/tutorials/)
   - [TensorFlow: A Comprehensive Guide](https://www.tensorflow.org/guide)

## Future Improvements

1. **Multi-Cloud Strategy**: Implement intelligent routing to choose the most cost-effective service for each request
2. **Edge Computing**: Add support for edge devices to reduce cloud dependency and costs
3. **Custom Model Training**: Train specialized models for specific detection scenarios
4. **Advanced Tracking**: Implement more sophisticated tracking algorithms for crowded scenes
5. **Automated Scaling**: Add infrastructure for automatically scaling based on demand
6. **Cost Forecasting**: Implement predictive analytics for cost forecasting
7. **Google Cloud Vision**: Add support for Google Cloud Vision API for more comprehensive comparison

### Frontend Dashboard

The frontend implements interactive visualization dashboards for comparing cloud services, built with React, TypeScript, and Recharts:

```tsx
// From frontend/src/components/dashboard/CloudComparison.tsx
const CloudComparison = () => {
    // Fetch cloud comparison data
    // Render charts for:
    // - Daily cost trends
    // - Total cost comparison
    // - Cost per request
    // - Performance comparison (latency)
    // - Detailed comparison table
}

Frontend Architecture Overview:

The frontend application follows a modern React architecture with a focus on component reusability, type safety, and responsive design:

  1. Component Structure:

    • Shared UI components (buttons, cards, inputs)
    • Feature-specific components (dashboard, comparison, authentication)
    • Layout components (header, sidebar, main content)
    • Page components that compose other components
  2. State Management:

    • React hooks for local component state
    • Context API for global application state

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages