Skip to content

Latest commit

 

History

38 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Advanced Information Systems & Big Data (SEBD)

Python MongoDB InfluxDB Neo4j Apache Spark

Author: Filippo Camossi Academic Year: 2025-2026

📖 Portfolio Overview

This repository contains a collection of four distinct Big Data analytics projects developed for the Advanced Information Systems and Big Data course. The portfolio explores the practical application of different NoSQL databases (Document, Time-Series, Graph) and Distributed Computing paradigms (MapReduce) on real-world and synthetic datasets.

The goal is to benchmark and compare data ingestion, complex querying, and machine learning integration (such as sentiment analysis, recommendation engines, and online clustering) across different technological stacks.

🗂️ Repository Structure & Projects

Each module is completely independent and contains its own specific documentation, source code, and Python requirements.

  • 📦 Homework #1: MongoDB
    • Domain: E-commerce (Amazon Fine Food Reviews, ~568k records).
    • Focus: Document Store analytics, VADER Sentiment Analysis, and a Hybrid Recommendation System (Collaborative + Content-Based + Sentiment).
  • 📈 Homework #2: InfluxDB
    • Domain: Urban Security (Los Angeles Crime Data, ~900k records).
    • Focus: Time-Series Database management, Flux analytical queries, spatial-temporal pattern recognition, and Incremental/Online Clustering using River.
  • 🕸️ Homework #3: Neo4j
    • Domain: Scientific Research (ArXiv AI Papers).
    • Focus: Graph Database modeling (Papers, Authors, Topics), Cypher queries, co-authorship network analysis, and weighted Jaccard (Ruzicka) similarity calculations.
  • 🚀 Homework #4: MapReduce
    • Domain: Transportation (Highway Tolls, 100k records).
    • Focus: Distributed Computing benchmark. Compares a pure Python implementation (mrjob / Hadoop Streaming) against Spark's DataFrame API (PySpark) to calculate inflation trends over a decade.

🚀 Quick Start

General Prerequisites

  • Python 3.8+
  • Jupyter Notebook / JupyterLab
  • Specific database engines depending on the module (MongoDB, InfluxDB 2.x, Neo4j).
  • Java 8+ (Required for PySpark in HW#4).

Basic Setup

# Clone the repository
git clone [https://github.com/yourusername/SEBD_homework.git](https://github.com/yourusername/SEBD_homework.git)
cd SEBD_homework

# Navigate to a specific project module
cd MongoDB  # or InfluxDB, Neo4j, MapReduce

# Install specific dependencies
pip install -r requirements.txt

# Launch Jupyter
jupyter notebook

About

A comprehensive Big Data analytics portfolio exploring NoSQL databases (MongoDB, InfluxDB, Neo4j) and Distributed Computing paradigms (MapReduce, PySpark) on real-world datasets.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Contributors

Languages