Author: Filippo Camossi Academic Year: 2025-2026
This repository contains a collection of four distinct Big Data analytics projects developed for the Advanced Information Systems and Big Data course. The portfolio explores the practical application of different NoSQL databases (Document, Time-Series, Graph) and Distributed Computing paradigms (MapReduce) on real-world and synthetic datasets.
The goal is to benchmark and compare data ingestion, complex querying, and machine learning integration (such as sentiment analysis, recommendation engines, and online clustering) across different technological stacks.
Each module is completely independent and contains its own specific documentation, source code, and Python requirements.
- 📦 Homework #1: MongoDB
- Domain: E-commerce (Amazon Fine Food Reviews, ~568k records).
- Focus: Document Store analytics, VADER Sentiment Analysis, and a Hybrid Recommendation System (Collaborative + Content-Based + Sentiment).
- 📈 Homework #2: InfluxDB
- Domain: Urban Security (Los Angeles Crime Data, ~900k records).
- Focus: Time-Series Database management, Flux analytical queries, spatial-temporal pattern recognition, and Incremental/Online Clustering using
River.
- 🕸️ Homework #3: Neo4j
- Domain: Scientific Research (ArXiv AI Papers).
- Focus: Graph Database modeling (Papers, Authors, Topics), Cypher queries, co-authorship network analysis, and weighted Jaccard (Ruzicka) similarity calculations.
- 🚀 Homework #4: MapReduce
- Domain: Transportation (Highway Tolls, 100k records).
- Focus: Distributed Computing benchmark. Compares a pure Python implementation (
mrjob/ Hadoop Streaming) against Spark's DataFrame API (PySpark) to calculate inflation trends over a decade.
- Python 3.8+
- Jupyter Notebook / JupyterLab
- Specific database engines depending on the module (MongoDB, InfluxDB 2.x, Neo4j).
- Java 8+ (Required for PySpark in HW#4).
# Clone the repository
git clone [https://github.com/yourusername/SEBD_homework.git](https://github.com/yourusername/SEBD_homework.git)
cd SEBD_homework
# Navigate to a specific project module
cd MongoDB # or InfluxDB, Neo4j, MapReduce
# Install specific dependencies
pip install -r requirements.txt
# Launch Jupyter
jupyter notebook