Skip to content

Repository files navigation

Itepee — Indonesia Tweet Prediction

Python TensorFlow Django Bootstrap

Hate speech detection in Indonesian tweets using 4 TensorFlow/Keras deep learning models.

Table of Contents


Overview

This system classifies Indonesian tweets into 4 aspects of hate speech using an ensemble of 4 deep learning models.

Pipeline:

  1. Preprocessing — Text cleaning, slang normalization, stopword removal
  2. Detection — Binary classification: hate speech or non-hate speech
  3. Multi-label Classification — Target (Individual/Group), Type (Religion/Race/Physical/Gender/Other), Level (Weak/Moderate/Strong)
  4. Web Interface — Django with Bootstrap 5

Tech Stack

Category Tools
Language Python 3.11
Web Framework Django 4.1
Deep Learning TensorFlow / Keras (4 models)
Frontend Bootstrap 5, Inter font
Database SQLite (dev) / MySQL
Data Processing Pandas, NumPy

Project Structure

Itepee/
├── Frontend/
│   ├── Layout/main.html       # Main layout (navbar, footer)
│   ├── index.html             # Landing page
│   ├── predict.html           # Tweet input form
│   ├── hate_speech.html       # Result: HS detected
│   ├── non_hate_speech.html   # Result: safe
│   └── detail.html            # Probability detail + progress bar
├── Itepee/
│   ├── settings.py            # Django configuration
│   ├── urls.py                # URL routing
│   ├── asgi.py / wsgi.py
├── ItepeeApp/
│   ├── views.py               # Prediction & rendering logic
│   ├── preprocessor.py        # Text preprocessing (8 steps)
│   ├── models.py              # Database model (TweetModel)
│   ├── admin.py               # Django admin
│   ├── model/                 # 4 .h5 models + tokenizer pickle
│   └── data/                  # Slang dictionary, stopwords, HTML entities
├── static/
│   └── style.css              # Custom styling
├── .env.example
├── manage.py
└── README.md

Dataset

Labels:

# Label Description
1 HS Hate speech
2 Abusive Abusive language
3 HS_Individual Targeted at individual
4 HS_Group Targeted at group
5 HS_Religion Religion
6 HS_Race Race/ethnicity
7 HS_Physical Physical/disability
8 HS_Gender Gender/sexual orientation
9 HS_Weak Weak level
10 HS_Moderate Moderate level
11 HS_Strong Strong level

Paper: Multi-label Hate Speech and Abusive Language Detection in Indonesian Twitter — ALW3 2019

Getting Started

1. Clone

git clone https://github.com/sintiasnn/Itepee.git
cd Itepee

2. Setup Environment

Using conda (recommended)
conda activate satelit

If satelit environment doesn't exist:

conda create -n satelit python=3.11 -y
conda activate satelit
pip install tensorflow django pandas numpy python-decouple
pip install tf_keras keras-preprocessing
Using venv
python -m venv venv
source venv/bin/activate    # macOS/Linux
# venv\Scripts\activate     # Windows
pip install -r requirements.txt

3. Migrate Database

python manage.py migrate

4. Run Server

python manage.py runserver

Open http://localhost:8000

Author

Ni Putu Sintia Wati

Contributors:

About

Multi-label hate speech & abusive language detection for Indonesian tweets using deep learning

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages