Hate speech detection in Indonesian tweets using 4 TensorFlow/Keras deep learning models.
This system classifies Indonesian tweets into 4 aspects of hate speech using an ensemble of 4 deep learning models.
Pipeline:
- Preprocessing — Text cleaning, slang normalization, stopword removal
- Detection — Binary classification: hate speech or non-hate speech
- Multi-label Classification — Target (Individual/Group), Type (Religion/Race/Physical/Gender/Other), Level (Weak/Moderate/Strong)
- Web Interface — Django with Bootstrap 5
| Category | Tools |
|---|---|
| Language | Python 3.11 |
| Web Framework | Django 4.1 |
| Deep Learning | TensorFlow / Keras (4 models) |
| Frontend | Bootstrap 5, Inter font |
| Database | SQLite (dev) / MySQL |
| Data Processing | Pandas, NumPy |
Itepee/
├── Frontend/
│ ├── Layout/main.html # Main layout (navbar, footer)
│ ├── index.html # Landing page
│ ├── predict.html # Tweet input form
│ ├── hate_speech.html # Result: HS detected
│ ├── non_hate_speech.html # Result: safe
│ └── detail.html # Probability detail + progress bar
├── Itepee/
│ ├── settings.py # Django configuration
│ ├── urls.py # URL routing
│ ├── asgi.py / wsgi.py
├── ItepeeApp/
│ ├── views.py # Prediction & rendering logic
│ ├── preprocessor.py # Text preprocessing (8 steps)
│ ├── models.py # Database model (TweetModel)
│ ├── admin.py # Django admin
│ ├── model/ # 4 .h5 models + tokenizer pickle
│ └── data/ # Slang dictionary, stopwords, HTML entities
├── static/
│ └── style.css # Custom styling
├── .env.example
├── manage.py
└── README.md
- Source: id-multi-label-hate-speech-and-abusive-language-detection by Okky Ibrohim & Indra Budi
- License: CC BY-NC-SA 4.0
Labels:
| # | Label | Description |
|---|---|---|
| 1 | HS | Hate speech |
| 2 | Abusive | Abusive language |
| 3 | HS_Individual | Targeted at individual |
| 4 | HS_Group | Targeted at group |
| 5 | HS_Religion | Religion |
| 6 | HS_Race | Race/ethnicity |
| 7 | HS_Physical | Physical/disability |
| 8 | HS_Gender | Gender/sexual orientation |
| 9 | HS_Weak | Weak level |
| 10 | HS_Moderate | Moderate level |
| 11 | HS_Strong | Strong level |
Paper: Multi-label Hate Speech and Abusive Language Detection in Indonesian Twitter — ALW3 2019
git clone https://github.com/sintiasnn/Itepee.git
cd ItepeeUsing conda (recommended)
conda activate satelitIf satelit environment doesn't exist:
conda create -n satelit python=3.11 -y
conda activate satelit
pip install tensorflow django pandas numpy python-decouple
pip install tf_keras keras-preprocessingUsing venv
python -m venv venv
source venv/bin/activate # macOS/Linux
# venv\Scripts\activate # Windows
pip install -r requirements.txtpython manage.py migratepython manage.py runserverNi Putu Sintia Wati
- GitHub: @sintiasnn
Contributors: