Skip to content

Repository files navigation

🔤 ar-text-normalizer

Topics: arabic-nlp text-normalization arabic-language llm-preprocessing nlp tashkeel-remover python

Release CI License: MIT Zero Dependency Python: 3.9+

Zero-dependency, high-speed pure Python package designed for Arabic text normalization, essential for preparing Arabic datasets for Large Language Models (LLM pretraining/fine-tuning), NLP tokenization, and Speech AI.


🏛 Architecture & Normalization Pipeline

flowchart TD
    A["Raw Arabic Text Input"] --> B["Strip Tashkeel (Diacritical Marks)"]
    B --> C["Strip Tatweel / Kashida (ـ)"]
    C --> D["Convert Eastern & Persian Numerals (٠-٩, ۰-۹ ➔ 0-9)"]
    D --> E["Standardize Alef & Hamza Forms (أ, إ, آ ➔ ا)"]
    E --> F{"Keep Teh Marbuta?"}
    F -- "No (Default)" --> G["Map ة ➔ ه and ى ➔ ي"]
    F -- "Yes (--keep-teh-marbuta)" --> H["Preserve ة and ى"]
    G --> I["Collapse Excessive Whitespace & Trim"]
    H --> I
    I --> J["Clean Normalized Output"]
Loading

🚀 Capabilities

  • Tashkeel Stripping: Eliminates diacritics efficiently (Fatha, Damma, Kasra, Sukun, Tanween, Shadda, Dagger Alef).
  • Tatweel Elimination: Removes elongation characters (ـ / Kashida).
  • Hamza & Alef Standardization: Normalizes [أ, إ, آ, ى, ة] for consistent NLP vectorization.
  • Eastern to Western Numerals: Standardizes both Eastern Arabic (٠-٩) and Persian (۰-۹) numerals into standard digits (0-9).
  • Whitespace Normalization: Collapses repeated whitespaces, tabs, and newlines into clean single spaces.
  • Robust Type Validation: Validates inputs with explicit type errors to prevent silent processing failures.
  • Interactive CLI & Rich Feedback: Supports single strings, pipe streaming, files, and processing statistics.


⚙️ Installation & Setup

git clone https://github.com/shadialhasan/ar-text-normalizer.git
cd ar-text-normalizer
pip install -r requirements.txt

📦 Python API Usage

1. Basic Full Normalization

from ar_normalizer import ArabicNormalizer

raw_text = "تَطْوِيرُ النُّظُمِ الذَّكِيَّةِ لِعَامِ ٢٠٢٦"
clean_text = ArabicNormalizer.full_normalize(raw_text)
print(clean_text)
# Output: تطوير النظم الذكيه لعام 2026

2. Preserving Teh Marbuta

from ar_normalizer import ArabicNormalizer

text = "مَدِينَةُ القَاهِرَةِ"
clean_preserved = ArabicNormalizer.full_normalize(text, keep_teh_marbuta=True)
print(clean_preserved)
# Output: مدينة القاهرة

3. Granular Step-by-Step Operations

from ar_normalizer import ArabicNormalizer

# Remove diacritics only
vowels_removed = ArabicNormalizer.strip_tashkeel("كِتَابٌ مُفِيدٌ")

# Remove tatweel only
tatweel_removed = ArabicNormalizer.strip_tatweel("تـــــطــــويــــر")

# Normalize numerals only
digits_fixed = ArabicNormalizer.convert_numerals("الهاتف: ٠١٢٣٤٥٦٧٨٩")

# Normalize hamza representations
hamza_fixed = ArabicNormalizer.normalize_hamza("إبراهيم وأحمد")

💻 Command Line Interface (CLI)

The package provides a built-in CLI interface:

Direct text processing:

python -m ar_normalizer "بِسْمِ اللَّهِ الرَّحْمَٰنِ الرَّحِيمِ - القَاهِرَةُ ١٢٣٤٥" --stats

File processing:

# Normalize an input text file and save the output
python -m ar_normalizer -i raw_dataset.txt -o cleaned_dataset.txt --stats

# Preserve Teh Marbuta in output
python -m ar_normalizer -i corpus.txt -o corpus_norm.txt --keep-teh-marbuta

⚙️ Configuration (.env)

A .env.example file is included for pipeline configurations:

cp .env.example .env

Available options:

  • AR_NORMALIZE_KEEP_TEH_MARBUTA: Default behavior for Teh Marbuta preservation.
  • AR_NORMALIZE_CONVERT_NUMERALS: Convert Eastern numerals to ASCII digits.
  • LOG_LEVEL: Logging verbosity level.

🧪 Running Automated Tests

Run the full automated test suite verifying all edge cases:

python -m unittest discover tests -v

👤 Author & Maintainer

Eng. MHD. Shadi AL-Hasan


📄 License

This project is licensed under the MIT License - see the LICENSE file for details.
Copyright (c) 2026 MHD. Shadi AL-Hasan. All rights reserved.

About

Zero-dependency high-speed Arabic text normalizer & preprocessing library for NLP and LLM training

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages