Topics:
arabic-nlptext-normalizationarabic-languagellm-preprocessingnlptashkeel-removerpython
Zero-dependency, high-speed pure Python package designed for Arabic text normalization, essential for preparing Arabic datasets for Large Language Models (LLM pretraining/fine-tuning), NLP tokenization, and Speech AI.
flowchart TD
A["Raw Arabic Text Input"] --> B["Strip Tashkeel (Diacritical Marks)"]
B --> C["Strip Tatweel / Kashida (ـ)"]
C --> D["Convert Eastern & Persian Numerals (٠-٩, ۰-۹ ➔ 0-9)"]
D --> E["Standardize Alef & Hamza Forms (أ, إ, آ ➔ ا)"]
E --> F{"Keep Teh Marbuta?"}
F -- "No (Default)" --> G["Map ة ➔ ه and ى ➔ ي"]
F -- "Yes (--keep-teh-marbuta)" --> H["Preserve ة and ى"]
G --> I["Collapse Excessive Whitespace & Trim"]
H --> I
I --> J["Clean Normalized Output"]
- Tashkeel Stripping: Eliminates diacritics efficiently (Fatha, Damma, Kasra, Sukun, Tanween, Shadda, Dagger Alef).
- Tatweel Elimination: Removes elongation characters (
ـ/ Kashida). - Hamza & Alef Standardization: Normalizes
[أ, إ, آ, ى, ة]for consistent NLP vectorization. - Eastern to Western Numerals: Standardizes both Eastern Arabic (
٠-٩) and Persian (۰-۹) numerals into standard digits (0-9). - Whitespace Normalization: Collapses repeated whitespaces, tabs, and newlines into clean single spaces.
- Robust Type Validation: Validates inputs with explicit type errors to prevent silent processing failures.
- Interactive CLI & Rich Feedback: Supports single strings, pipe streaming, files, and processing statistics.
git clone https://github.com/shadialhasan/ar-text-normalizer.git
cd ar-text-normalizer
pip install -r requirements.txtfrom ar_normalizer import ArabicNormalizer
raw_text = "تَطْوِيرُ النُّظُمِ الذَّكِيَّةِ لِعَامِ ٢٠٢٦"
clean_text = ArabicNormalizer.full_normalize(raw_text)
print(clean_text)
# Output: تطوير النظم الذكيه لعام 2026from ar_normalizer import ArabicNormalizer
text = "مَدِينَةُ القَاهِرَةِ"
clean_preserved = ArabicNormalizer.full_normalize(text, keep_teh_marbuta=True)
print(clean_preserved)
# Output: مدينة القاهرةfrom ar_normalizer import ArabicNormalizer
# Remove diacritics only
vowels_removed = ArabicNormalizer.strip_tashkeel("كِتَابٌ مُفِيدٌ")
# Remove tatweel only
tatweel_removed = ArabicNormalizer.strip_tatweel("تـــــطــــويــــر")
# Normalize numerals only
digits_fixed = ArabicNormalizer.convert_numerals("الهاتف: ٠١٢٣٤٥٦٧٨٩")
# Normalize hamza representations
hamza_fixed = ArabicNormalizer.normalize_hamza("إبراهيم وأحمد")The package provides a built-in CLI interface:
python -m ar_normalizer "بِسْمِ اللَّهِ الرَّحْمَٰنِ الرَّحِيمِ - القَاهِرَةُ ١٢٣٤٥" --stats# Normalize an input text file and save the output
python -m ar_normalizer -i raw_dataset.txt -o cleaned_dataset.txt --stats
# Preserve Teh Marbuta in output
python -m ar_normalizer -i corpus.txt -o corpus_norm.txt --keep-teh-marbutaA .env.example file is included for pipeline configurations:
cp .env.example .envAvailable options:
AR_NORMALIZE_KEEP_TEH_MARBUTA: Default behavior for Teh Marbuta preservation.AR_NORMALIZE_CONVERT_NUMERALS: Convert Eastern numerals to ASCII digits.LOG_LEVEL: Logging verbosity level.
Run the full automated test suite verifying all edge cases:
python -m unittest discover tests -vEng. MHD. Shadi AL-Hasan
- Role: Executive CTO & Enterprise Solutions Architect
- Email: mhd.shadi.alhasan@gmail.com
- Phone / WhatsApp: +963934005922
- Location: Damascus, Syria
- GitHub: shadialhasan
This project is licensed under the MIT License - see the LICENSE file for details.
Copyright (c) 2026 MHD. Shadi AL-Hasan. All rights reserved.