Trained spaCy pipelines for Persian, installable with pip. spaCy has never shipped an official
one, and spacy.blank("fa") only gives you a tokenizer and stop words. These add POS tags,
morphology, lemmas, dependency parses and named entities, trained on the UD Persian PerDT
treebank.
pip install https://huggingface.co/Phazel/fa_core_news_sm/resolve/main/fa_core_news_sm-3.8.0-py3-none-any.whl>>> import spacy
>>> nlp = spacy.load("fa_core_news_sm")
>>> doc = nlp("محمدرضا شجریان در مشهد به دنیا آمد.")
>>> [(t.text, t.pos_, t.lemma_, t.dep_) for t in doc][:2]
[('محمدرضا', 'PROPN', 'محمدرضا', 'nsubj'), ('شجریان', 'PROPN', 'شجریان', 'flat:name')]
>>> doc.ents
(محمدرضا شجریان, مشهد)
>>> doc = nlp("وزارت نفت روز شنبه قیمت بنزین را ۱۰ درصد افزایش داد.")
>>> [(e.text, e.label_) for e in doc.ents]
[('وزارت نفت', 'ORG'), ('روز شنبه', 'DAT'), ('۱۰ درصد', 'PCT')]Ten packages in four tiers. Three run on CPU, one wants a GPU:
sm: no static vectors. Smallest and fastest.md: adds floret vectors (fastText-style vectors that hash subwords, so no word is out of vocabulary), 50k rows trained on 400,000 Persian Wikipedia articles.lg: floret vectors, 200k rows trained on the full Persian Wikipedia.trf: fine-tuned ParsBERT (a Persian BERT). Most accurate, needs a GPU.
| You need | sm |
md / lg |
trf |
|---|---|---|---|
| Tags, lemmas, parse | fa_dep_news_sm |
fa_dep_news_md, fa_dep_news_lg |
fa_core_news_trf |
| Entities only | fa_ent_news_sm |
fa_ent_news_md, fa_ent_news_lg |
fa_core_news_trf |
| Both | fa_core_news_sm |
fa_core_news_md, fa_core_news_lg |
fa_core_news_trf |
| Package | Size |
|---|---|
fa_dep_news_sm |
7.9 MB |
fa_core_news_sm |
13.5 MB |
fa_ent_news_sm |
5.9 MB |
fa_dep_news_md |
62.6 MB |
fa_core_news_md |
68.5 MB |
fa_ent_news_md |
60.6 MB |
fa_dep_news_lg |
229.3 MB |
fa_core_news_lg |
235.2 MB |
fa_ent_news_lg |
227.3 MB |
fa_core_news_trf |
608.2 MB |
HF=https://huggingface.co/Phazel
# syntax + NER, 13.5 MB
pip install $HF/fa_core_news_sm/resolve/main/fa_core_news_sm-3.8.0-py3-none-any.whl
# syntax only, 7.9 MB
pip install $HF/fa_dep_news_sm/resolve/main/fa_dep_news_sm-3.8.0-py3-none-any.whl
# NER only, 5.9 MB
pip install $HF/fa_ent_news_sm/resolve/main/fa_ent_news_sm-3.8.0-py3-none-any.whlFor the vector tiers, swap sm for md or lg in both the repo name and the filename:
pip install $HF/fa_core_news_md/resolve/main/fa_core_news_md-3.8.0-py3-none-any.whl
pip install $HF/fa_core_news_lg/resolve/main/fa_core_news_lg-3.8.0-py3-none-any.whlfa_core_news_trf is versioned 1.0.0, not 3.8.0, and needs spacy-transformers:
pip install spacy-transformers
pip install $HF/fa_core_news_trf/resolve/main/fa_core_news_trf-1.0.0-py3-none-any.whlSyntax scores are on the held-out test split of PerDT (UD_Persian-PerDT), the Persian
Dependency Treebank in Universal Dependencies form: 29,107 sentences with gold UPOS, features,
lemmas and dependencies. Entity scores use a separate NER layer over the same sentences, see
Named entity recognition. Scores come from
spacy benchmark accuracy and also ship in each package's meta.json.
What the metrics measure (percentages, higher is better):
TOKEN_ACC/TOKEN_F: how well the tokenizer finds word boundaries.TAG_ACC: fine-grained part-of-speech tags (XPOS, PerDT's own tagset).POS_ACC: coarse part-of-speech tags (UPOS, the 17 universal tags such asNOUN,VERB).MORPH_ACC: morphological features such as number, person and tense.LEMMA_ACC: the dictionary form of each word (کتابها→کتاب).SENTS_F: sentence boundaries.DEP_UAS(unlabeled attachment score): share of words attached to the correct head word.DEP_LAS(labeled attachment score): the same, but the relation label (nsubj,obj, …) must also be right. Always at or below UAS.ENTS_P/ENTS_R/ENTS_F: entity precision (share of predicted entities that are right), recall (share of real entities that were found), and F, the balance of the two.
The md tier's config differs from sm by exactly one line (include_static_vectors), so
the columns below isolate what the vectors buy. Full breakdown in docs/MODELS.md §6.
| Metric | sm |
md |
lg |
trf |
Reference |
|---|---|---|---|---|---|
TOKEN_ACC / TOKEN_F |
99.96 / 99.11 | 99.96 / 99.11 | 99.96 / 99.11 | 99.96 / 99.11 | |
TAG_ACC (XPOS) |
95.96 | 96.25 | 96.55 | 97.62 | |
POS_ACC (UPOS) |
96.24 | 96.64 | 96.68 | 97.63 | |
MORPH_ACC |
96.29 | 96.64 | 96.70 | 97.82 | |
LEMMA_ACC |
97.91 | 97.96 | 98.08 | 97.31 | |
SENTS_F |
99.25 | 99.28 | 99.18 | 97.35 | |
DEP_UAS |
89.69 | 90.52 | 90.96 | 93.87 | hazm-bert-dependency-parser: 92.46 |
DEP_LAS |
85.15 | 86.34 | 86.60 | 90.79 | hazm-bert-dependency-parser: 89.34 |
ENTS_P |
77.67 | 76.56 | 81.51 | 84.06 | |
ENTS_R |
66.87 | 72.95 | 71.09 | 81.76 | |
ENTS_F |
71.87 | 74.71 | 75.94 | 82.89 | |
| Speed (940MX, batch 32) | 10,235 words/s | 9,058 words/s | 9,215 words/s | 1,106 words/s | |
| Wheel size | 13.5 MB | 68.5 MB | 235.2 MB | 608.2 MB |
Reference cells are hazm-bert-dependency-parser's
own meta.json, trained on the same treebank (docs/MODELS.md §4).
trf leads on every metric except lemmatization and sentence segmentation. It is also the only
tier to clear the hazm-bert-dependency-parser DEP_LAS reference of 89.34. It needs a GPU in
production (187 words/s on the laptop CPU), and its HooshvareLab/bert-base-parsbert-uncased
encoder states no licence, so the published wheel carries a redistribution warning in its
meta.json and its terms are unknown (docs/MODELS.md §8).
Entity scores are fa_core_news_* on the PerDT NER test split; per-label breakdown and
caveats are in Named entity recognition.
Compared against Hazm (https://github.com/roshan-research/hazm), the most widely used Python
toolkit for Persian, and en_core_web_sm (English reference).
| Metric | This projectfa_core_news_trf |
Hazm (Persian toolkit) |
en_core_web_sm(English reference) |
|---|---|---|---|
| POS accuracy | 97.63% UPOS | 98.8% own tagset¹ | 97.29% PTB XPOS² |
| Lemma accuracy | 97.31% | 89.9%¹ | not reported² |
| Dependency UAS / LAS | 93.87% / 90.79% | 92.46% / 89.34%¹ | 91.77% / 89.92%² |
| NER F-score | 82.89% | not reported¹ | 84.33%² |
¹ Hazm's own README, https://github.com/roshan-research/hazm#evaluation:
POSTagger98.8% on Hazm's own tagset (augmented with ezafe markers), which is not UPOS;Lemmatizer89.9%. It reports no NER score. Parser scores are fromhazm-bert-dependency-parser'smeta.json, as in the tier comparison; Hazm's README quotes UAS 92.30 / LAS 89.15 for itsSpacyDependencyParser.²
en_core_web_sm3.8.0meta.json, https://github.com/explosion/spacy-models/blob/master/meta/en_core_web_sm-3.8.0.json:tag_acc0.9729,dep_uas0.9177,dep_las0.8992,ents_f0.8433. That pipeline reports nopos_accand nolemma_acc, because OntoNotes has no UPOS or lemma layer.The three columns were scored on different test data.
Median of repeated nlp.pipe passes over the 146-document PerDT test split (23,825 words),
timing the pipe only, warmup discarded. Reproduce with
python scripts/benchmark_throughput.py <model> --gpu-id <n>, which writes the raw records to
metrics/throughput-*.json.
| Tier | CPU, i5-7200U | GPU, GeForce 940MX | CPU, Xeon @ 2.00GHz | GPU, Tesla T4 |
|---|---|---|---|---|
sm |
5,484 | 10,235 | ||
md |
5,408 | 9,058 | ||
lg |
4,715 | 9,215 | ||
trf |
187 | 1,106 | 336 | 8,320 |
trf is 29x slower than sm on the same CPU. The Xeon and T4 columns come from one Colab VM,
a 25x GPU speedup. The CPU tiers sit within 15% of each other, so the bottleneck is the parser
and lemmatizer, not the tok2vec lookup. Laptop numbers vary by about 10% with thermal state.
Running trf on the 940MX needs a cu126 torch build, see docs/MODELS.md §9.
The same 146 documents as raw text, full pipeline, same laptop CPU. Reproduce with
scripts/benchmark_toolkit.py.
| Toolkit | Components | Words/s |
|---|---|---|
fa_dep_news_sm |
tokenizer, tagger, morphologizer, lemmatizer, parser | 13,073 |
fa_core_news_sm |
the same, plus NER | 8,831 |
UDPipe 1.4, persian-seraji-ud-2.5 model³ |
tokenizer, tagger, lemmatizer, parser | 1,929 |
| DadmaTools 2.3.6¹ | tokenizer, lemmatizer, POS tagger, parser, on XLM-RoBERTa | 96 |
Stanza 1.14, perdt models¹ |
tokenizer, MWT, POS tagger, lemmatizer, parser | 82 |
| Hazm 0.12.1² | normalizer, tokenizers, POS tagger, lemmatizer, MaltParser | 25 |
fa_dep_news_sm is 7x faster than UDPipe, 136x faster than DadmaTools and 160x faster than
Stanza. md and lg stay within 15% of sm, so they lead too. trf, at 187 words/s, is
slower than UDPipe but still faster than DadmaTools, Stanza and Hazm.
¹ Timed on the first 16 documents (2,396 words); the other rows use all 146.
² Timed on 2026-08-11 in one 933-second pass. Every other row was measured on 2026-09-28 with one thread per process.
³ Speed only. The UDPipe 1 models (UD 2.5) predate PerDT, so this row uses Seraji (
UD_Persian-Seraji), a smaller Persian UD treebank.This table and the one above come from different days with different background load, which is why
fa_core_news_smreads 5,484 there and 8,831 here. Compare within a table, not across.
Seven labels: LOC (location), PER (person), ORG (organisation), DAT (date), MON
(money), TIM (time), PCT (percent). They come from PerDT's own
not-to-release/Dadegan with NER tag/ layer, transferred onto this pipeline's tokenization
by difflib at a 99.86% alignment rate (scripts/transfer_perdt_ner.py). Spans that could not
be aligned exactly were dropped rather than guessed. That layer is silver: PerDT's README
states it was produced by the BERT-based Beheshti-NER tagger with manual corrections for
recall, so the ENTS_F numbers below partly reflect agreement with that tagger, not with
human annotation.
Both that realigned layer and a four-label LLM relabelling of the same sentences
(annotation/, guideline v2.2) are published as
Phazel/fa-perdt-ner, CC BY-SA 4.0,
keyed by PerDT sent_id; spacy project run hub-dataset rebuilds it. The relabelling is
measured against the silver layer in docs/MODELS.md §10 and ships no model yet.
ner runs standalone with its own embedded tok2vec (fa_ent_news_sm, fa_ent_news_md,
fa_ent_news_lg), or bundled into fa_core_news_sm, fa_core_news_md and fa_core_news_lg
alongside the syntax pipeline. In trf it is trained jointly against the shared transformer
instead, so there is no standalone trf variant.
| Label | sm F |
md F |
lg F |
trf F |
Train examples |
|---|---|---|---|---|---|
LOC |
80.24 | 84.05 | 83.66 | 87.78 | 4,954 |
PER |
65.29 | 68.18 | 72.63 | 81.88 | 4,847 |
ORG |
68.77 | 70.25 | 71.01 | 78.50 | 2,643 |
DAT |
74.45 | 76.19 | 70.83 | 82.52 | 1,323 |
MON |
73.68 | 84.21 | 88.89 | 88.89 | 205 |
TIM |
66.67 | 66.67 | 61.54 | 50.00 | 135 |
PCT |
57.14 | 33.33 | 57.14 | 33.33 | 121 |
MON, TIM and PCT have single-digit support in the test split, so their deltas are one or
two entities changing hands, not signal. PER, LOC and ORG carry the split. The md gain
over sm (ENTS_F 71.87 to 74.71) is almost entirely recall (+6.08): static vectors already
know rare proper nouns, which sm can only learn from the training data. trf adds another
+6.95 F over lg, again mostly recall (71.09 to 81.76), and its largest per-label gains are
PER (+9.25) and DAT (+11.69).
- Some lemmas contain a space. Multiword tokens were merged, so
کتابهایشis one token taggedN_IANM_PR_JOPERwith lemmaکتاب او. This affects about 1.5% of tokens. doc.noun_chunksunder-fires.spacy/lang/fa/syntax_iterators.pyupstream matches ClearNLP labels that do not exist in Universal Dependencies. Bug analysis and proposed upstream patch indocs/upstream/fa-noun-chunks.md.
If you use these pipelines or the fa-perdt-ner dataset, please cite:
@software{fazeli_2026_persian_spacy,
author = {Fazeli, Mohammad},
title = {Persian (Farsi) pipelines for spaCy},
year = {2026},
url = {https://github.com/Fazel94/spacy-persian},
note = {Models: https://huggingface.co/Phazel}
}The syntax and NER training data come from the PerDT treebank (UD_Persian-PerDT); cite it
too when you report results.
The corpora download as checksummed assets; the md and lg tiers also need a floret wheel,
described below. Python 3.12:
python -m venv .venv
.venv/bin/python -m pip install -U pip
.venv/bin/python -m pip install -r requirements.txt
.venv/bin/python -m spacy project assets # download + checksum the corpora
.venv/bin/python -m spacy project run all # -> fa_dep_news_sm + fa_core_news_sm
.venv/bin/python -m spacy project run ent # -> fa_ent_news_sm, NER alone| Command | What it does |
|---|---|
inspect |
annotation coverage of the treebanks (scripts/inspect_treebanks.py) |
convert-ud |
CoNLL-U to DocBin with --merge-subtokens, plus the tokenizer-agreement report |
transfer-ner |
align PerDT's NER layer onto that tokenization by difflib (scripts/transfer_perdt_ner.py) |
convert-ner |
transferred IOB2 to DocBin |
debug-data, debug-data-ner |
spacy debug data on both corpora before spending CPU |
train-dep |
tagger + morphologizer + trainable_lemmatizer + parser |
train-ner |
the ner component, with its own embedded tok2vec |
finalize-dep |
write fa_dep_news_sm metadata: sources, licence, notes (scripts/finalize_pipeline.py) |
evaluate-dep |
spacy benchmark accuracy on the held-out UD test split |
assemble-core |
source ner into the dep pipeline to produce fa_core_news_sm |
evaluate-core |
score the assembled pipeline on both test splits |
finalize-meta |
re-run finalize on both, folding test scores into meta.json["performance"] |
package |
build wheels + sdists for both |
smoke |
run both pipelines over Persian text and print every annotation layer |
finalize-ent |
write fa_ent_news_sm metadata from the standalone NER run |
evaluate-ent |
spacy benchmark accuracy for the NER-only package |
package-ent |
build the fa_ent_news_sm wheel + sdist |
The two training runs are single-threaded and independent, so they can run concurrently.
The vector and transformer tiers are separate workflows:
.venv/bin/python -m spacy project run md # -> fa_dep_news_md, fa_core_news_md
.venv/bin/python -m spacy project run lg # -> fa_dep_news_lg, fa_core_news_lg
.venv/bin/python -m spacy project run trf # -> fa_core_news_trf, GPU onlyThe md and lg workflows stop at the dep and core packages. fa_ent_news_lg is built by
finalize-ent-lg, evaluate-ent-lg and package-ent-lg, which no workflow calls; run them
by name.
md and lg start by unpacking a floret vector wheel that spacy project assets does not
download, because it is built by this project rather than fetched. Put it in the repo root
under the exact filename project.yml expects (vars.floret_wheel, vars.floret_lg_wheel):
curl -L -o fa_floret-0.1.0-py3-none-any-400k-documents.whl \
https://huggingface.co/Phazel/fa_floret_400k/resolve/main/fa_floret_400k-0.1.0-py3-none-any.whl
curl -L -o fa_floret-0.1.0-py3-none-any-full-wiki-200k-5epoch.whl \
https://huggingface.co/Phazel/fa-floret-wiki-vectors/resolve/main/fa_floret_wiki_200k-0.1.0-py3-none-any.whltrf additionally needs spacy-transformers and a real GPU (vars.gpu_trf is 0); the
940MX needs a cu126 torch build in a separate venv, see docs/MODELS.md §9.
sm, md and lg were trained on a 4-core i5-7200U, CPU only: sm took 1h27m for syntax
plus 17 min for NER, md 1h54m plus 25 min (the two md runs overlapped, so wall clock
overstates each), lg about 2h08m plus 13 min. trf took 1h58m on a rented Colab T4.
Standalone floret vector packages (vectors only, pipeline: []), usable as
--paths.vectors for your own training or as a plain embedding table:
# 50k rows x 300d, first 400,000 Persian Wikipedia articles (the md tier's table)
pip install https://huggingface.co/Phazel/fa_floret_400k/resolve/main/fa_floret_400k-0.1.0-py3-none-any.whl
# 50k rows x 300d, full Persian Wikipedia dump
pip install https://huggingface.co/Phazel/fa_floret_full_wiki/resolve/main/fa_floret_full_wiki-0.1.0-py3-none-any.whl
# 200k rows x 300d, full Persian Wikipedia dump, 5 epochs (the lg tier's table)
pip install https://huggingface.co/Phazel/fa-floret-wiki-vectors/resolve/main/fa_floret_wiki_200k-0.1.0-py3-none-any.whlRaw fa.floret and fa.vec exports of the lg tier's 200k-row table are in
fa-floret-wiki-vectors.
--merge-subtokens. spaCy has no multiword-token layer, and PerDT splits pronominal clitics (پدرمintoپدر+م). Measured on dev, merging gives token F 0.9887 against 0.9823 for the split version, costing 34 composite XPOS tags on 1.5% of tokens. Without it, 1.5% of gold tokens are boundaries the shipped tokenizer can never produce. Seescripts/tokenization_report.py.nercarries its own tok2vec. ATok2VecListeneronly resolves inside the pipeline it was trained in, so a listener-based component cannot be sourced elsewhere.configs/fa_ner_sm.cfgembeds the tok2vec instead, asen_core_web_smdoes.morphologizer+trainable_lemmatizerinstead ofattribute_ruler+ rule lemmatizer. The English pipelines derive UPOS from PTB tags by rule because OntoNotes has no UPOS. UD gives gold UPOS, FEATS and lemmas, which yields realpos_acc,morph_accandlemma_accnumbers instead of unmeasurable rule coverage.- PerDT, not Seraji: 3.7x more tokens, and Seraji has no
PROPNtag.
Hazm publishes spaCy-format pipelines on the HF Hub, so it was the obvious starting point. Four problems:
- Its trainable models are pycrfsuite CRFs (
hazm/sequence_tagger.py). The repo contains noconfig.cfgand nospacy train; theSpacy*classes only download pretrained pipelines. - Those pipelines are three single-task models (
transformer + tagger,transformer + parser,transformer + chunker), eachversion: 0.0.0with an emptylicensefield, pinned to spaCy 3.6. Using all three costs three ParsBERT forward passes and gives no sharedDoc. - Its tokenizer is incompatible with UD tokenization: the normaliser fuses ZWNJ (zero-width
non-joiner) affixes and
join_verb_parts()glues multi-word verb chains into single tokens. - Most corpora it reads (Bijankhan, Peykare, Hamshahri, raw PerDT) sit behind
peykaregan.irordadegan.irunder research-only terms.
It did confirm the corpus choice: Hazm's own spaCy parser was trained on
modified_fa_perdt-ud-train.spacy, the same treebank used here.
- Pipeline inventory, corpus and licence analysis:
docs/MODELS.md - How spaCy models get published, and what upstream
faalready has:docs/CONTRIBUTING-GUIDE.md - The build:
project.yml - Language data comes from
spacy/lang/faupstream, whose stop word list came from Hazm. - خلاصهٔ فارسی:
README.fa.md