Skip to content

fix(deps): update dependency stanza to v1.15.0 - #41

Open
renovate[bot] wants to merge 1 commit into
mainfrom
renovate/stanza-1.x
Open

renovate[bot] wants to merge 1 commit into
mainfrom
renovate/stanza-1.x

Conversation

@renovate

@renovate renovate Bot commented Sep 19, 2026 •

Copy link
Copy Markdown
Contributor

This PR contains the following updates:

Package Change Age Confidence
stanza ==1.13.0 → ==1.15.0 age confidence

Release Notes

stanfordnlp/stanza (stanza)

v1.15.0: Stanza v1.15.0 Release Notes - Updated for UD 2.18

Compare Source

Stanza v1.15.0 Release Notes

Updated Models

All tokenizer, MWT, POS, lemmatizer, and dependency parser models have been rebuilt using UD 2.18 datasets. The combined English, Spanish, and French packages have also been refreshed from the most recent development-branch snapshots to reflect recent improvements.

  • Add an NER model for Shahmukhi Punjabi (the Persian-script variety of Punjabi). #​1666

Security Fixes

Bugfixes

  • Fix multi-word token IDs becoming lists instead of tuples after a JSON round-trip through to_serialized() / from_serialized(), which produced invalid CoNLL-U output (e.g. [3, 4] instead of 3-4) and broke downstream dictionary operations. Introduced in v1.14.0; addresses #​1662. Thanks @​Arthur031221! #​1663

  • Fix empty words (enhanced UD) being misidentified as multi-word tokens when reconstructing a Document from dicts, causing incorrect token structure after a round-trip through to_dict(). Thanks @​Arthur031221! #​1664

  • Fix start/end character offsets not being assigned to words and tokens when reading a CoNLL-U Document. The fix populates offsets either by aligning tokens against the sentence text or by reading SpaceAfter annotations already present. #​1656

  • Fix a KeyError when loading the English MIMIC no-CharLM lemmatizer package: charlm_forward_file and charlm_backward_file are now treated as optional keys rather than required ones. Addresses #​1651. #​1668

Tokenizer

  • Embed external segmentation dictionaries (used by Thai, Japanese, and Chinese tokenizers) directly in the tokenizer model files, so they are not lost if models are rebuilt without them. The dictionaries are stored compressed. Adds unit tests to verify that models which should include a dictionary do. #​1688

  • Allow en-dashes and em-dashes to function as standalone tokens without suppressing comma→dash augmentation, improving the tokenizer's ability to learn dash-connected word patterns from diverse training data. #​1687

Dependency Parser

  • Add a warning when the constraint-repair loop reaches its maximum iteration count without fully resolving all violations, making incomplete repairs visible during debugging. #​1643

  • Refactor the dependency parser DataLoader to separate PyTorch and non-PyTorch code paths, enabling dynamic augmentation of individual training examples on-the-fly rather than at initialization time. This gives more balanced augmentation across training and extends coverage to silver-annotated datasets. #​1650

Training Infrastructure

  • Generalize mixed_odia_dataset.py into mixed_indic_dataset.py, which can now build combined training datasets for any low-resource Indic target language. Originally developed for Sindhi, this release also uses it for Bhojpuri. The interface replaces five language-specific flags with a single --donors parameter; a DONOR_CONFIGS dictionary at the top of the script makes adding new donor languages a one-line change. #​1655

  • Add a script for building tokenizer training sets from a mixture of multiple languages, useful for training tokenizers on related-language groups such as the Indic family. #​1672

  • Move initial-punctuation stripping from data preparation into the DataLoader for tokenizer, POS tagger, and dependency parser, so it is applied dynamically during training rather than baked into data files. #​1653

  • Integrate speaker information from ingested UDCoref documents into the coreference training data pipeline. #​1645

  • Add a conversion script for the IIT (BHU) Bhojpuri POS corpus, transforming it from its mixed flat/SSF-bracket format into a standardized one-token-per-line layout. Thanks @​abhiprd2000! #​1675

  • Add multi-column xpos tagging support, allowing the tagger to train across multiple datasets with differing xpos schemes simultaneously. Applied to Bhojpuri (BHTB + IIT corpus), where it yields substantial improvements. Experiments on English (ParTUT and LinES) showed no benefit — English xpos accuracy is already saturated around 97.4%, which makes it a poor test case for a technique aimed at low-resource settings with small main treebanks. #​1680

CoreNLP Integration

  • Upgrade the CoreNLP semgrex communication protocol to support enhanced queries. #​1685

  • Update the CoreNLP installation script to report what was installed, be more conservative about which version to download, and fix a bug in the DEFAULT_CORENLP_URL constant. #​1686

Contributors

v1.14.0: - Security fixes and Lemmatizer efficiency updates

Compare Source

Stanza v1.14.0 Release Notes

Security Fixes

  • Fix a potential zip slip vulnerability when extracting downloaded model archives. While low-risk given that Stanza controls the resources being downloaded, extraction now validates that no file paths escape the target directory. See GHSA-2fwf-f686-7p34. #​1621

  • Restrict the unpickler used when deserializing annotated Documents, and add a deprecation warning: in a future release, Document serialization will move to JSON entirely, removing the pickle dependency. See GHSA-487q-m798-cp85. #​1626

  • Remove shell subprocess calls from make_lm_data.py, addressing GHSA-c9h2-qmqw-qf6h. As a side benefit, the charlm data preparation script is now fully portable to Windows. #​1623

Bugfixes

  • Fix a bug where coreference chain annotations were written under the ner= MISC key instead of coref_chains= when serializing a Document to CoNLL-U, causing collision with real NER labels on the same token. Thank you @​devteamaegis! #​1628

New / Updated Models

  • Substantially updated Slovenian models: the new default sl_combined package mixes the SSJ and SST treebanks (reported by Kaja Dobrovoljc to be highly compatible), augments lemma and POS training with SUK 1.1 data, builds a lemma dictionary from Sloleks 3.1, and adds contextual lemma classifiers for the ambiguous pairs del/delo and rok/roka. #​1625

Lemmatizer Improvements

  • Reorganize the lemmatizer dictionary to use a pos → word → lemma layout and store it gzip-compressed. This dramatically reduces load time for large models — Slovenian drops from 30+ seconds to under 5 seconds — and shrinks model sizes considerably. A conversion script for updating locally trained 1.13.0 models is included.

    Note: lemmatizer models from v1.13.0 are not compatible with v1.14.0. Please re-download or convert existing models. #​1627

  • Reduce the hidden dimension of the contextual lemma classifier, making models smaller and faster without hurting accuracy. #​1629

Dependency Parser

  • Post-process dependency parses to enforce uniqueness constraints on nsubj/csubj and obj relations: if the graph parser produces a node with multiple subjects or direct objects, the parser now reruns Chu-Liu-Edmonds iteratively (reusing the original neural scores) to find the best-scoring repair. This is on by default in the Pipeline. Addresses #​1340. #​1638

Tokenizer

  • Move comma-transposition augmentations from the data preparation script into the DataLoader, so that augmentation is applied on-the-fly during training rather than being baked in once at preprocessing time. This produces more balanced training and avoids accidentally affecting other annotators' data files. #​1624

  • Add new structural feature functions to the tokenizer to help distinguish address-line formatting from normal running text, laying groundwork for fixing sentence-splitting errors on non-prose inputs. Addresses #​1640. #​1642

Interface Improvements

  • Add a stanza.utils.list_installed script that lists all locally cached Stanza models and their versions, without modifying anything on disk. Addresses #​1542. #​1632

  • Add a tokenize_with_speakers() convenience function for processing transcript-style text where each line begins with a speaker label, automatically assigning speaker metadata to sentences before passing them to the coref annotator. #​1631

CharLM Training Infrastructure

For researchers building character language models for new languages, this release includes updated tooling for collecting and deduplicating training data from OSCAR. The previous OSCAR 2023 source is no longer accessible to new users and is broken with datasets >= 4.0; the new scripts target the OSCAR Community Crawl instead. Addresses #​1622.

  • Add a download script for the OSCAR Community Crawl that bypasses load_dataset (which has a known bug with OSCAR), along with an inventory script to inspect the language breakdown of downloaded chunks. Also adds OSCAR language codes to constant.py and fixes a bug where extra language name aliases were being silently clobbered. #​1633

  • Switch the near-deduplication strategy from TLSH to MinHash LSH. MinHash is faster, retains more content, and still achieves satisfactory deduplication rates as verified by the diagnostic script included in this PR. #​1639

Contributors


Configuration

📅 Schedule: (UTC)

  • Branch creation
    • At any time (no schedule defined)
  • Automerge
    • At any time (no schedule defined)

🚦 Automerge: Disabled by config. Please merge this manually once you are satisfied.

♻ Rebasing: Whenever PR is behind base branch, or you tick the rebase/retry checkbox.

🔕 Ignore: Close this PR and you won't be reminded about this update again.


  • If you want to rebase/retry this PR, check this box

This PR was generated by Mend Renovate. View the repository job log.

@codecov

codecov Bot commented Sep 19, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 82.72%. Comparing base (909354f) to head (ed45432).

Additional details and impacted files

Impacted file tree graph

@@           Coverage Diff           @@
##             main      #41   +/-   ##
=======================================
  Coverage   82.72%   82.72%           
=======================================
  Files          71       71           
  Lines        3207     3207           
=======================================
  Hits         2653     2653           
  Misses        554      554           
🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@renovate
renovate Bot force-pushed the renovate/stanza-1.x branch 2 times, most recently from d89b1cd to 3442d7d Compare September 28, 2026 10:42
@stefano81
stefano81 enabled auto-merge (rebase) September 28, 2026 10:43
@renovate
renovate Bot force-pushed the renovate/stanza-1.x branch from 3442d7d to c8b2588 Compare September 28, 2026 23:25
@renovate
renovate Bot force-pushed the renovate/stanza-1.x branch from c8b2588 to ed45432 Compare October 1, 2026 06:04
@renovate renovate Bot changed the title fix(deps): update dependency stanza to v1.14.0 fix(deps): update dependency stanza to v1.15.0 Oct 1, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

0 participants