Skip to content

Latest commit

 

History

4 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Russian Word Segmentation with ruBERT

A token-classification pipeline for restoring spaces in Russian text after whitespace has been removed.

Given an input such as:

книгавхорошемсостоянии

the model predicts the character positions after which spaces should be inserted and reconstructs the original phrase:

книга в хорошем состоянии

Problem formulation

The task is treated as binary classification over the input character sequence:

  • 0 — do not insert a space after the current character;
  • 1 — insert a space after the current character.

Because positive labels are much less frequent than negative labels, accuracy alone is not representative. The training pipeline evaluates precision, recall, and F1 and applies a weighted cross-entropy loss to give additional weight to whitespace positions.

Dataset

Training examples are generated from the open Russian-PD corpus:

  1. Source text is split into shorter segments.
  2. Whitespace is removed to create the model input.
  3. Character-level labels preserve the original whitespace positions.
  4. Examples are shuffled and split into training and validation sets.

The current script loads 1,997 source rows from one Russian-PD Parquet shard before generating training examples.

Model and training

  • Base model: cointegrated/rubert-tiny
  • Task head: BertForTokenClassification with two labels
  • Frameworks: PyTorch and Hugging Face Transformers
  • Maximum sequence length: 256
  • Training epochs: 3
  • Training batch size: 128
  • Learning rate: 3e-5
  • Class weights: [1.0, 3.0]
  • Precision mode: BF16

After training, the script searches a grid of logit thresholds and selects the threshold with the highest validation F1. The selected value is written to spacer_model/threshold.json together with the saved model.

Repository structure

train.py              Data preparation, training, evaluation, and threshold search
predict.py            Minimal inference example
spacer_model/         Saved training artifacts and checkpoint
dataset_1937770_3.txt Additional local text sample

Installation

git clone https://github.com/Gregory-mcbit/avitomodel.git
cd avitomodel

python -m venv .venv
source .venv/bin/activate  # Windows: .venv\Scripts\activate

pip install torch transformers datasets scikit-learn numpy accelerate

Training

python train.py

The supplied configuration uses BF16 and was designed for an NVIDIA A100 environment. Reduce the batch size and disable BF16 when training on unsupported hardware.

Inference

After producing a complete model directory, run:

python predict.py

About

Transformer-based Russian word segmentation and whitespace restoration using ruBERT-tiny and PyTorch.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages