A token-classification pipeline for restoring spaces in Russian text after whitespace has been removed.
Given an input such as:
книгавхорошемсостоянии
the model predicts the character positions after which spaces should be inserted and reconstructs the original phrase:
книга в хорошем состоянии
The task is treated as binary classification over the input character sequence:
0— do not insert a space after the current character;1— insert a space after the current character.
Because positive labels are much less frequent than negative labels, accuracy alone is not representative. The training pipeline evaluates precision, recall, and F1 and applies a weighted cross-entropy loss to give additional weight to whitespace positions.
Training examples are generated from the open Russian-PD corpus:
- Source text is split into shorter segments.
- Whitespace is removed to create the model input.
- Character-level labels preserve the original whitespace positions.
- Examples are shuffled and split into training and validation sets.
The current script loads 1,997 source rows from one Russian-PD Parquet shard before generating training examples.
- Base model:
cointegrated/rubert-tiny - Task head:
BertForTokenClassificationwith two labels - Frameworks: PyTorch and Hugging Face Transformers
- Maximum sequence length: 256
- Training epochs: 3
- Training batch size: 128
- Learning rate:
3e-5 - Class weights:
[1.0, 3.0] - Precision mode: BF16
After training, the script searches a grid of logit thresholds and selects the threshold with the highest validation F1. The selected value is written to spacer_model/threshold.json together with the saved model.
train.py Data preparation, training, evaluation, and threshold search
predict.py Minimal inference example
spacer_model/ Saved training artifacts and checkpoint
dataset_1937770_3.txt Additional local text sample
git clone https://github.com/Gregory-mcbit/avitomodel.git
cd avitomodel
python -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install torch transformers datasets scikit-learn numpy acceleratepython train.pyThe supplied configuration uses BF16 and was designed for an NVIDIA A100 environment. Reduce the batch size and disable BF16 when training on unsupported hardware.
After producing a complete model directory, run:
python predict.py