A decoder-only language model built from scratch in PyTorch, featuring a custom BPE tokenizer, Rotary Positional Embeddings (RoPE), Grouped-Query Attention (GQA), SwiGLU activations, Flash Attention, and sharded dataset streaming. The project covers the full training pipeline, from large-scale pretraining to supervised fine-tuning (SFT).
Try PlutoAI directly in your browser:
Decoder-only Transformer
12 Layers
768 Hidden Size
12 Attention Heads
4 KV Heads (GQA)
SwiGLU
RoPE
Flash Attention
Weight Tying
32K Vocabulary
1,024 Context Length
The model uses a 2944-dimensional SwiGLU feed-forward layer with a GPU-friendly multiple-of-128 hidden dimension.
A custom BPE tokenizer:
Name : pluto_ai_bpe
Vocabulary : 32,000
Context : 1,024
Special tokens support chat structure and reasoning-style delimiters such as:
<|start_header_id|>
<|end_header_id|>
<|end_turn|>
<|think|>
<|end_think|>
Raw Data
↓
Filtering + Tokenization
↓
Sharded Dataset
↓
Pretraining
↓
Supervised Fine-Tuning
↓
Inference
Currently based on FineWeb-Edu, with approximately 2.36B tokens after processing.
Quality Score : ≥ 2.5
Token Count : ≥ 128
Language Score : ≥ 0.85
Validation : 1%
SFT currently uses:
Dataset 1 : 50%
Dataset 2 : 40%
Replay : 10%
Training Tokens : 100M
Validation : 10M
Epochs : 3
Learning Rate : 1e-5
Datasets are stored as binary shards instead of requiring the entire processed corpus to reside in RAM.
50M tokens / shard
1,024 tokens / sequence
.bin → token data
.idx → sequence/index metadata
.mask → SFT loss masks
Built with PyTorch and supports:
- Mixed precision
- Gradient accumulation
- Flash Attention
- Sequence packing
- Sharded data loading
- Gradient clipping
- Checkpoint recovery
- TensorBoard
- Multi-GPU DDP
PlutoAI/
├── model/
├── tokenizer/
├── training/
├── fine_tuning/
├── inference/
├── shards/
├── scripts/
├── testing/
├── data/
├── checkpoints/
├── logs/
└── config.py
The repository keeps the model, data pipeline, training system, fine-tuning, inference, and tests separated into their own modules.
Build the model.
Understand the data.
Own the training loop.
Optimize where it matters.
PlutoAI is primarily an engineering and research project focused on understanding how a modern language model is built end-to-end.
Requires uv.
git clone https://github.com/Shaahmir/PlutoAI.git
cd PlutoAI
uv syncStart training with the configured pipeline:
uv run python -m training.trainRun the inference pipeline:
uv run python -m inferenceTraining supports checkpoint recovery, gradient accumulation, cosine learning-rate scheduling, mixed precision, and TensorBoard logging.
uv run tensorboard --logdir logsPlutoAI follows the same from-scratch philosophy as my earlier conversational model project:
ChatRNN — a modular Seq2Seq conversational model implemented in PyTorch with a Bidirectional LSTM encoder, Bahdanau Attention, LSTM decoder, SentencePiece tokenization, sharded datasets, AMP, and beam-search inference.
This project is licensed under the MIT License.
