Codename Ferry — the pipeline name used throughout the source.
Make a smaller student model produce the same answers as a larger teacher model — even when they differ in layer count, hidden dimension, and vocabulary — without any training data. Targets come only from the teacher, driven by synthetic probes (random tensors / token ids). No datasets, no loaders, no disk I/O for data.
This is a proof of concept. The same-answer guarantee is conditional (a rank condition): it holds only when the student is wide enough to linearly reconstruct the teacher's output map. When it cannot, WeightForge reports the residual honestly instead of faking a match.
| Stage | What | Method |
|---|---|---|
| 0 | Vocabulary reconciliation (LLM only) | reconcile_vocab → VocabMap (teacher↔student token map + projection) |
| 1 | Weight transfer | name-matched Copy / CropPad / SvdProject / Skip (deterministic) |
| 2 | Output alignment | closed-form least squares (torch.linalg.lstsq) re-fits the last linear layer |
| 2b | Hidden alignment | closed-form forward sweep over hidden linears (nonlinear teachers) |
| 3 | Gradient distillation | Adam loop, fresh synthetic probe each step, teacher output as target (data-free) |
Stages 1–2b are deterministic closed-form algebra. Gradient training is confined
to Stage 3 (see theory.html for the full derivation and honest limits).
ferry.py # core pipeline + runnable 6-part toy demo (MLP / ActMLP / TinyLM)
test_ferry.py # pytest suite (42 cases, torch-only)
ferry_qwen.py # real-model: Qwen3-0.6B -> architecture-changed ferry-?B (CPU, data-free)
test_ferry_qwen.py # gated tests (skip if model unavailable)
transfer_gemma_to_aster.py # real-model: Gemma-2 -> Aster weight transfer + vocab/byte-compose
ferry_aster.py # Aster PyTorch reproduction + data-free KD
align_aster_embed.py # closed-form Procrustes embed-basis alignment
grow_aster_1b_to_3b.py # stage-1 growth: aster-1b ckpt -> aster-3b initial seed (CPU, data-free)
theory.html # self-contained theory write-up (no dependencies)
AGENTS.md # project conventions and hard constraints
.agents/ # project docs (GOAL/PLAN/TODO/PROGRESS/DECISION/MEMORY)
python -m pytest test_ferry.py -q # toy tests (42 cases, fast, no extra deps)
python ferry.py # toy demo, prints transfer report
# real-model extension (CPU-only, data-free) — needs transformers + accelerate
python ferry_qwen.py # distill Qwen3-0.6B -> ferry-0.1B, report
python -m pytest test_ferry_qwen.py -q
# grow a real local checkpoint: aster-1b -> aster-3b initial seed (CPU, data-free)
python grow_aster_1b_to_3b.py --dry-run # print the 1b->3b transfer plan only
python grow_aster_1b_to_3b.py # write ../SLM_FROM_BEGIN/.../aster-3b-init
python -m pytest test_grow_aster_1b_to_3b.py -qgrow_aster_1b_to_3b.py applies Ferry's Stage-1 weight transfer (SVD-project +
zero-pad) to grow SLM_FROM_BEGIN's trained aster-1b into an aster-3b initial
seed checkpoint. It writes a fully resume-ready 4-file checkpoint
(params.safetensors + state.json at global_step=0 + zero-valued AdamW
optimizer.safetensors / optimizer_state.json), so the SLM trainer can start a
3B run directly: slm pretrain --resume-from <…>/aster-3b-init (begins at step 1;
Muon momentum starts from zero). Pass --no-optimizer-sidecars for a model-only
checkpoint.
torch is required (toy core has no other deps). ferry_qwen.py additionally
needs transformers + accelerate and downloads Qwen/Qwen3-0.6B (~1.2GB) on
first run. GPU is disabled by design — everything runs on CPU in float32.
- Same-answer is guaranteed only under the rank condition; narrow students plateau
below 100% agreement (capacity sweep in
theory.html§6). - Nonlinear teachers need Stage 2b; shallower students cannot fully close the gap.
- Autoregressive residual is flattened by Stage 3, not erased.
- Real-model runs are a directional PoC ceiling, not a fluency claim.
Source-available, publication-reserved (strict) — see LICENSE.
A narrow, revocable privilege for private, non-public, non-commercial evaluation
and research only. Publication, research-credit, redistribution, commercial use,
patenting, and using the Work or its outputs to train/distill or build other
models are reserved to the Author and require prior written consent.