Skip to content

About

Data-free network-level distillation: weight transfer + closed-form output/hidden alignment + gradient distillation make a smaller student match a teacher across different depth/width/vocabulary. No datasets, synthetic probes only; same-answer guarantee is conditional and honestly reported.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

WeightForge

Codename Ferry — the pipeline name used throughout the source.

Make a smaller student model produce the same answers as a larger teacher model — even when they differ in layer count, hidden dimension, and vocabulary — without any training data. Targets come only from the teacher, driven by synthetic probes (random tensors / token ids). No datasets, no loaders, no disk I/O for data.

This is a proof of concept. The same-answer guarantee is conditional (a rank condition): it holds only when the student is wide enough to linearly reconstruct the teacher's output map. When it cannot, WeightForge reports the residual honestly instead of faking a match.

How it works

Stage What Method
0 Vocabulary reconciliation (LLM only) reconcile_vocab → VocabMap (teacher↔student token map + projection)
1 Weight transfer name-matched Copy / CropPad / SvdProject / Skip (deterministic)
2 Output alignment closed-form least squares (torch.linalg.lstsq) re-fits the last linear layer
2b Hidden alignment closed-form forward sweep over hidden linears (nonlinear teachers)
3 Gradient distillation Adam loop, fresh synthetic probe each step, teacher output as target (data-free)

Stages 1–2b are deterministic closed-form algebra. Gradient training is confined to Stage 3 (see theory.html for the full derivation and honest limits).

Layout

ferry.py                    # core pipeline + runnable 6-part toy demo (MLP / ActMLP / TinyLM)
test_ferry.py               # pytest suite (42 cases, torch-only)
ferry_qwen.py               # real-model: Qwen3-0.6B -> architecture-changed ferry-?B (CPU, data-free)
test_ferry_qwen.py          # gated tests (skip if model unavailable)
transfer_gemma_to_aster.py  # real-model: Gemma-2 -> Aster weight transfer + vocab/byte-compose
ferry_aster.py              # Aster PyTorch reproduction + data-free KD
align_aster_embed.py        # closed-form Procrustes embed-basis alignment
grow_aster_1b_to_3b.py      # stage-1 growth: aster-1b ckpt -> aster-3b initial seed (CPU, data-free)
theory.html                 # self-contained theory write-up (no dependencies)
AGENTS.md                   # project conventions and hard constraints
.agents/                    # project docs (GOAL/PLAN/TODO/PROGRESS/DECISION/MEMORY)

Run

python -m pytest test_ferry.py -q   # toy tests (42 cases, fast, no extra deps)
python ferry.py                     # toy demo, prints transfer report

# real-model extension (CPU-only, data-free) — needs transformers + accelerate
python ferry_qwen.py                # distill Qwen3-0.6B -> ferry-0.1B, report
python -m pytest test_ferry_qwen.py -q

# grow a real local checkpoint: aster-1b -> aster-3b initial seed (CPU, data-free)
python grow_aster_1b_to_3b.py --dry-run   # print the 1b->3b transfer plan only
python grow_aster_1b_to_3b.py             # write ../SLM_FROM_BEGIN/.../aster-3b-init
python -m pytest test_grow_aster_1b_to_3b.py -q

grow_aster_1b_to_3b.py applies Ferry's Stage-1 weight transfer (SVD-project + zero-pad) to grow SLM_FROM_BEGIN's trained aster-1b into an aster-3b initial seed checkpoint. It writes a fully resume-ready 4-file checkpoint (params.safetensors + state.json at global_step=0 + zero-valued AdamW optimizer.safetensors / optimizer_state.json), so the SLM trainer can start a 3B run directly: slm pretrain --resume-from <…>/aster-3b-init (begins at step 1; Muon momentum starts from zero). Pass --no-optimizer-sidecars for a model-only checkpoint.

torch is required (toy core has no other deps). ferry_qwen.py additionally needs transformers + accelerate and downloads Qwen/Qwen3-0.6B (~1.2GB) on first run. GPU is disabled by design — everything runs on CPU in float32.

Honest limits

  • Same-answer is guaranteed only under the rank condition; narrow students plateau below 100% agreement (capacity sweep in theory.html §6).
  • Nonlinear teachers need Stage 2b; shallower students cannot fully close the gap.
  • Autoregressive residual is flattened by Stage 3, not erased.
  • Real-model runs are a directional PoC ceiling, not a fluency claim.

License

Source-available, publication-reserved (strict) — see LICENSE. A narrow, revocable privilege for private, non-public, non-commercial evaluation and research only. Publication, research-credit, redistribution, commercial use, patenting, and using the Work or its outputs to train/distill or build other models are reserved to the Author and require prior written consent.

About

Data-free network-level distillation: weight transfer + closed-form output/hidden alignment + gradient distillation make a smaller student match a teacher across different depth/width/vocabulary. No datasets, synthetic probes only; same-answer guarantee is conditional and honestly reported.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages