Skip to content
MadB0iPublic

About

VLM fine-tuning on consumer GPUs via CPU<->GPU layer streaming through the backward pass

Resources

Stars

1 star

Watchers

0 watching

Forks

Latest commit

 

History

14 Commits

Folders and files

Repository files navigation

🔥 Wick

VLM fine-tuning on low-VRAM consumer GPUs via CPU↔GPU layer streaming — through the backward pass.

Wick lets you fine-tune Vision-Language Models on hardware that was never meant for it (think 4 GB GTX 1650s) by streaming the vision encoder's layers in and out of VRAM, one layer at a time — not just for inference, but for training itself.

Inference-side streaming was already solved (AirLLM, oLLM). Wick solves the harder half nobody's cracked yet: computing gradients through streamed layers.

65 tests passing | GTX 1650 4GB | fp16 stable | real SigLIP streamed


🚀 Quick Demo

========================================================================
WICK DEVICE PROFILER
========================================================================
Your GPU: NVIDIA GeForce GTX 1650
  VRAM      4.00 GB (4.00 GB)
  compute   7.5 (no bf16, fp16)
  system RAM 15 GB (host masters / optimizer offload)

  Peak-VRAM model (Wick streaming): 1 vision layer fp16 + resident LLM
  + LoRA fp32 + LoRA optimizer + ~1-layer activation + 10% headroom

  ✅ MiniCPM-V 1.3B     best(fp16)   2.60 GB of 4.00 GB VRAM
          fp16 needs   2.60 GB, int4 needs   0.82 GB
  ✅ MiniCPM-V 2.6B     best(int4)   1.69 GB of 4.00 GB VRAM
          fp16 needs   5.56 GB, int4 needs   1.69 GB
  ✅ Qwen2-VL 2B        best(int4)   1.35 GB of 4.00 GB VRAM
          fp16 needs   4.32 GB, int4 needs   1.35 GB
  ❌ LLaVA-1.5 7B       best(int4)   4.49 GB of 4.00 GB VRAM
          fp16 needs  14.89 GB, int4 needs   4.49 GB
  ✅ Phi-3.5 Vision     best(int4)   2.69 GB of 4.00 GB VRAM
          fp16 needs   8.93 GB, int4 needs   2.69 GB

  ✅ Can fine-tune:
     - MiniCPM-V 1.3B
     - MiniCPM-V 2.6B
     - Qwen2-VL 2B
     - Phi-3.5 Vision
  ❌ Cannot fit:
     - LLaVA-1.5 7B
========================================================================

Run it yourself with python -m wick.profiler.


✨ Why Wick?

| | AirLLM / oLLM (inference-only) | Wick (training) | |---|---| | Forward pass streaming | ✅ one layer at a time | ✅ one layer at a time | | Backward pass streaming | ❌ not supported | ✅ **bit-exact gradients | | Peak VRAM residency | ~1 layer | ~1 layer | | Fine-tuning | ❌ | ✅ LoRA / QLoRA |


🧪 What's Proven

  • ✅ Phase 2 — backward-pass streaming is gradient-exact. rel_diff_norm = 0.0 vs a fully-resident baseline (bitwise), fp64, CPU sim.
  • ✅ Peak residency = 1 layer. No device storage leaks into the autograd graph (40 saved tensors walked via data_ptr, non-vacuous check).
  • ✅ AirLLM-style hooks rejected with evidence. Hooks retain device allocations and produce NaN gradients under storage pressure. A custom torch.autograd.Function is required — and proven.
  • ✅ Phase 3 — streamed LoRA training (frozen base + resident adapters) yields bit-identical loss trajectories vs full-resident (fp64, CPU).
  • ✅ Sub-gate 4A — real GTX 1650, fp16 + GradScaler. Stable scale (final 4e3, init 2^12), zero overflow across 100 streamed steps / 4800 real PCIe load-evict cycles; peak weight residency 1.00 fp16 layer; loss 7.3e-4 → 2.8e-6.
  • ✅ Sub-gate 4B — real SigLIP SO-400M encoder, real GTX 1650. All 27 encoder layers (15.24 M params, 29.07 MB fp16 each) streamed layer-by-layer through both forward and backward. Peak VRAM 84.8 MiB (≪ 2339 MiB budget); residency 1.00 layer; GradScaler stable (65536, 0 overflow); loss 1.38 → 9.68e-6 over 100 steps; 43,200 real PCIe load/evict cycles; device empty after.
  • ✅ The gate can fail — injected fault types are all caught.

🚧 What's Not Proven Yet

  • ⏳ Sub-gate 4B 1000-step full gate. The 100-step smoke gate passes on real SigLIP; the 1000-step loss-curve-within-5% run vs a full-VRAM baseline is pending.
  • ⏳ LLM backbone not yet resident. int4 (bitsandbytes) quantization of the MiniCPM LLM + LoRA wiring to the streamed encoder is not built. 4B used a small stand-in trainable head, not the real LLM.
  • ⏳ Realistic-scale peak VRAM. The 84.8 MiB peak is frozen-SigLIP + small head only. With int4 LLM (~1 GB) + LoRA resident, the honest number is the ~2339 MiB fit-check budget — not yet measured end-to-end on the GPU.
  • ⏳ Real dataset / task loss not tested (synthetic random target used in 4B).

🎯 Target

MiniCPM-V 1.3B — stream only the vision encoder (~400 M params), keep the LLM

  • LoRA adapters resident in VRAM.

GPU floor: GTX 1650 · 4 GB VRAM · TU117 die — no bf16, no tensor cores. Mixed precision is fp16 + GradScaler only; fp32 master weights for streamed layers live on CPU always.


🛠️ Stack

Python + PyTorch (torch.autograd.Function) · pure local, no cloud calls. Real-GPU runs need a CUDA build — for the GTX 1650 (sm_75) that is pip install torch==2.6.0+cu124 --index-url https://download.pytorch.org/whl/cu124.


📅 Roadmap

  • Phase 2 ✅ — backward-pass streaming proof (gate passes)
  • Phase 3 ✅ — LoRA training via streamed backward (CPU sim, exact)
  • Phase 4A ✅ — real GTX 1650: fp16 + GradScaler stability, 1-layer residency
  • Phase 4B ✅ (smoke gate) — real SigLIP encoder streamed on GTX 1650 (100-step gate passes; 1000-step full gate next)
  • Phase 5 ⏳ — int4 LLM + LoRA wiring, 1000-step gate vs full-VRAM baseline, real dataset, benchmarks

📊 Results — Sub-gate 4A (real GTX 1650)

Measured on a real NVIDIA GeForce GTX 1650 (4 GB, sm_75, fp16 only), 100 streamed training steps with Adam + GradScaler.

Metric Value
GradScaler final scale 4e3 (init 2^12 = 4096)
Overflow / skipped steps 0
Peak weight residency 1.00 fp16 layer (97.6 KiB; all-resident would need 195.2 KiB)
Loss trajectory 7.35e-4 → 2.77e-6 over 100 steps
Real PCIe load / evict cycles 4800 / 4800
Device bytes resident after run 0.00 MiB
==============================================================================
SUB-GATE 4A PASSED --
  final GradScaler scale 4e+03 (finite, stable, 0 overflow steps)
  peak weight residency 97.6 KiB = 1.00 fp16 layers (bounded at ~1; all-resident would need 195.2 KiB)
  loss 7.349e-04 -> 2.775e-06 over 100 steps
  4800 real PCIe loads, 4800 evicts, device empty at end
==============================================================================

Run it yourself with python scripts/run_phase4a.py.


📊 Results — Sub-gate 4B (real SigLIP encoder, real GTX 1650)

Real SigLIP SO-400M vision encoder (MiniCPM-V 1.3B's tower) — 27 layers, 15,239,504 params each (29.07 MB fp16) — streamed layer-by-layer through both forward and backward, 100 training steps, fp16 + GradScaler.

Metric Value
Peak VRAM 84.8 MiB (≪ 2339 MiB budget)
Peak weight residency 1.00 SigLIP layer
GradScaler final scale 65536 (finite, stable)
Overflow / skipped steps 0
Loss trajectory 1.38 → 9.68e-6 over 100 steps
Real PCIe load / evict cycles 43200 / 43200
Device bytes resident after run 36.5 MiB (CUDA context only; 0 weight bytes)
==============================================================================
SUB-GATE 4B PASSED --
  peak VRAM 84.8 MiB < 2339 MiB ceiling
  peak residency 1.00 layers
  GradScaler 65536 (finite, stable)
  overflow=0 skipped=0
  loss 1.3265e+00 -> 9.6845e-06 over 100 steps
  device empty: True
==============================================================================

Note: 4B trains a small stand-in head on the frozen SigLIP output to exercise GradScaler end-to-end. The int4 LLM + LoRA integration (Phase 5) is the next step.

Run it yourself with python -m wick.phase4b (needs the local SigLIP weights in models/siglip-so400m-patch14-384/).


📊 Results — Phase 5 Stage 4, Finding 1 (T=129, realistic length)

End-to-end on the real GTX 1650, seed=1234, grad checkpointing ON for both arms:

Arm Result
Streamed (1 layer at a time) 1000/1000 steps, loss 0.479 → 0.274, peak flat 2870 MiB, GradScaler 65536, 0 overflows
Full-resident (27 layers) Tripped the 3686 MiB tripwire at step 2 (peak 3700 MiB), stopped per protocol — no retry, no line-moving

Recorded as: at realistic sequence length, streaming trains reliably where full-resident approaches the VRAM ceiling. The apples-to-apples loss-curve gate (<5%) moves to T=79, a length where both arms fit (Finding 2).

Finding 2 (T=79, seed=1234, both arms fresh 0→1000)

Arm Result
Streamed 1000/1000 steps, loss 2.1035 → 1.2304, peak flat 2766 MiB, GradScaler 65536, 0 overflows. No trip.
Full-resident 1000/1000 steps, loss 2.1035 → 1.2304, peak flat 3596 MiB, GradScaler 65536, 0 overflows. No trip (90 MiB under the line).

Max relative loss difference across all 20 logged points: 0.0000% — bit-identical trajectories on the real GPU (gate needed < 5%). Four gate numbers: (1) 0.0000% PASS · (2) peaks 2766/3596, both under the 3686 tripwire · (3) overflows 0/0 · (4) neither tripped. Verdict: FULL 1000-step gate PASSES at T=79.


🚀 Get Started

pip install -e .
# Phase 2 gate report (exit 0 = pass)
python scripts/run_gate.py
# Phase 3 gate report (LoRA via streamed backward, CPU sim)
python scripts/run_phase3.py
# Sub-gate 4A report (real GPU: fp16 + GradScaler stability) -- needs CUDA build
python scripts/run_phase4a.py
# Sub-gate 4B report (real GPU: real SigLIP streaming) -- needs CUDA + local weights
python -m wick.phase4b
# Device profiler: which VLMs fit your hardware (no model download)
python -m wick.profiler
# Full assertion suite
python -m pytest tests/ -v

Live Monitor

streamlit run scripts/monitor.py

Opens the Wick Monitor browser tab: pick one or more run logs from logs/ in the sidebar and see summary cards (peak VRAM, steps, final loss, overflows, pass/fail vs the tripwire), loss-over-step charts (select streamed + resident to overlay both), and VRAM-over-step with the tripwire line. If a log is still being written, the page auto-refreshes every few seconds, so it works as a live viewer during training too.

Note: this is for viewing logs — it works alongside or after the terminal dashboard (--live-dashboard), it doesn't replace it.


📄 License

Apache-2.0 · built for the low-VRAM community.

About

VLM fine-tuning on consumer GPUs via CPU<->GPU layer streaming through the backward pass

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages