VLM fine-tuning on low-VRAM consumer GPUs via CPU↔GPU layer streaming — through the backward pass.
Wick lets you fine-tune Vision-Language Models on hardware that was never meant for it (think 4 GB GTX 1650s) by streaming the vision encoder's layers in and out of VRAM, one layer at a time — not just for inference, but for training itself.
Inference-side streaming was already solved (AirLLM, oLLM). Wick solves the harder half nobody's cracked yet: computing gradients through streamed layers.
65 tests passing | GTX 1650 4GB | fp16 stable | real SigLIP streamed
========================================================================
WICK DEVICE PROFILER
========================================================================
Your GPU: NVIDIA GeForce GTX 1650
VRAM 4.00 GB (4.00 GB)
compute 7.5 (no bf16, fp16)
system RAM 15 GB (host masters / optimizer offload)
Peak-VRAM model (Wick streaming): 1 vision layer fp16 + resident LLM
+ LoRA fp32 + LoRA optimizer + ~1-layer activation + 10% headroom
✅ MiniCPM-V 1.3B best(fp16) 2.60 GB of 4.00 GB VRAM
fp16 needs 2.60 GB, int4 needs 0.82 GB
✅ MiniCPM-V 2.6B best(int4) 1.69 GB of 4.00 GB VRAM
fp16 needs 5.56 GB, int4 needs 1.69 GB
✅ Qwen2-VL 2B best(int4) 1.35 GB of 4.00 GB VRAM
fp16 needs 4.32 GB, int4 needs 1.35 GB
❌ LLaVA-1.5 7B best(int4) 4.49 GB of 4.00 GB VRAM
fp16 needs 14.89 GB, int4 needs 4.49 GB
✅ Phi-3.5 Vision best(int4) 2.69 GB of 4.00 GB VRAM
fp16 needs 8.93 GB, int4 needs 2.69 GB
✅ Can fine-tune:
- MiniCPM-V 1.3B
- MiniCPM-V 2.6B
- Qwen2-VL 2B
- Phi-3.5 Vision
❌ Cannot fit:
- LLaVA-1.5 7B
========================================================================
Run it yourself with python -m wick.profiler.
| | AirLLM / oLLM (inference-only) | Wick (training) | |---|---| | Forward pass streaming | ✅ one layer at a time | ✅ one layer at a time | | Backward pass streaming | ❌ not supported | ✅ **bit-exact gradients | | Peak VRAM residency | ~1 layer | ~1 layer | | Fine-tuning | ❌ | ✅ LoRA / QLoRA |
- ✅ Phase 2 — backward-pass streaming is gradient-exact.
rel_diff_norm = 0.0vs a fully-resident baseline (bitwise), fp64, CPU sim. - ✅ Peak residency = 1 layer. No device storage leaks into the autograd graph
(40 saved tensors walked via
data_ptr, non-vacuous check). - ✅ AirLLM-style hooks rejected with evidence. Hooks retain device allocations
and produce NaN gradients under storage pressure. A custom
torch.autograd.Functionis required — and proven. - ✅ Phase 3 — streamed LoRA training (frozen base + resident adapters) yields bit-identical loss trajectories vs full-resident (fp64, CPU).
- ✅ Sub-gate 4A — real GTX 1650, fp16 + GradScaler. Stable scale (final 4e3, init 2^12), zero overflow across 100 streamed steps / 4800 real PCIe load-evict cycles; peak weight residency 1.00 fp16 layer; loss 7.3e-4 → 2.8e-6.
- ✅ Sub-gate 4B — real SigLIP SO-400M encoder, real GTX 1650. All 27 encoder layers (15.24 M params, 29.07 MB fp16 each) streamed layer-by-layer through both forward and backward. Peak VRAM 84.8 MiB (≪ 2339 MiB budget); residency 1.00 layer; GradScaler stable (65536, 0 overflow); loss 1.38 → 9.68e-6 over 100 steps; 43,200 real PCIe load/evict cycles; device empty after.
- ✅ The gate can fail — injected fault types are all caught.
- ⏳ Sub-gate 4B 1000-step full gate. The 100-step smoke gate passes on real SigLIP; the 1000-step loss-curve-within-5% run vs a full-VRAM baseline is pending.
- ⏳ LLM backbone not yet resident. int4 (bitsandbytes) quantization of the MiniCPM LLM + LoRA wiring to the streamed encoder is not built. 4B used a small stand-in trainable head, not the real LLM.
- ⏳ Realistic-scale peak VRAM. The 84.8 MiB peak is frozen-SigLIP + small head only. With int4 LLM (~1 GB) + LoRA resident, the honest number is the ~2339 MiB fit-check budget — not yet measured end-to-end on the GPU.
- ⏳ Real dataset / task loss not tested (synthetic random target used in 4B).
MiniCPM-V 1.3B — stream only the vision encoder (~400 M params), keep the LLM
- LoRA adapters resident in VRAM.
GPU floor: GTX 1650 · 4 GB VRAM · TU117 die — no bf16, no tensor cores. Mixed precision is fp16 + GradScaler only; fp32 master weights for streamed layers live on CPU always.
Python + PyTorch (torch.autograd.Function) · pure local, no cloud calls.
Real-GPU runs need a CUDA build — for the GTX 1650 (sm_75) that is
pip install torch==2.6.0+cu124 --index-url https://download.pytorch.org/whl/cu124.
- Phase 2 ✅ — backward-pass streaming proof (gate passes)
- Phase 3 ✅ — LoRA training via streamed backward (CPU sim, exact)
- Phase 4A ✅ — real GTX 1650: fp16 + GradScaler stability, 1-layer residency
- Phase 4B ✅ (smoke gate) — real SigLIP encoder streamed on GTX 1650 (100-step gate passes; 1000-step full gate next)
- Phase 5 ⏳ — int4 LLM + LoRA wiring, 1000-step gate vs full-VRAM baseline, real dataset, benchmarks
Measured on a real NVIDIA GeForce GTX 1650 (4 GB, sm_75, fp16 only), 100 streamed training steps with Adam + GradScaler.
| Metric | Value |
|---|---|
| GradScaler final scale | 4e3 (init 2^12 = 4096) |
| Overflow / skipped steps | 0 |
| Peak weight residency | 1.00 fp16 layer (97.6 KiB; all-resident would need 195.2 KiB) |
| Loss trajectory | 7.35e-4 → 2.77e-6 over 100 steps |
| Real PCIe load / evict cycles | 4800 / 4800 |
| Device bytes resident after run | 0.00 MiB |
==============================================================================
SUB-GATE 4A PASSED --
final GradScaler scale 4e+03 (finite, stable, 0 overflow steps)
peak weight residency 97.6 KiB = 1.00 fp16 layers (bounded at ~1; all-resident would need 195.2 KiB)
loss 7.349e-04 -> 2.775e-06 over 100 steps
4800 real PCIe loads, 4800 evicts, device empty at end
==============================================================================
Run it yourself with python scripts/run_phase4a.py.
Real SigLIP SO-400M vision encoder (MiniCPM-V 1.3B's tower) — 27 layers, 15,239,504 params each (29.07 MB fp16) — streamed layer-by-layer through both forward and backward, 100 training steps, fp16 + GradScaler.
| Metric | Value |
|---|---|
| Peak VRAM | 84.8 MiB (≪ 2339 MiB budget) |
| Peak weight residency | 1.00 SigLIP layer |
| GradScaler final scale | 65536 (finite, stable) |
| Overflow / skipped steps | 0 |
| Loss trajectory | 1.38 → 9.68e-6 over 100 steps |
| Real PCIe load / evict cycles | 43200 / 43200 |
| Device bytes resident after run | 36.5 MiB (CUDA context only; 0 weight bytes) |
==============================================================================
SUB-GATE 4B PASSED --
peak VRAM 84.8 MiB < 2339 MiB ceiling
peak residency 1.00 layers
GradScaler 65536 (finite, stable)
overflow=0 skipped=0
loss 1.3265e+00 -> 9.6845e-06 over 100 steps
device empty: True
==============================================================================
Note: 4B trains a small stand-in head on the frozen SigLIP output to exercise GradScaler end-to-end. The int4 LLM + LoRA integration (Phase 5) is the next step.
Run it yourself with python -m wick.phase4b (needs the local SigLIP weights in
models/siglip-so400m-patch14-384/).
End-to-end on the real GTX 1650, seed=1234, grad checkpointing ON for both arms:
| Arm | Result |
|---|---|
| Streamed (1 layer at a time) | 1000/1000 steps, loss 0.479 → 0.274, peak flat 2870 MiB, GradScaler 65536, 0 overflows |
| Full-resident (27 layers) | Tripped the 3686 MiB tripwire at step 2 (peak 3700 MiB), stopped per protocol — no retry, no line-moving |
Recorded as: at realistic sequence length, streaming trains reliably where full-resident approaches the VRAM ceiling. The apples-to-apples loss-curve gate (<5%) moves to T=79, a length where both arms fit (Finding 2).
| Arm | Result |
|---|---|
| Streamed | 1000/1000 steps, loss 2.1035 → 1.2304, peak flat 2766 MiB, GradScaler 65536, 0 overflows. No trip. |
| Full-resident | 1000/1000 steps, loss 2.1035 → 1.2304, peak flat 3596 MiB, GradScaler 65536, 0 overflows. No trip (90 MiB under the line). |
Max relative loss difference across all 20 logged points: 0.0000% — bit-identical trajectories on the real GPU (gate needed < 5%). Four gate numbers: (1) 0.0000% PASS · (2) peaks 2766/3596, both under the 3686 tripwire · (3) overflows 0/0 · (4) neither tripped. Verdict: FULL 1000-step gate PASSES at T=79.
pip install -e .
# Phase 2 gate report (exit 0 = pass)
python scripts/run_gate.py
# Phase 3 gate report (LoRA via streamed backward, CPU sim)
python scripts/run_phase3.py
# Sub-gate 4A report (real GPU: fp16 + GradScaler stability) -- needs CUDA build
python scripts/run_phase4a.py
# Sub-gate 4B report (real GPU: real SigLIP streaming) -- needs CUDA + local weights
python -m wick.phase4b
# Device profiler: which VLMs fit your hardware (no model download)
python -m wick.profiler
# Full assertion suite
python -m pytest tests/ -vstreamlit run scripts/monitor.pyOpens the Wick Monitor browser tab: pick one or more run logs from
logs/ in the sidebar and see summary cards (peak VRAM, steps, final
loss, overflows, pass/fail vs the tripwire), loss-over-step charts
(select streamed + resident to overlay both), and VRAM-over-step with the
tripwire line. If a log is still being written, the page auto-refreshes
every few seconds, so it works as a live viewer during training too.
Note: this is for viewing logs — it works alongside or after the
terminal dashboard (--live-dashboard), it doesn't replace it.
Apache-2.0 · built for the low-VRAM community.