Skip to content

adapters: merge LoRA and LoKr adapters into either backbone half at load - #10

Open
timoncool wants to merge 14 commits into
ServeurpersoCom:masterfrom
timoncool:adapters
Open

timoncool wants to merge 14 commits into
ServeurpersoCom:masterfrom
timoncool:adapters

Conversation

@timoncool

Copy link
Copy Markdown

Community adapters for YuE2 load at request time. The server takes --adapters <dir> and a request names entries of it:

"adapters": [{ "name": "instrumental.safetensors", "ar_scale": 1.0, "nar_scale": 0.0 }]
  • The factors are merged into the staged weights of the half they touch, before wctx_alloc and before QKV / gate-up fusion, every contribution to a tensor summed in one backend graph and encoded back to the GGUF type once. An adapter list is part of the model store key, so a half only reloads when its own list changes, and under --keep-loaded another list replaces the idle variant instead of stacking a copy.
  • Read: PEFT / trainer-native keys, ComfyUI and AI Toolkit keys (text_encoders. for AR, diffusion_model. for NAR, fused qkv_proj / gate_up_proj split back by rows, .diff), GGUF-style blk.N.attn_q keys and LyCORIS LoKr (lokr_w1 with lokr_w2 or lokr_w2_a/lokr_w2_b). Alpha comes from the per-tensor .alpha, then __metadata__, then adapter_config.json; use_rslora scales by alpha / sqrt(rank).
  • DoRA, LoHa and PiSSA-delta files are refused by name: their update is not B @ A, and merging one as plain LoRA renders plausible, wrong audio.
  • GET /props lists the adapter directory with the halves each entry touches; yue-synth takes --adapters.

Tested with the Mothersuperior instrumental adapter (vocal stem at -89.8 dB against -25.3 dB without it), ntc-ai sliders and adapters trained with HOT-Step's ace-train yue2-joint-train.

Requests name entries of a --adapters directory, each with a scale per
half. The factors merge into the staged weights before QKV and gate/up
fusion, every contribution to a tensor summed in one backend graph and
encoded back to the GGUF type once. The model store keys each half on its
own adapter list, so an AR only adapter never reloads the NAR half, and
under --keep-loaded a new list replaces the idle variant.

Reads PEFT and trainer native keys, ComfyUI and AI Toolkit keys with fused
qkv_proj and gate_up_proj split back by rows and .diff weights, GGUF style
blk.N keys, whole replacement flow heads and LoKr factors. GET /props
lists the directory with the halves each entry touches.
adapter_config.json's use_rslora scales a LoRA by alpha / sqrt(rank), which
a merge at alpha / rank would apply 16 times too weakly at rank 256. A DoRA
magnitude vector, LoHa factors or a HOT-Step method marker other than plain
LoRA mean the update is not B @ A; such a file is refused by name instead of
merged into plausible, wrong audio.
…ile bounds, strict request entries

K-quant bases are decoded on the host when the backend has no cast to F32
(the CUDA copy kernels do not decode Q4_K/Q5_K/Q6_K), and the staging buffer
holds only the tensor's own bytes. LyCORIS w1_a/w1_b are multiplied out like
w2_a/w2_b; a w2 without its w1, an adapter that changes nothing in its half,
rank_pattern and alpha_pattern are refused instead of merging nothing or the
wrong scale. Keys naming nothing in the model are logged. A lone file no longer
takes the adapter_config.json of the folder it sits in. Safetensors tensors
must fit the file and their shape. A request's adapters entry may be a bare
name; any other malformed entry fails the request. The store key carries each
file's size and time, and a failed merge frees what the load had allocated.
@timoncool

Copy link
Copy Markdown
Author

Pushed 177d9b9 and 16c42d2 on top: K-quant bases decoded on the host when the backend cannot cast them to F32 (a LoRA on the Q6_K set stopped on CUDA), staging sized to the tensor, LyCORIS w1_a/w1_b, refusals instead of silent no-ops (orphan w2, an adapter that changes nothing, rank_pattern/alpha_pattern), unknown keys logged, safetensors bounds checked, strict adapters entries in requests, store key with file size and time. Verified on the Q6_K set with a LoRA (112 tensors) and a LoKr (16).

A 130 s song under guidance allocated 5.4 GB of cache for 24576 positions;
it now takes the prefix, the budget and the acoustic chunk, padded to the
attention window, so the output is unchanged. The batch graph keys on an
allocation number: a freed cache can come back at the same address.
A release ships a CUDA 12 and a CUDA 13 ggml-cuda, each in a folder of
its own, and the launcher names the one the card and its driver run.
Loaded before the rest so CUDA still outranks Vulkan; a second CUDA left
beside the executable by an older install is unloaded.
Rows gathered from an F16 table and a matrix product run on the device,
checked against their known sum. A device that initialises but fails its
first kernel - a CUDA driver older than the toolkit, a broken Vulkan
driver - stops the engine at startup, so the launcher moves on to the
next device instead of failing the first song.
…y they sound in, as SheetSage2 (55bfe14) and ComfyUI (66b68a3) write them
@ServeurpersoCom

Copy link
Copy Markdown
Owner

Hi!
I need to look into this PR to ensure consistency with what I did in acestep.cpp (there is LoRA / LoKr support also)

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants