Not an issue per se since I don't have the full bugfix yet, but I wanted to say it's a really nice effort - after patching a few bugs I plugged redline into llama.cpp's ROCm backend and immediately managed to halve the decoding performance gap between ROCm and Vulkan (from 62 t/s to 66 t/s with Vulkan at 70 t/s for my tested model). I'll report once I'm done with my little experiment and then I'll also submit the bugfix PRs.
Not an issue per se since I don't have the full bugfix yet, but I wanted to say it's a really nice effort - after patching a few bugs I plugged redline into llama.cpp's ROCm backend and immediately managed to halve the decoding performance gap between ROCm and Vulkan (from 62 t/s to 66 t/s with Vulkan at 70 t/s for my tested model). I'll report once I'm done with my little experiment and then I'll also submit the bugfix PRs.