feat: sync llama.cpp to b11010 - #395
Merged
Merged
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
🤖 Automated llama.cpp sync
This PR was automatically created/updated by the daily sync workflow.
Changes:
b10951tob11010Verification:
auto/sync-llama.cppdirectly: while this PR stays open, later sync runs build on top of the branch instead of resetting it.📋 llama.cpp changes (b10951 → b11010)
4bc272fvulkan: work around NV bug with argsort_large.comp (#28975)fb27a52TP: fix split state and granularity for fused QKV gemma4, qwen35 (#28965)c6824a9ci: switch fast jobs back to github (#28959)2f3fd02Enable CUDA graph for MTP draft (#28549)1ec8188hexagon: Support for K-Quants Q4_K and Q6_K (#28994)82324fchexagon: accept the zeroed rope probe in supports_op (#28995)7ceed87models : allow Nemotron-H models to only define layer_norm_epsilon (#28989)7d6f5d0model : add support for HrmTextForCausalLM (DFM Mimir 1B) (#27625)83078feCUDA/HIP: improve access patterns in im2col (#28013)f266648spacemit : fix wrong transpose function for int16 data (#25161)6019933rpc : invalidate cached compute graph when a referenced buffer is freed (#24292)b04d4e5Change max context length for auto-fitting with unified KV (#28849)37b53fdqwen4exp: add hc ops (#28901)fccf716HIP: broaden MoE ncols_opt tile heuristic on RDNA3.5 architecture (#28935)0bec16echat : force\n</think>on reasoning budget end for qwen3-coder (#28869)d4365d9vulkan: make MUL_MAT_ID BN/2 tail unconditional (#28923)0a8b29ametal: fix NaN in mul_mm_id when activations exceed f16 range (#26223)583926eci : add self-hosted webgpu to hf-jobs (#28712)e13469allama-bench: support --version to print build info (#28971)930e2fahexagon: add back missing contiguous fast-path and hvx_copy_uu for each run (#28886)72b590dhex-cpy: use dma if src and dst are contiguous (#28906)38a5b42HIP: Enable AllReduce for ROCm (#27825)9f31776opencl: choose the MoE expert matmul by batch size for speculative decoding/MTP (#27637)d1d3c33ci: build MUSA for only 1 arch (#28944)6011c34docs: Rule of thumb for AI review time [no ci] (#28945)7609846rpc : hash-cache only weights (#28789)5431581cuda: support row-contiguous SUM_ROWS (#26308)9e71716models : move build_arch_graph() after graph() template specialization (#28934)fc82583vulkan: support sparse Flash Attention (#28105)77d554bOpenVINO: optimize stateful decode and GPU MoE inference (#28638)6ec1a7eopencl: add generic ssm_scan (#28881)1af6c65ci: bump kleidiai runners from 22.04 to 24.04 (#28885)1e7bcf3metal : add FA kernels for HSK=96, HSV=64 (MiniCPM3) (#28599)0ecb159ci: Bump CUDA Windows x64 builds to 13.4.1 (#28930)987498fci : fix android release (#28936)4c9233ccuda : enable i16 and i32 for DUP (#28897)69eb250cmake : use PROJECT_SOURCE_DIR instead of CMAKE_SOURCE_DIR (#28771)1bc7a5awebui: stop re-probing disabled /tools endpoint on every message (#28646)7cf1c54ci : reuse build tag name when used instead of safe one (#28911)96ffdc4CI: hip-quality-check: ignore spill added in bfdc32183d57f1e35bacf35c47d6311e2028bbbc (#28909)bfdc321HIP: fattn-mma: use fp32 accumulation on MFMA devices (#28576)391fac1ci : add ubuntu-cuda builds to release (#28186)41abbfdqwen4exp: enable rms_norm + mul fusion (#28896)b4fa47drelease : added gfx1103 to ubuntu rocm build (#28423)f3a184bcmake : remove precompiled headers (#28892)dfe4516scripts: Add script to verify API/ABI compatibility (#28579)b29c606llama.cpp : bump version to 0.4.1 (#28900)d9e03f1sync : ggmleeea731ggml : bump version to 0.24.0 (ggml/1627)bbdd9f2tests : add fusion baseline README and broaden fusion CI triggers (#28893)Please review and merge if all checks pass.