forked from ggml-org/llama.cpp
-
Notifications
You must be signed in to change notification settings - Fork 0
All issues
Issue creation is restricted in this repository
Issues
is:issue state:open
is:issue state:open
Search results
HYDRA E1: armed env + --verbose deterministically hangs llama-perplexity (lazy call_once arm on compute path)
review-findingFinding created from code reviewFinding created from code reviewStatus: Open.#155 In ddvnguyen/llama.cpp;Certified-base secondary KV/checkpoint-snapshot cache is 3.0x smaller than pilot base (12.76 vs 38.26 MiB) — invisible to cold single-turn legs; continuation/cache-reuse leg required before epic merge
review-findingFinding created from code reviewFinding created from code reviewStatus: Open.#153 In ddvnguyen/llama.cpp;Campaign doc drift: CLAUDE.md names a model that does not exist on the rig (nearly voided a replicate)
review-findingFinding created from code reviewFinding created from code reviewStatus: Open.#150 In ddvnguyen/llama.cpp;--n-cpu-moe silently discards user -ot overrides for expert tensors (all-host placement, no warning)
review-findingFinding created from code reviewFinding created from code reviewStatus: Open.#149 In ddvnguyen/llama.cpp;- Status: Open.#148 In ddvnguyen/llama.cpp;
- Status: Open.#147 In ddvnguyen/llama.cpp;
R1: expert slab is raw cudaMalloc outside ggml accounting, never freed (~3.2 GB at N=38)
review-findingFinding created from code reviewFinding created from code reviewStatus: Open.#145 In ddvnguyen/llama.cpp;M2 (E0): LOOKUP-basis per-path hit counters, decode-only labelled (epic base)
review-findingFinding created from code reviewFinding created from code reviewStatus: Open.#144 In ddvnguyen/llama.cpp;M1 (E0): per-invocation timing on the ENGAGED path as a hook parameter — p50/p99 µs per site (epic base)
review-findingFinding created from code reviewFinding created from code reviewStatus: Open.#143 In ddvnguyen/llama.cpp;M3 (E0): pin-load must fail loudly on every failure + load summary + engagement counter (epic base)
review-findingFinding created from code reviewFinding created from code reviewStatus: Open.#142 In ddvnguyen/llama.cpp;hydra_gather_mmvq: miss experts re-fetched from host every token (~498 MiB/token, no cross-token reuse)
review-findingFinding created from code reviewFinding created from code reviewStatus: Open.#141 In ddvnguyen/llama.cpp;hydra_gather_mmvq: full stream sync per call suppresses CUDA graph capture (~144 host syncs/token)
review-findingFinding created from code reviewFinding created from code reviewStatus: Open.#140 In ddvnguyen/llama.cpp;