Summary
glm53-spark-tp2 (GLM-5.3-Flash-NVFP4-Spark, TP2/DCP2, MTP3) works well overall on 2x RTX PRO 6000, but occasionally degenerates into a repetition loop — usually inside the reasoning block, once in the final answer text. Nothing in the engine logs coincides with it (no preemption, OOM, allocator retries or exceptions), so this looks like model/checkpoint behaviour rather than a runtime fault. Posting in case it's useful for qualifying the Spark checkpoint.
Setup
- Image:
ghcr.io/local-inference-lab/vllm:karmic-kraken-beta-20260922-3221ccacf71002ea
PRESET=glm53-spark-tp2, CACHE_MODE=vram, MAX_NUM_SEQS=4 (preset value, pinned); no MAX_MODEL_LEN override (auto-fit resolved 995,328)
- Resolved argv matches the published recipe: TP2 + DCP2, MTP
num_speculative_tokens: 3 (probabilistic draft / standard rejection), fp8 KV, kv-cache-memory-bytes 4190109696, B12X backends
- Sampling: preset defaults only —
temperature 1.0, top_p 0.95 (same as the checkpoint's generation_config.json); client sends no sampling overrides, no repetition/presence penalty. reasoning_effort: high, clear_thinking: false
- Hardware: 2x RTX PRO 6000 Blackwell 96 GB, PCIe, driver 610.43.03 (recipe notes 615.71.09 tested — loads and serves fine on 610, flagging in case it matters)
- Client: agentic coding harness via an OpenAI-compatible proxy, multi-turn with tool calls
Observed
- The loop occurred on a single request at ~140k prompt tokens (well inside the context cap), one request running, KV usage ~30%.
- The looping turn produced 24,805 completion tokens (~54k chars of reasoning) before stopping.
- Reasoning loop example — the model re-quotes the same line over and over with small variations:
But then re-read result 2.4.1: "profile hard limits: true", then profile injection sets the effective env.
Wait — review result 2.4.1: "profile hard limits: true", then profile injection sets the effective env.
But then re-read result 2.4.1: "profile hard limits: true", then profile injection sets the effective env.
... (dozens more)
- A separate occurrence in the answer text degenerated into a short token-level cycle,
, or in [ ] -/ repeated for many hundreds of tokens.
- MTP signature: during the loop, SpecDecoding metrics showed sustained 97–99.9% avg draft acceptance (mean acceptance length ~3.9–4.0/3+1) at ~290 t/s generation, versus 50–90% during normal reasoning. Sustained ~99% acceptance may be a cheap loop detector.
Possibly relevant context
- The conversation was started on a different model and switched to Spark mid-session (history compacted twice). With
clear_thinking: false, earlier reasoning stays in the prompt — not established as causal, but noting it.
Questions
- Has Spark been through the same verifier-backed behavioral-fidelity run as the NVFP4 / QAD checkpoints? Those report 46–73 length-limited responses out of 21,504 at temp 1 / top_p 0.95 — is Spark's length-limit (looping) rate known?
- Is there a recommended mitigation that stays within the qualified sampling contract (e.g. a small presence/repetition penalty), or would that invalidate the qualification?
- Any known interaction between DCP2 / MTP3 on this preset and repetition behaviour?

Summary
glm53-spark-tp2(GLM-5.3-Flash-NVFP4-Spark, TP2/DCP2, MTP3) works well overall on 2x RTX PRO 6000, but occasionally degenerates into a repetition loop — usually inside the reasoning block, once in the final answer text. Nothing in the engine logs coincides with it (no preemption, OOM, allocator retries or exceptions), so this looks like model/checkpoint behaviour rather than a runtime fault. Posting in case it's useful for qualifying the Spark checkpoint.Setup
ghcr.io/local-inference-lab/vllm:karmic-kraken-beta-20260922-3221ccacf71002eaPRESET=glm53-spark-tp2,CACHE_MODE=vram,MAX_NUM_SEQS=4(preset value, pinned); noMAX_MODEL_LENoverride (auto-fit resolved 995,328)num_speculative_tokens: 3(probabilistic draft / standard rejection), fp8 KV,kv-cache-memory-bytes 4190109696, B12X backendstemperature 1.0,top_p 0.95(same as the checkpoint'sgeneration_config.json); client sends no sampling overrides, no repetition/presence penalty.reasoning_effort: high,clear_thinking: falseObserved
, or in [ ] -/repeated for many hundreds of tokens.Possibly relevant context
clear_thinking: false, earlier reasoning stays in the prompt — not established as causal, but noting it.Questions