Skip to content

glm53-spark-tp2: occasional repetition loops (reasoning and answer) at default sampling #107

Description

@mmssix

Summary

glm53-spark-tp2 (GLM-5.3-Flash-NVFP4-Spark, TP2/DCP2, MTP3) works well overall on 2x RTX PRO 6000, but occasionally degenerates into a repetition loop — usually inside the reasoning block, once in the final answer text. Nothing in the engine logs coincides with it (no preemption, OOM, allocator retries or exceptions), so this looks like model/checkpoint behaviour rather than a runtime fault. Posting in case it's useful for qualifying the Spark checkpoint.

Setup

  • Image: ghcr.io/local-inference-lab/vllm:karmic-kraken-beta-20260922-3221ccacf71002ea
  • PRESET=glm53-spark-tp2, CACHE_MODE=vram, MAX_NUM_SEQS=4 (preset value, pinned); no MAX_MODEL_LEN override (auto-fit resolved 995,328)
  • Resolved argv matches the published recipe: TP2 + DCP2, MTP num_speculative_tokens: 3 (probabilistic draft / standard rejection), fp8 KV, kv-cache-memory-bytes 4190109696, B12X backends
  • Sampling: preset defaults only — temperature 1.0, top_p 0.95 (same as the checkpoint's generation_config.json); client sends no sampling overrides, no repetition/presence penalty. reasoning_effort: high, clear_thinking: false
  • Hardware: 2x RTX PRO 6000 Blackwell 96 GB, PCIe, driver 610.43.03 (recipe notes 615.71.09 tested — loads and serves fine on 610, flagging in case it matters)
  • Client: agentic coding harness via an OpenAI-compatible proxy, multi-turn with tool calls

Observed

  • The loop occurred on a single request at ~140k prompt tokens (well inside the context cap), one request running, KV usage ~30%.
  • The looping turn produced 24,805 completion tokens (~54k chars of reasoning) before stopping.
  • Reasoning loop example — the model re-quotes the same line over and over with small variations:
    But then re-read result 2.4.1: "profile hard limits: true", then profile injection sets the effective env.
    Wait — review result 2.4.1: "profile hard limits: true", then profile injection sets the effective env.
    But then re-read result 2.4.1: "profile hard limits: true", then profile injection sets the effective env.
    ... (dozens more)
    
  • A separate occurrence in the answer text degenerated into a short token-level cycle, , or in [ ] -/ repeated for many hundreds of tokens.
  • MTP signature: during the loop, SpecDecoding metrics showed sustained 97–99.9% avg draft acceptance (mean acceptance length ~3.9–4.0/3+1) at ~290 t/s generation, versus 50–90% during normal reasoning. Sustained ~99% acceptance may be a cheap loop detector.

Possibly relevant context

  • The conversation was started on a different model and switched to Spark mid-session (history compacted twice). With clear_thinking: false, earlier reasoning stays in the prompt — not established as causal, but noting it.

Questions

  1. Has Spark been through the same verifier-backed behavioral-fidelity run as the NVFP4 / QAD checkpoints? Those report 46–73 length-limited responses out of 21,504 at temp 1 / top_p 0.95 — is Spark's length-limit (looping) rate known?
  2. Is there a recommended mitigation that stays within the qualified sampling contract (e.g. a small presence/repetition penalty), or would that invalidate the qualification?
  3. Any known interaction between DCP2 / MTP3 on this preset and repetition behaviour?
Image

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions