test: cover load-path defect shapes with synthetic checkpoints - #143
Merged
Conversation
Points at bfc2462, which brings two things: - SharpAI/mlx-swift-lm#48 — glm_moe_dsa / deepseek_v3_2 load and run with dense attention (stage 1 of #111). GLM-5.2 is DeepSeek V3.2, whose indexer is inert below index_topk (2048), so output is exact for the first 2048 positions of context and diverges beyond them. That is enough to exercise --stream-experts against the 308GB checkpoint, which is what the issue actually asks for. - SharpAI/mlx-swift-lm#47 — the all-KV-shared assistant regression tests, which had not been picked up by a bump yet. #48 also generalises a latent trap in DeepseekV3.sanitize, which dropped `model.layers.61` by string literal. That number is just numHiddenLayers; on GLM-5.2's 78 layers it would have deleted a real layer while keeping the MTP block. Verified past the registry: pointing the binary at a glm_moe_dsa config constructs the model and fails only on absent weights — Key model.embed_tokens.weight not found in DeepseekV32Model.DeepseekV3ModelInner.Embedding so the architecture is reachable end to end, not merely registered. No real weights have been run: the smallest glm_moe_dsa checkpoint is 308GB. Refs #111 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Dependency Automation has failed all 12 times it has run since 2026-04-27 —
it has never once succeeded. Every failure is the same:
##[error]Input 'token' not supplied. Unable to continue.
The Create Pull Request step reads secrets.SWIFTLM_PR_TOKEN, which is not set
in this repository. The dispatch side is fine: mlx-swift-lm's auto_release
does hold a token that can dispatch cross-repo, so the event arrives and the
job runs, does its work, and dies at the last step.
Rather than add the secret, stop trying to open the PR. A workflow needs a
personal access token to open one usefully because GitHub does not start
workflow runs for events raised by GITHUB_TOKEN — a bot-opened PR would arrive
with no checks at all, permanently pending rather than green, and release.yml
gates releases on CI concluding successfully. A pushed branch plus a compare
link in the job summary costs one click and gets real CI, because the PR event
is then the human's.
Keeping a human in that loop is not a consolation prize. Bumps here have
needed a pointer check, an umbrella build and a smoke test before they were
trustworthy; this does the mechanical part and leaves the judgement.
Three further problems fixed while in here:
- The mlx-swift branch ran `swift package update mlx-swift`, which does
nothing: both dependencies are `.package(path: "./…")` local paths backed by
submodules, and SwiftPM takes whatever is on disk for a path dependency. It
could only ever have produced an empty commit. Both are now handled the same
way, as the pointer move they are.
- client_payload was interpolated straight into run blocks, so a crafted
new_tag would have been executed rather than compared. Values are now
validated (source_repo against an allowlist, new_tag against a plain-tag
pattern) and passed through the environment. Verified rejecting
`b554; rm -rf /`, `$(whoami)`, `b554 && curl evil.sh`, `../../../etc/passwd`,
`-x` and empty, while accepting b554, b459 and v1.2.3.
- A re-dispatch for a tag already checked out produced an empty commit; that
case now reports and stops.
Exercised against the real submodule: an already-current tag (b500) takes the
no-op path, a nonexistent tag (b99999) fails with a clear message, and a real
older tag (b497) computes bfc2462 → b320bc4.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Main went red at 19:28 today with nothing changed on our side — the merge that
preceded it touched only a workflow file. The failing job was
integration_matrix (vision), and it reproduced on re-run, so it was not a flake.
LiquidAI republished LFM2.5-VL-450M-MLX-4bit at 19:23, five minutes earlier.
The new revision's chat template is one brace short of valid:
old: {{- bos_token -}}
new: {- bos_token -}}
Every request against it returns HTTP 500,
`parser('Unexpected token type: closeExpression')`. Confirmed by reproducing
locally against the new revision, then restoring that single brace in a copy —
same weights, same request, HTTP 200 with identical token counts. The fault is
upstream, not a compatibility gap on our side, and no code change here would be
the right response to a malformed template.
CI never noticed the substitution because the vision job did not prefetch this
model at all: the server fetched it mid-test and resolved the floating id to
whatever was newest. So the job's result depended on what a third party
published that afternoon.
Pins the revision, prefetches it, and teaches ci-download-models.sh a
`repo@revision` spec so any model can be pinned the same way. The test resolves
the pinned snapshot on disk and falls back to the floating id with a printed
note, so a local run without a prefetch still works but cannot quietly test a
different revision than CI did.
The test-vision.sh edit rotates the job's model cache key, so CI re-downloads
rather than restoring a cache that now holds the broken revision.
Verified: the vision test passes locally with the pin, both cases; the
`repo@revision` split parses correctly for pinned and unpinned specs; the
fallback path triggers and warns when the pinned snapshot is absent.
Worth reporting upstream — LiquidAI's template is broken for every consumer,
not just this repository.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The first version of the revision-pinning change assembled the optional `--revision` flag into an array and expanded it unconditionally. The runners are macOS, which ships bash 3.2, where expanding an *empty* array under `set -u` is an unbound-variable error rather than expanding to nothing. Every unpinned download therefore failed, which took out every job that prefetches a model — speculative-decoding, dflash, ssd-draft-memory-guard — while the pinned path would have worked fine. Spelled the two calls out instead. Verified by running the script under /bin/bash 3.2 with `set -u` for both shapes: unpinned resolves to the current snapshot, `repo@revision` resolves to the pinned one. CI caught this, which is the system working; worth noting the local `bash -n` syntax check could not have, since the failure is a runtime expansion. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
When a checkpoint ships a chat template the Jinja parser rejects, SwiftLM used
to load cleanly, report ready, open the port, and then return HTTP 500 on every
request with
parser('Unexpected token type: closeExpression')
That message names neither the chat template, nor the model, nor the fact that
the offending file came out of someone else's checkpoint. It reads as a broken
server. Diagnosing the real instance of this — LiquidAI republishing
LFM2.5-VL-450M-MLX-4bit with `{- bos_token -}}`, one brace short — took CI logs
and a bisect across two model revisions, and that was with far more to work
with than a user reporting it would have.
Two changes:
- Template failures now surface as MalformedChatTemplate, which names the model,
points at chat_template.jinja / tokenizer_config.json, says the defect belongs
to whoever publishes the checkpoint, and mentions pinning as the workaround.
- The template is rendered once during load, before the port opens. A checkpoint
that cannot produce a prompt now refuses to start rather than serving 500s
indefinitely across restarts.
A model with no chat template at all stays legitimate — base models ship without
one and /v1/completions does not need it — so only a template that exists and
fails to parse is treated as fatal. The startup probe is shaped like the
simplest real request (one user turn, add_generation_prompt) rather than a bare
minimum, so a failure is the template's rather than the probe's.
Verified: the broken revision now exits 1 with the diagnostic and never opens the
port. Five cached models covering both modalities, thinking and non-thinking, and
two model families all still start normally — LFM2.5-VL-450M (good revision),
Qwen2-VL-2B, Qwen2.5-0.5B, Qwen3-1.7B, LFM2-VL-1.6B. Contract suite: 10 passed,
0 failed, 2 skipped.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
#128 observed that of the nine defects found in the #108/#110/#112 cycle, real checkpoints caught four, code review caught four, and the ~250-test unit suite caught none — because the bugs lived in weight-and-config-shape assumptions that only a checkpoint on disk exercises. The obvious response, running CI against real models, does not fit: gemma-4-e2b alone is 3.6 GB and GitHub allows 10 GB of cache for the entire repository. llama.cpp solved the same problem by publishing purpose-built tiny models — ggml-org/test-model-stories260K is 1.2 MB — rather than shrinking real ones. Their files are GGUF and unusable here, but the technique transfers: a checkpoint with the same config fields, the same weight keys, random values and a ~300-token vocabulary runs the same loading code at a few hundred kilobytes. Four shapes, each one a defect that reached users: dense baseline stray-shard #118 — a .safetensors beside the index but absent from it kv-shared-absent #120 — gemma-4-e4b shape, shared layers ship no k/v kv-shared-present b674 — gemma-4-e2b shape, shared layers ship k/v anyway 1.2 MB committed in total; the suite runs in 9 seconds with no network, no model cache and no download. The CI entry declares no models at all. Red-green verified against the real history rather than asserted. Building the submodule at 717d77f — #44 landed, #45 not yet, which is the state that shipped the b674 regression — kv-shared-present fails with Unable to set model.layers.2.self_attn.v_proj while kv-shared-absent still loads, exactly reproducing the asymmetry that made that regression possible. Both load at current main. Two things these fixtures do not do. They say nothing about numerical correctness, because the weights are noise — real checkpoints remain the only way to judge output quality. And the MoE-config shape behind #112 is not covered yet; it needs a MoE architecture fixture and is worth a follow-up. Shapes were mirrored from a real gemma-4-e2b checkpoint rather than guessed, after the loader rejected several hand-written attempts. Notes for whoever extends this: swift-transformers rejects a WordLevel tokenizer with "BPETokenizer requires merges", and merges must be spelled in the byte-level alphabet. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This was referenced Aug 14, 2026
solderzzc
added a commit
that referenced
this pull request
Aug 14, 2026
… faults (#145) Review follow-ups on #143/#144. All three findings were in code I wrote. **The MoE fixture tested nothing it claimed.** It used `qwen3_moe`, whose Swift configuration decodes from the root of config.json, so the expert count had to be at the top level; the nested copy underneath was never reached, because findExpertCounts is breadth-first and returns on the first hit. Both advertised properties — the nesting, and the `vision_config` decoy that "must not be mistaken" — were dead weight. Deleting the nested walk entirely would not have failed it. Rebuilt on `gemma4`, whose Gemma4Configuration decodes `text_config` and nothing else, so the count exists only one level down and the decoy is genuinely reached. `gemma4` also contains no "moe", so modelTypeImpliesMoE cannot rescue it — which is what made the qwen3_moe version untestable. Red-green verified, the check the previous version could not pass: reverting detection to the pre-#114 top-level single-key form fails the new assertion, and restoring it passes. The fixture now reproduces #112. The assertion also moved to the right gate. There are two: a config-level MoE check, and a model-level StreamableMoE conformance check. gemma4 passes the first and legitimately declines the second, so asserting "streaming enabled" would have tested the wrong thing. It now asserts only that detection did not reject. Incidentally covers the fused-expert remap — real gemma4 checkpoints ship `experts.gate_up_proj` as one tensor that sanitize splits into `switch_glu.gate_proj`/`up_proj`. A wrong split axis is a silent numerical fault. **Two harness faults.** cleanup() killed the server without waiting, and the readiness loop probed health before checking liveness. If a teardown outlived the 1s sleep, the next fixture would fail to bind, and its first probe would be answered by the previous server — assertions then run against the wrong checkpoint and report a false pass, not a flake. cleanup now waits, and liveness is checked first. The generator wrote into existing directories without clearing them, so files a builder stopped emitting survived and the fixture kept testing a shape the source no longer described. Each fixture's own directory is now cleared first — scoped to one known directory, which is also what keeps regeneration away from siblings like tests/fixtures/omni, whose assets belong to test-omni.sh. Suite: 6 passed, 0 failed, twice consecutively. Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Addresses the Tier 1 half of #128, using the technique llama.cpp uses rather than the one #128 originally proposed.
The problem #128 identified
Of the nine defects in the #108/#110/#112 cycle: real checkpoints caught four, code review caught four, the ~250-test unit suite caught none. The bugs lived in weight-and-config-shape assumptions that only a checkpoint on disk exercises.
The obvious response — run CI against real models — doesn't fit the budget.
gemma-4-e2balone is 3.6 GB against a 10 GB per-repository GitHub cache limit, which is why the MTP CI job was ruled out earlier in this cycle.What llama.cpp does instead
They don't shrink real models; they publish purpose-built tiny ones:
ggml-org/test-model-stories260Kggml-org/tinygemma3-GGUFggml-org/stories15M_MOETheir files are GGUF, so unusable directly. The technique transfers exactly: same config fields, same weight keys, random values, ~300-token vocabulary instead of 150k.
What this adds
Four shapes, each one a defect that reached users:
densestray-shard.safetensorsbeside the index but absent from itkv-shared-absentkv-shared-present1.2 MB total, 9 seconds, no network, no model cache, no download. The CI matrix entry declares no models at all.
Red-green verified against real history
Not asserted — reproduced. Built the submodule at
717d77f(#44 landed, #45 not yet — the state that shipped the b674 regression):717d77fbfc2462kv-shared-absentkv-shared-presentUnable to set model.layers.2.self_attn.v_projThat asymmetry is the bug: #44 assumed shared layers never ship k/v weights, which holds for e4b and not for e2b. A 394 KB file catches it.
What this does not do
No numerical coverage. The weights are noise, so these prove shapes load and plumbing runs — nothing about whether the arithmetic is right. Real checkpoints remain the only way to judge output quality; this replaces the shape half of Tier 1, not the whole of it.
#112's MoE-config shape is not covered. It needs a MoE architecture fixture — worth a follow-up.
Notes for extending
Shapes were mirrored from a real gemma-4-e2b checkpoint after the loader rejected several hand-written attempts. Two things that cost time: swift-transformers rejects a WordLevel tokenizer (
BPETokenizer requires merges), and merges must be spelled in the byte-level alphabet (Ġ, not a raw0x20).scripts/make-test-fixtures.pyregenerates everything deterministically.🤖 Generated with Claude Code