test: cover the third config-overlay pattern in a VLM architecture (GlmOcr) - #53
Merged
Conversation
…lmOcr) Follow-up to ml-explore#147-ml-explore#149. MLXVLM/Models/GlmOcr.swift (938 lines) was the largest remaining zero-coverage file in the model layer. A third distinct config pattern: BaseConfiguration (model_type, image_token_id, vocab_size, ...) decodes from the same top-level container as text_config/vision_config rather than nesting inside either. GlmOcrConfiguration.init decodes TextConfiguration and VisionConfiguration from their keyed children, then re-runs BaseConfiguration(from: decoder) against the same decoder to pick up the flat fields. That split has a real trap the tests name rather than work around: GlmOcr's reported vocabularySize reads config.baseConfiguration.vocabularySize (a top-level field defaulting to 59392), while the language model's actual output width is text_config.vocab_size. A checkpoint that sets one and not the other produces a model whose declared vocabulary size does not match its logits' last dimension. The tiny config sets both to the same value on purpose, which establishes the ordinary path — the mismatch itself is not asserted, just documented for whoever touches this next. Verified with coverage: GlmOcr.swift moves from 0% to 37.9% line coverage. Full suite: 128 tests, 23 suites, passing (120 before, +8 across this PR and its Qwen25VL sibling). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
solderzzc
added a commit
to SharpAI/SwiftLM
that referenced
this pull request
Aug 16, 2026
#150) Points at 512dcde, which carries both: - SharpAI/mlx-swift-lm#53 — MLXVLM/Models/GlmOcr.swift, whose base config overlays the same top-level container as text_config/vision_config - SharpAI/mlx-swift-lm#54 — MLXVLM/Models/Qwen25VL.swift, whose text config does the same, with several fields duplicated across two parallel structs Both were, alongside Gemma4/Qwen3VL/Qwen35/LFM2VL (#147-#149), among the largest zero-coverage files in the model layer. Coverage: GlmOcr 0% → 37.9%, Qwen25VL 0% → 26.7%. No SwiftLM-side code changes; verified the umbrella builds clean against the bumped pointer. Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Companion PR opening next for Qwen25VL. Follow-up to ml-explore#147-ml-explore#149 —
MLXVLM/Models/GlmOcr.swift(938 lines) was the largest remaining zero-coverage file in the model layer.What's different about it
A third distinct config pattern:
BaseConfiguration(`model_type`, `image_token_id`, `vocab_size`, ...) decodes from the same top-level container as `text_config`/`vision_config`, rather than nesting inside either. `GlmOcrConfiguration.init` decodes `TextConfiguration`/`VisionConfiguration` from their keyed children, then re-runs `BaseConfiguration(from: decoder)` against the same decoder to pick up the flat fields.A real trap, named rather than worked around
`GlmOcr`'s reported `vocabularySize` reads `config.baseConfiguration.vocabularySize` (a top-level field, defaulting to 59392), while the language model's actual output width is `text_config.vocab_size`. A checkpoint that sets one and not the other produces a model whose declared vocabulary size doesn't match its logits' last dimension. The tiny config sets both to the same value on purpose — this establishes the ordinary path, and the type doc records the mismatch for whoever touches this next.
Verification
```
GlmOcr.swift: 0% → 37.9% line coverage
Full suite: 128 tests, 23 suites, passing (120 before, +8 across this PR and its Qwen25VL sibling)
```
🤖 Generated with Claude Code