Skip to content

[MLX][Spec] Add Gemma 4 MTP E2E and docs - #6

Draft
wirybeaver wants to merge 1 commit into
prototype/mlx-gemma4-mtp-enablefrom
prototype/mlx-gemma4-mtp-e2e-docs
Draft

wirybeaver wants to merge 1 commit into
prototype/mlx-gemma4-mtp-enablefrom
prototype/mlx-gemma4-mtp-e2e-docs

Conversation

@wirybeaver

@wirybeaver wirybeaver commented Jul 23, 2026

Copy link
Copy Markdown
Owner

Stack

PR 6 of 6. Depends on #5.

  • Base: prototype/mlx-gemma4-mtp-enable
  • Head: prototype/mlx-gemma4-mtp-e2e-docs

Scope

Completes the experimental prototype with:

  • minimal JSON-safe /server_info activation and lifecycle counters;
  • one pinned real-model exact-token parity gate;
  • documented launch flags, limitations, and verification steps.

The test surface deliberately omits the umbrella prototype’s endpoint/horizon matrix, memory soak, reload fingerprints, and unrelated OpenAI response API changes.

Validation

Passed on Apple Silicon:

  • targeted pre-commit hooks;
  • accumulated offline suite: 54 passed, 65 subtests passed;
  • pinned Stage B E2E: 1 passed.

The E2E launches fresh target-only and MTP servers, compares 64 exact cross-window token IDs, validates per-request counter deltas and cleanup, flushes without changing the assistant generation, and verifies an 8-token post-flush prefix.


CI States

Latest PR Test (Base): ❌ Run #30053131640
Latest PR Test (Extra): ❌ Run #30053131384

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

apple-silicon documentation Improvements or additions to documentation speculative-decoding

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant