Skip to content

Complete cross-platform validation cases, device coverage and reports #514

Description

@leehack

Current evidence and remaining work — 2026-09-18

PR #515 head e2404b38b8d36384fed3183942deaa54dc3a8074. The full roadmap and platform matrix are not complete or green. The PR remains draft/unmerged. Native v0.4.1 and LiteRT 0.17.0-5 are the tested runtime pins. Runtime artifacts retain their source independently from reporter/implementation revisions.

Implemented and verified

  • Locked Gemma 4 E2B/Qwen3.5 GGUF and native LiteRT profiles; platform/backend/use-case inventory; quick/focused/release selection; explicit unsupported/unimplemented obligations.
  • Portable desktop/Flutter/Web harness with model/config/source hashes, download/checksum progress, transcripts, latency/TPS, JSON/CSV/JUnit/HTML and fail-closed placement/provenance qualification. CI builds all six desktop/app targets; no cloud/model run is implicitly launched.
  • Catalog 3 adds public stop-marker and unloaded-engine guards/recovery. GPU/NPU obligations include the extra model loads/generations. Thinking, tools and Unicode generation remain three explicit NOT_RUN release cases; Unicode tokenization is a separate implemented case.
  • Firebase polling/recovery fix (Preserve Firebase polling diagnostics and collect terminal iOS results #518), bounded read retries, no submission retries, fresh terminal collection and safe errors. Windows GCE upload/collection uses verified legacy SCP; desktop backend discovery resolves the standard bin/../lib layout (Windows compiled CLI bundles skip sibling lib backend modules #525). Vulkan reporting requires consistent indexed physical identity plus positive offload/compute buffers for every expected load; software rendering is never GPU coverage.
  • Local diagnostic STT/TTS/Moonshine and file voice-round-trip packs. Diagnostic passes remain qualified=false; mobile/Web packaging, full quality fixtures and physical microphone/playback remain open.
  • Private suite: 111 tests, provider: 44 tests (one existing x64 ABI skip on Mac), Windows service: 129 tests, root format/analyze pass. Exact final CI and independent audit evidence are recorded in PR Add portable validation suite and guarded Firebase/GCE runners #515.

Actual execution results

  • Original catalog-3 Mac arm64/Linux x64 bundles: source 89097b7. Gemma GGUF CPU/Metal/CUDA rows record17 PASS/1 stop FAIL/3 NOT_RUN; Qwen GGUF records16 PASS/2 FAIL/3 NOT_RUN. Mac Metal and Linux CUDA have native placement proof. Qwen history returned Cedar17, which retained the code but failed the strict exact-output oracle; no history-loss claim or predicate relaxation.
  • Corrected Windows x64 clean bundle f1a13f35: both primary GGUF models run on CPU/CUDA/Vulkan, with the same counts above. CUDA/Vulkan prove all five loads on NVIDIA L4. Gemma LiteRT CPU: 18 PASS/3 NOT_RUN. Qwen LiteRT CPU times out; both LiteRT GPU profiles remain unqualified with D3D12/DXC initialization errors. Every requested Windows profile produced retained results; this is not all-green platform qualification.
  • Linux hardware-only Vulkan also executes both primary GGUF models with five verified loads. Gemma LiteRT GPU passes 18 implemented cases with native L4/Vulkan selection, but automated placement remains unverified; Qwen LiteRT GPU times out. All teardown results are linked in the current PR report. The earlier loader-only llvmpipe attempt and driver-version-mismatch setup attempt stopped before inference and do not establish hardware GPU support.
  • Qwen3-ASR bytes fix Fix encoded audio transport in string chat templates #523 is integrated. Clean source 0c0a379 reran locked file and bytes inputs on CPU/Metal: WER 0, functional_pass=true, qualified=false. Earlier source 21f89f94 TTS CPU/Metal and all four file voice chains passed functional checks; Moonshine WER 0.22727 failed its strict oracle. Those older TTS/voice/Moonshine runs were not relabelled as a fresh rerun.
  • Firebase repaired infrastructure: physical iPhone 16 Pro / iOS 18.3 tiny GGUF CPU, 43 seconds, assertions PASS, qualified=true, collection COMPLETE, cleanup VERIFIED; matrix matrix-2j9mb3xdi1y7a. Original failed iOS matrix was re-collected as terminal INCONCLUSIVE/CANCELLED and unqualified. This tiny CPU control does not qualify primary-model Metal.
  • Firebase S24 / API 36 Gemma LiteRT GPU still has a 600-second large-model download failure before inference. Primary Android/iOS/NPU coverage remains open. Latest conservative allowance accounting: 20 measured rounded minutes plus 9 reserved for missing-duration cancellation; no additional Firebase run was launched.
  • The user explicitly authorized temporary GCE use with available credits and teardown. VMs use bounded deadlines/ownership-checked deletion. Final VM/disk/firewall absence is verified in the PR report; no idle resources are intentionally retained. Bootstrap execution is separate from qualification of the maintained GCE adapter/custom image.

Required next work

  1. Improve/stage large-model delivery for Firebase Android; rerun primary Android/iOS and device rotation only with a fresh quota/funding preflight. Preserve Firebase polling diagnostics and collect terminal iOS results #518 is fixed in PR Add portable validation suite and guarded Firebase/GCE runners #515 and live-tested; keep it open until merge.
  2. Complete thinking, tools and Unicode-generation cases and representative media packs. Resolve GGUF public generation emits caller stop markers before stopping #520 stop-marker emission and Qwen3.5 LiteRT GPU reload times out on macOS and leaves CLI alive #521 Qwen Mac LiteRT GPU reload hang; retain strict output-format results and investigate timeouts without asserting unsupported root causes.
  3. Establish reference-qualified Moonshine/speech quality fixtures, then finish portable mobile/Web speech adapters, report ingestion, microphone/playback and broader language/noise/voice boundaries. ASR bytes defect Native Qwen3-ASR byte input fails while the same WAV file transcribes successfully #517 is fixed by merged Fix encoded audio transport in string chat templates #523 and the rerun above.
  4. Establish immutable driver-ready GPU images and qualify the maintained GCE adapter end-to-end. Desktop portable execution now has real Windows/Linux CPU/CUDA evidence; Linux/Windows LiteRT GPU prerequisites/runtime failures remain distinct from packaging or GGUF GPU success.
  5. Acquire matched primary Gemma 4 Tensor G5 / SM8750 inputs and kits before NPU execution. S24 SM8650 is not interchangeable; Qwen3.5 NPU remains unestablished and typed speech NPU unsupported. Legacy Gemma 3 controls do not qualify primary families.
  6. Complete Web primary-model hardware tests, remaining device/CPU-architecture rotation (including macOS x64 and Linux/Windows arm64) and aggregate historical reporting.

Detailed exact profiles, configurations, outputs, TPS/TTFA, placement and resource-cleanup evidence are in PR #515. The acceptance requirements below remain open unless explicitly qualified above; older dated evidence is historical.

Required speech use cases: STT and TTS (2026-09-17)

STT and TTS are required follow-up acceptance scope, not optional/deferred coverage. Gemma 4 and Qwen3.5 remain primary chat/multimodal families; dedicated speech models are explicit exceptions. This section supersedes earlier wording deferring all specialty speech packs. The acceptance matrix below remains the full target. PR #515 now implements local diagnostic baselines for file/bytes STT, TTS, dedicated streaming ASR, and a file voice round trip; platform packaging, broader quality fixtures and qualification remain incomplete.

Models and runtime support

Select one small reference-qualified STT model and one TTS model first. Qwen3-ASR, Moonshine and Qwen3-TTS now have immutable locks and public-API diagnostic runners. Results and remaining failures are recorded below; they are not fully qualified references. If no candidate works through the public API, record a product/runtime gap and track the owning implementation; an external speech service or direct upstream-only run cannot count as llamadart coverage. Gemma audio understanding is a separate test and is not a substitute for dedicated transcription or synthesized audio output.

Acceptance cases

ID Case Required result
S01 Short clean, licensed or generated speech fixture through the public API Nonempty normalized transcript; compare against a committed reference with a model-qualified WER threshold, including a numeric/name fixture. Save raw and normalized transcripts.
S02 Silence, noise and malformed/unsupported audio Silence/noise behavior matches a reference-qualified expectation; invalid format fails with actionable error and the next valid request succeeds. No hangs or invented success.
S03 Input format and duration boundaries Supported sample rates/channels are accepted or explicitly converted by the app; unsupported inputs and over-limit audio are rejected clearly. Include short and bounded longer clips.
T01 Short sentence, punctuation, Unicode and numbers through public TTS API Nonempty decodable audio with valid declared codec, sample rate, channels and duration; reference-qualified intelligibility and pronunciation check. Do not use nonempty bytes alone as speech-quality proof.
T02 Longer text, empty/invalid input and supported voice/language controls Valid bounded output; documented behavior for empty input; unsupported controls fail explicitly. Only claim languages/voices actually tested.
S/T03 Streaming where supported Ordered text/audio chunks, clean completion, and valid assembled output; first partial transcript / first playable audio latency. Unsupported streaming is explicit, not silently counted as passed.
S/T04 Cancellation, second request, unload/reload and failure recovery No stale chunks, mixed requests or stuck state; a subsequent request succeeds; resources are released. Include repeated execution under the same model.
E01 Record/import → STT → Gemma 4 or Qwen3.5 response → TTS → playback One bounded voice round trip using the actual public API and app path; retain intermediate transcript, response and output audio. Exercise permission denial, stop/cancel and recovery where capture/playback exist.

Use fixed short fixture text/audio, deterministic settings where supported, immutable model/media hashes, exact prompts and resolved inference configuration. Establish correctness thresholds against the selected model's native reference before qualification; never loosen them merely to turn a failing device green. Listening checks are recorded separately from automated checks. Use non-personal fixtures and do not retain microphone recordings by default.

Platform/backend matrix

For each selected speech artifact, enumerate macOS, Linux, Windows, Android, iOS and Web against every backend the pinned runtime actually supports. Start with Mac CPU reference, then one representative Firebase Android and iOS device, then desktop/browser supported lanes. CPU/GPU/NPU are separate evidence rows: hardware availability does not establish speech-runtime support. Unsupported combinations have a reason; unexecuted supported combinations remain NOT_RUN. Existing general chat/backend results do not qualify speech.

Firebase fixture-based inference can qualify speech processing without proving real microphone capture, speaker output or acoustic quality. Test app permissions and audio routing where the environment permits, and retain an explicit separate manual/device check for actual microphone-to-speaker behavior when the lab cannot provide it. Browser microphone/playback restrictions require browser evidence, not native iOS test evidence.

Metrics, logs and reports

STT: WER (CER where appropriate), audio duration, transcription latency, first partial latency if streamed, and real-time factor = processing seconds / input audio seconds. TTS: time to first playable audio, total generation time, output duration and real-time factor = generation seconds / output audio seconds. Report load time, failures and peak memory where observable. Token throughput may be supplementary if native counters exist; it must not replace speech metrics.

Every JSON/CSV/HTML report retains model/artifact/runtime identity, platform/device, requested and observed backend, fixture hashes, resolved configuration, status and reason, timing source, raw output transcript or non-personal generated audio artifact, and correctness/listening results. Missing counters remain null. Graph speech separately: STT WER versus real-time factor, TTS first-audio latency and real-time factor, and a platform/backend coverage heatmap filtered by exact model/fixture/config. Never mix chat decode TPS with audio throughput.

Completion requires a runnable registered speech pack and app round-trip scenario, model locks and reference evidence, supported/unsupported matrix, real execution reports for the selected baseline platforms, and clearly tracked remaining rows. Integrate scenarios with the existing test matrix/local E2E runner; no orphan scripts. Cloud runs remain bounded and require freshly verified free allowance or covered credit; this plan update starts no paid resources.

Primary model focus (2026-09-17)

Gemma 4 and Qwen3.5 are the primary validation families. This priority supersedes older candidate-model priorities in the checked-in plan; historical results and known failures below remain valid evidence for their original models only. New artifacts listed here are candidates, not completed llamadart qualification.

Primary model Artifact / runtime lanes Initial scope
Qwen3.5 0.8B Existing pinned GGUF Q4_0; separate LiteRT-LM INT8 text artifact Lightweight text, streaming, history, cancellation/reload, batching and TPS. Thinking/tools/structured output require artifact-specific support and reference fixtures; do not infer parity from family name.
Gemma 4 E2B IT GGUF with matching projector; native LiteRT-LM; separate Web LiteRT-LM artifact Text core first, then image/audio where each artifact and public runtime support it. The Web artifact is currently text-only.
Gemma 4 E2B IT NPU Tensor G5-specific LiteRT-LM candidate for Pixel 10 Verify exact artifact, vendor kit, SoC, installed-app execution and native/public/CPU comparison before claiming NPU coverage.

Artifact sources: Qwen3.5 0.8B LiteRT-LM, Gemma 4 E2B LiteRT-LM files. The Qwen text export uses a simplified template without thinking/tool sections; its separate VL export is a later candidate. GPU memory must be checked per device. Gemma's listed Qualcomm mobile artifact targets SM8750, not the existing S24's SM8650; it is not an interchangeable S24 test input.

Keep tiny GGUF only for fast packaging/lifecycle smoke. Retain Gemma 3 and Qwen 3 profiles/results as legacy diagnostic and regression controls, including #509/#513; do not count them as primary-family coverage or discard unresolved failures. Larger Gemma 4/Qwen3.5 variants are opt-in stress tests. Specialty embedding and other-family fixtures remain deferred exceptions. STT and TTS are required dedicated-model exceptions as specified above.

Execution order: lock immutable revisions, sizes, hashes, templates and capability expectations; qualify Mac CPU/Metal and LiteRT CPU/GPU where supported; run selected Firebase Android/iOS CPU/GPU cases; qualify browser WASM/WebGPU and LiteRT Web; attempt compatible Gemma 4 NPU and Linux/Windows GPU rows only with supported artifacts and available resources. Each platform/backend row records PASS/FAIL/NOT_RUN/UNSUPPORTED separately, actual accelerator evidence, exact model/runtime identity, output/predicate results, latency and comparable decode TPS. No full Cartesian product is required.

Zero out-of-pocket remains mandatory: verify current free allowance or covered credit before each cloud dispatch. This tracker update does not migrate runnable profiles, change PR #515's audited head, or start cloud resources.

Current implementation: draft PR #515, head c29999eff6901baa92c3b29a0ef56ed250580628.

R02/R03 progress: C11 native batching parity and LiteRT Web per-option rejection/recovery are implemented, along with the earlier second lifecycle cycle. Catalog version 2 preserves older case definitions and keeps unimplemented or unqualified combinations explicit. Five extended catalog cases remain NOT_RUN; tool-bearing C11, deterministic NPU parity and GGUF Web controls still need qualification.

R11 progress: imported reports now require verified model SHA256/size matching the lock. Desktop qualification verifies the runtime payload, rejects overrides/sidecars, anchors discovery to the bundle and records its hash. Collection matches all uploaded source/runtime identities. JIT and older desktop journals without payload evidence remain diagnostic.

Validation: 82 private tests, 46 provider/NPU tests, 57 pin tests and 2 actual Chrome public-engine tests pass. Runtime source 8c5ed9401bca2d226bdbefafdc9d4015ec6067bd passes tiny GGUF CPU and Metal 12/12 each, with C11 chunks 5/32/5 and identical output. Gemma LiteRT CPU is 14 PASS / 4 known history FAIL, C11 PASS. The final CI Mac artifact 10522918241, run 35284547940, passes 12/12 CPU and 12/12 Metal outside the checkout. All 33 files were verified; its synthetic merge has the exact final head/base parents and identical tree. All 22 current-head CI checks pass, all six bundle jobs pass, and the exact-head independent audit is accepted with zero remaining PR-caused P1 blockers or unresolved review threads. R01's current-head build/review checks are complete. PR remains draft; no merge was performed.

Independent audit reproduced two false-qualification gaps; the fixes above close both. Seven branch-removal/bypass mutation checks prove the relevant assertions fail when protections are removed. The repository-local evaluator validates the evidence and returns the documented external-prerequisites limitation; external enforcement was not configured. Full runtime/device qualification remains incomplete; no semantic predicate was weakened. Qwen's strict Cedar17 vs cedar17 history mismatch also predates C11, as shown by the previous CI artifact control.

#516 is merged and #511 is closed. No Firebase test, VM or paid resource was started for this increment. Keep product failures in #509, #513 and their native-owner investigations.

13. Remaining work checklist (2026-09-17)

GitHub tracker: #514.

This is the remaining scope from the full plan, not a claim that all rows belong
in the first PR. The initial PR delivers the quick core, reports, portable build
and cloud adapters, NPU diagnostics and the discovered system-message correction.
Its CI and independent review must finish before merge readiness. Device/model
failures remain visible and are investigated separately from harness completion.

ID Remaining work Completion evidence
R01 Current-head PR/build qualification complete at c29999ef; rerun after substantive head/base changes Exact-head CI green, extracted bundles executable outside a checkout, checksums/manifests retained; independent high-risk review before ready. iOS physical signing remains on the Mac.
R02 Complete C05 thinking/budgets, C07 tool auto/required/none and continuation, C10 stop-marker semantics, C12 guard/recovery subcases, C02 separate Unicode generation and C11 tool-bearing parity C09's second cycle and C11 native text/thinking parity plus LiteRT Web option rejection are implemented. The five unimplemented catalog cases remain NOT_RUN. C11 tool-bearing fixtures, NPU deterministic sampling and GGUF Web worker controls remain unqualified; do not substitute text-only evidence.
R03 Core focused selection and versioned metadata implemented; extend the catalog as future model/media packs land Schema 2 binds case/feature versions, resolved prompts/tools/predicates and fixture hashes; omitted cases have explicit reasons. tiny-gguf-lifecycle and tiny-gguf-batching are runnable focused examples. Catalog version 2 imports older version-1 reports against their original definitions. Future pack media hashes and reference qualifications remain with R04/R05.
R04 Add the eleven targeted packs in section 6 Structured output; state/prompt reuse; embeddings; vision; audio understanding; ASR; TTS; LoRA; speculative decoding; runtime controls; app/device/browser lifecycle. Reuse existing registered tests and keep large models opt-in.
R05 Lock and reference-qualify pack models/media Prioritize Qwen3.5 0.8B GGUF/LiteRT text and Gemma 4 E2B GGUF/native/Web/Tensor G5 candidates with matching projectors/media. Verify exact revisions/hashes/access, templates, runtime compatibility and memory limits. Lock one supported STT and one TTS artifact for the required speech packs; defer other specialty and larger-model packs. Current quick/NPU locks do not qualify new candidates.
R06 Finish Firebase core device rotation Isolated GGUF CPU/GPU and LiteRT CPU/GPU on S24, Tab P12, iPhone 16 Pro and SE 3; targeted iPad 10 GPU runs. Existing pilots are partial evidence, not a completed rotation. A05s full/compact, iPhone 8 and Pixel 5 remain later compatibility rows.
R07 Prioritize compatible Gemma 4 NPU; preserve S24 diagnostic work Gemma 4 Tensor G5 is a Pixel 10 candidate pending immutable lock, vendor-kit compatibility and installed-app native/public/CPU runs. SM8750 artifacts do not qualify S24 SM8650. Preserve #513 and existing Gemma 3 S24 controls as legacy diagnostics. Require Unicode/tokenizer, coherent outputs, lifecycle and placement evidence; a compiled dispatch library is not execution proof.
R08 Fill remaining platform/packaging rows Android arm64 virtual 4K/16K and separate backcompat, Android x64 emulator, Apple simulators, Linux arm64, Windows arm64, macOS x64 as available; full/compact and lower-ISA physical coverage. Record unavailable hardware explicitly.
R09 Qualify browser and GPU evidence paths LiteRT GPU adapters with actual driver/delegate proof; Chrome WASM/WebGPU, Safari and Firefox capability rows; a genuine LiteRT Web model bundle and negative native-only contracts. Native Firebase XCTest does not qualify iPadOS Safari.
R10 Exercise the GCE lifecycle and desktop CUDA runs One bounded Linux then Windows GGUF/CUDA run proving upload, execution, retrieval and deletion of instance/disks; inspect actual NVIDIA execution. Recheck current credit before provisioning. LiteRT desktop GPU is Vulkan/D3D12, not CUDA. No available credit means NOT_RUN, not personal charges.
R11 Complete the required evidence envelope and missing measurements Model hash/size evidence, portable desktop payload verification, override rejection and collected identity binding are implemented. Remaining: structured preparation failures before a model manifest; explicit provenance/availability for artifact and companion hashes, device memory/page size, cold first response, prefill/native TTFT and first-thinking timing where observable. Keep unsupported counters null; add ASR/TTS WER/real-time factor with their packs.
R12 Add aggregate and historical reporting after core evidence is stable Paired run comparison, device/backend coverage heatmap, comparable-cohort filters, trend and quota views. Existing per-run JSON/JUnit/CSV/HTML and three-sample TPS are usable; a dashboard or performance threshold is not required for the first PR.

Prioritize R01, then bounded R02/R03 work. R06/R07 use only freshly verified free
Firebase allowance or covered credit; never dispatch the entire rotation at once.
R10 is optional while credit is unavailable. R04/R05 are change-focused feature
coverage, not every-model-by-every-device permutations. R12 visual polish comes
after the mandatory evidence, not before correctness.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or requestpriority:P2Planned next: useful unblocked work or validation after P1 items

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions