Add portable validation suite and guarded Firebase/GCE runners - #515
Merged
Merged
Conversation
Contributor
|
Chat app preview removed for |
This was referenced Sep 17, 2026
This was referenced Sep 18, 2026
…cross-platform-validation
…cross-platform-validation # Conflicts: # CHANGELOG.md
…cross-platform-validation # Conflicts: # CHANGELOG.md
…m-validation # Conflicts: # CHANGELOG.md # doc/testing_matrix.md # tool/testing/test_matrix.dart # website/docs/changelog/recent-releases.md
…m-validation # Conflicts: # CHANGELOG.md # doc/testing_matrix.md # website/docs/changelog/recent-releases.md
This was referenced Sep 18, 2026
# Conflicts: # CHANGELOG.md # website/docs/changelog/recent-releases.md
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Add a private, opt-in suite that executes the public llamadart API through portable desktop, Flutter and Web adapters. Locked model/runtime inputs and prompts produce JSON/JUnit/CSV/HTML reports with output assertions, latency/TPS, accelerator placement and cleanup evidence. Gemma 4 E2B and Qwen3.5 0.8B are the primary chat families; local speech diagnostics cover Qwen3-ASR, Qwen3-TTS, Moonshine and file STT → primary chat → TTS.
Guarded Firebase/GCE tooling provides plan, run, status, collect, reconcile and cleanup commands. Firebase polling now records safe typed failures, bounds read retries under the original deadline and collects terminal evidence after failure cleanup. Manual recovery invalidates stale collections and refreshes status. Fixes #518. Windows VM upload and collection explicitly select legacy SCP after same-VM default/legacy transfer probes reproduced the default protocol failure. Linux retains its default protocol. Submission is never retried automatically, and provider success cannot replace missing assertions.
Catalog 4 adds exact Unicode generation, thinking on/off controls, and tool choice/result roundtrips to the existing stop-sequence and unloaded-engine readiness/reload cases. GPU/NPU evidence accounts for their additional model loads and generations. Historical catalog-1/2/3 reports preserve their original obligations. Vulkan placement accepts indexed discovery or selected-model device identity, rejects conflicting sources, and retains raw software-device checks before discovery annotation normalization; software, missing, conflicting and unknown device identities remain unverified. Independent guard-bypass mutation probes fail as expected. Windows backend discovery now recognizes the standard compiled CLI
bin/../liblayout, retaining override and adjacent-module precedence (fixes #525). Desktop x64 bundles include CPU/Vulkan/CUDA modules (and Windows CUDA dependencies), rejecting incomplete archives before sealing.Current validation and scope
Head
051fbd233290f8387fb5aff197eac3fac87c6d23, integrated main699969b070d75443784ed9bb4d933f5a2988ea2a. Private maintainer tooling; draft infrastructure review, not universal model/platform qualification.Cedar17versus requiredcedar17. No strict assertion was relaxed and no production fix was justified by this bounded result.libdiscovery. Separate merged runtime fixes are consumed from main.Cedar17versuscedar17; C07 tool roundtrips pass. Both reports have complete provenance and verified payload. Artifact10576610456from workflow run35416263110was built at synthetic mergec1879c33c8937af1ea2d36ad108b2a663e1e7c56, whose parents are the exact base/head above and whose tree equals PR head. This verifies macOS CPU portability, not other platform or accelerator lanes. Inherited library-path overrides were rejected as designed and removed for these runs.Historical hardware evidence
The following records predate catalog4 and retain their original sources/failures. They are not current-head reruns; historical NOT_RUN entries are not upgraded by new test implementations. Earlier setup failures coexist with later successful retries. Accelerator UNVERIFIED means missing execution proof, not a compatibility verdict.
Latest coverage at a glance
These historical release rows remain unqualified: thinking, tools and Unicode generation were NOT_RUN in their original catalog. Build success, execution, placement and full qualification are different claims. These desktop runs do not cover macOS x64 or Linux/Windows arm64.
Detailed run evidence
89097b7089097b7089097b7089097b7089097b7089097b7089097b7089097b7089097b7089097b70unknownunknownunknownunknownunknownunknownunknown89097b7089097b7089097b7089097b7089097b7089097b7089097b7089097b7089097b7089097b70f1a13f35f1a13f35f1a13f35f1a13f3589097b7089097b7089097b7089097b7089097b7089097b7089097b7089097b70f1a13f35f1a13f35f1a13f35f1a13f35f1a13f35f1a13f35f1a13f35f1a13f35f1a13f35f1a13f35TPS values are three short B01 samples, not sustained benchmarks or equivalent cross-runtime workloads. Wall TPS uses retokenized visible output; native decode TPS uses runtime counters. GPU native counters report a much narrower interval than wall time and must not be treated as end-to-end throughput. Manifests preserve prompts, inference settings, model/runtime hashes, original outputs and environment. CPU fallback is never GPU coverage.
Firebase and lifecycle
Physical iPhone16Pro/iOS18.3 tiny GGUF CPU run: 43 seconds, provider SUCCESS, public assertions PASS, qualified=true, collection COMPLETE, cleanup VERIFIED. Matrix result. This verifies infrastructure, not primary models or Metal. Final recovery code also re-collected the original failed matrix with saved terminal FINISHED/INCONCLUSIVE/CANCELLED and qualified=false (#518).
The earlier S24 Gemma 4 LiteRT GPU attempt hit its 600-second download deadline before inference; primary Android/iOS/NPU qualification remains open. Firebase retry used fresh conservative free-allowance accounting; missing duration for the cancelled test was reserved at full timeout plus rounding, not counted as zero. No further Firebase run is claimed.
Cloud lifecycle
The earlier loader-only Linux Vulkan follow-up discovered only llvmpipe software rendering. It was stopped before inference and cannot establish physical Vulkan or LiteRT GPU coverage. The reporter now requires matching physical-device identity, with negative and mutation tests.
Final corrected Linux run completed all four requested profiles; Windows completed all ten. Both final VMs, boot disks and temporary firewalls were deleted. A fresh project-wide inventory at 2026-09-18T14:33:41Z contains no instances, disks or task firewalls.
The first exact-driver Linux graphics attempt rejected the newer live Ubuntu package before inference and verified VM/disk/firewall deletion. The corrected attempt uses Canonical snapshot 20260801T000000Z, whose signed apt metadata supplies matching 580.173.02 graphics libraries; status is recorded below.
Earlier Windows attempts were cleaned up: one failed during initial boot probing; the next reached NVIDIA driver 582.53 but failed SCP before inference. A subsequent zone-capacity failure also left no VM or disk. Current retry status is recorded below.
linux:
qa-linux-1789733133, us-central1-c, provider DELETE deadline 2026-09-18T13:05:41.916945+00:00; VM/disk cleanup VERIFIED, temporary firewall deleted=True.windows:
qa-windows-1789736203, us-central1-a, provider DELETE deadline 2026-09-18T13:56:50.921966+00:00; VM/disk cleanup VERIFIED, temporary firewall deleted=True.linux-vulkan:
qa-linux-vulkan-1789735442, us-central1-c, provider DELETE deadline 2026-09-18T13:44:11.810229+00:00; VM/disk cleanup VERIFIED, temporary firewall deleted=True.windows-fixed-f1a13f3/gce-windows:
qa-windows-1789738859, us-central1-a, provider DELETE deadline 2026-09-18T14:41:07.823201+00:00; VM/disk cleanup VERIFIED, temporary firewall deleted=True.linux-graphics-f1a13f3/gce-linux-vulkan:
qa-linux-vulkan-1789741139, us-central1-a, provider DELETE deadline 2026-09-18T15:19:06.059690+00:00; VM/disk cleanup VERIFIED, temporary firewall deleted=True.The VM runs use bounded bootstrap scripts with pinned public images, exact archive checksums, provider deletion deadlines and ownership-checked teardown. They are actual portable public-API execution evidence, but do not claim the maintained GCE adapter has a qualified custom image. Windows/Linux preparation or partial runs are not platform qualification.
Prior speech evidence (separate source)
Local source 21f89f9 diagnostics: Qwen3-TTS CPU/Metal functional checks and all four file STT→Gemma/Qwen×GGUF/LiteRTCPU→TTS chains passed. Qwen3-ASR file WER 0 but identical bytes failed (#517); Moonshine WER 0.22727 failed the strict oracle. The #523 fix is now integrated: locked Qwen3-ASR file and bytes diagnostics were repeated on clean source 0c0a379, CPU and Metal, with WER 0 and functional_pass=true. They retain qualified=false. TTS, Moonshine and complete voice chains were not rerun; this is not complete speech/device qualification.
Current high-risk review evidence
Independent review verified remote head
051fbd233290f8387fb5aff197eac3fac87c6d23against actual main699969b070d75443784ed9bb4d933f5a2988ea2a: zero known PR-caused P1 regressions and zero unresolved review threads. Independent validation: 129 private suite tests and 253 relevant root tests passed (one existing skip). No independent hardware qualification claimed.Evaluator: unverifiedPrerequisites / externalPrerequisitesUnavailable (exit 2). Local evidence is internally consistent; auditor authentication, GitHub App publication, protected-environment provenance and conditional ruleset enforcement are unavailable. This is not operational merge readiness. Keep draft.
Exact-head independent audit and evaluator result
{ "schema": "llamadart.high-risk-readiness-evidence", "schema_version": "1.0.0", "timestamp": "2026-09-19T02:39:22Z", "correlation_id": "pr515-051fbd23-current-main-audit", "repository": "leehack/llamadart", "pr_number": 515, "expected_pr_head_sha": "051fbd233290f8387fb5aff197eac3fac87c6d23", "current_base_sha": "699969b070d75443784ed9bb4d933f5a2988ea2a", "pr_author": "leehack", "classification": "high-risk", "surfaces": [ "artifactConsumer", "backendRuntime", "regressionPolicy" ], "required_matrix_row_ids": [ "high-risk-exact-head-independent-qa" ], "matrix_row_evidence": { "high-risk-exact-head-independent-qa": { "row_id": "high-risk-exact-head-independent-qa", "result": "pass", "command": "Fresh risk-bearing call-site inspection and byte-continuity verification; private dart test129pass; root VM LiteRTservice/conversation, llamaCppservice, remote and NPU provider tests253pass1MacABI skip.", "evidence_notes": "Fresh full-risk local review and Git blob comparison confirm suite/provider/mobile hosts/remaining runtime deltas unchanged from independently acceptedb13. Currentmain529/531 integrations preserve their independently reviewed/postmerge-tested implementations. Re-inspected report/catalog/config immutability, all-pass qualification, GPU/NPU proofs, payload sealing/collector identity, remote deadline/terminal recovery/cleanup, speech diagnostics and LiteRTsystem/Windows directory callsites. Independently129private and253root tests pass with1existingMacABI skip. Docs retain C06FAIL/qualifiedfalse and distinguish bootstrap execution/implementation from platform qualification. Current GitHub PRhead/actualmain/author and zero unresolved review threads independently verified. Model/C06reference/full broad gates are author evidence, not independent hardware execution." } }, "independent_audit": { "auditor_identity": "codex-adversarial-current-main-pr515-051fbd23", "audit_kind": "codex-adversarial", "audit_head_sha": "051fbd233290f8387fb5aff197eac3fac87c6d23", "audit_base_sha": "699969b070d75443784ed9bb4d933f5a2988ea2a", "decision": "accepted", "unresolved_review_threads": 0, "known_pr_caused_p1_regressions": 0, "summary": "Fresh full-risk local review and Git blob comparison confirm suite/provider/mobile hosts/remaining runtime deltas unchanged from independently acceptedb13. Currentmain529/531 integrations preserve their independently reviewed/postmerge-tested implementations. Re-inspected report/catalog/config immutability, all-pass qualification, GPU/NPU proofs, payload sealing/collector identity, remote deadline/terminal recovery/cleanup, speech diagnostics and LiteRTsystem/Windows directory callsites. Independently129private and253root tests pass with1existingMacABI skip. Docs retain C06FAIL/qualifiedfalse and distinguish bootstrap execution/implementation from platform qualification. Current GitHub PRhead/actualmain/author and zero unresolved review threads independently verified. Model/C06reference/full broad gates are author evidence, not independent hardware execution." }, "structured_output_evidence": null, "affected_test_paths": [ "test/unit/backends/litert_lm/litert_lm_service_test.dart", "test/unit/backends/litert_lm/litert_lm_backend_web_test.dart", "test/unit/tooling/validation_remote_test.dart", "test/unit/tooling/validation_npu_test.dart", "test/unit/backends/llama_cpp/llama_cpp_service_test.dart" ], "evaluation": { "evaluated_at": "2026-09-19T02:39:22.650769Z", "changed_files": [ { "path": ".github/workflows/ci.yml", "status": "modified" }, { "path": ".github/workflows/validation_bundles.yml", "status": "added" }, { "path": ".gitignore", "status": "modified" }, { "path": "CHANGELOG.md", "status": "modified" }, { "path": "doc/cross_platform_validation.md", "status": "added" }, { "path": "doc/cross_platform_validation_plan.md", "status": "added" }, { "path": "doc/testing_matrix.md", "status": "modified" }, { "path": "example/chat_app/android/app/build.gradle.kts", "status": "modified" }, { "path": "example/chat_app/android/app/src/main/kotlin/com/example/llamadart_chat_example/MainActivity.kt", "status": "modified" }, { "path": "example/chat_app/android/app/src/main/kotlin/com/example/llamadart_chat_example/ValidationNpuHost.kt", "status": "added" }, { "path": "example/chat_app/android/gradle.properties", "status": "modified" }, { "path": "example/chat_app/integration_test/validation_test.dart", "status": "added" }, { "path": "example/chat_app/ios/Runner.xcodeproj/project.pbxproj", "status": "modified" }, { "path": "example/chat_app/ios/RunnerTests/RunnerTests.m", "status": "added" }, { "path": "example/chat_app/ios/RunnerTests/RunnerTests.swift", "status": "deleted" }, { "path": "example/chat_app/lib/validation/controller.dart", "status": "added" }, { "path": "example/chat_app/lib/validation/host.dart", "status": "added" }, { "path": "example/chat_app/lib/validation/host_native.dart", "status": "added" }, { "path": "example/chat_app/lib/validation/host_web.dart", "status": "added" }, { "path": "example/chat_app/lib/validation_main.dart", "status": "added" }, { "path": "example/chat_app/pubspec.lock", "status": "modified" }, { "path": "example/chat_app/pubspec.yaml", "status": "modified" }, { "path": "example/chat_app/test/validation_app_test.dart", "status": "added" }, { "path": "example/chat_app/test/validation_controller_test.dart", "status": "added" }, { "path": "lib/src/backends/litert_lm/litert_lm_service.dart", "status": "modified" }, { "path": "lib/src/backends/llama_cpp/llama_cpp_service.dart", "status": "modified" }, { "path": "packages/llamadart_validation/README.md", "status": "added" }, { "path": "packages/llamadart_validation/analysis_options.yaml", "status": "added" }, { "path": "packages/llamadart_validation/assets/profiles/chat-gguf-cpu.json", "status": "added" }, { "path": "packages/llamadart_validation/assets/profiles/chat-gguf-cuda.json", "status": "added" }, { "path": "packages/llamadart_validation/assets/profiles/chat-gguf-metal.json", "status": "added" }, { "path": "packages/llamadart_validation/assets/profiles/chat-gguf-vulkan.json", "status": "added" }, { "path": "packages/llamadart_validation/assets/profiles/chat-litert-cpu.json", "status": "added" }, { "path": "packages/llamadart_validation/assets/profiles/chat-litert-gpu.json", "status": "added" }, { "path": "packages/llamadart_validation/assets/profiles/gemma3-litert-cpu.json", "status": "added" }, { "path": "packages/llamadart_validation/assets/profiles/gemma4-gguf-cpu.json", "status": "added" }, { "path": "packages/llamadart_validation/assets/profiles/gemma4-gguf-cuda.json", "status": "added" }, { "path": "packages/llamadart_validation/assets/profiles/gemma4-gguf-metal.json", "status": "added" }, { "path": "packages/llamadart_validation/assets/profiles/gemma4-gguf-vulkan.json", "status": "added" }, { "path": "packages/llamadart_validation/assets/profiles/gemma4-litert-cpu.json", "status": "added" }, { "path": "packages/llamadart_validation/assets/profiles/gemma4-litert-gpu.json", "status": "added" }, { "path": "packages/llamadart_validation/assets/profiles/npu-qualcomm-sm8650.json", "status": "added" }, { "path": "packages/llamadart_validation/assets/profiles/npu-tensor-g5.json", "status": "added" }, { "path": "packages/llamadart_validation/assets/profiles/qwen35-litert-cpu.json", "status": "added" }, { "path": "packages/llamadart_validation/assets/profiles/qwen35-litert-gpu.json", "status": "added" }, { "path": "packages/llamadart_validation/assets/profiles/tiny-gguf-batching.json", "status": "added" }, { "path": "packages/llamadart_validation/assets/profiles/tiny-gguf-cpu.json", "status": "added" }, { "path": "packages/llamadart_validation/assets/profiles/tiny-gguf-cuda.json", "status": "added" }, { "path": "packages/llamadart_validation/assets/profiles/tiny-gguf-lifecycle.json", "status": "added" }, { "path": "packages/llamadart_validation/assets/profiles/tiny-gguf-metal.json", "status": "added" }, { "path": "packages/llamadart_validation/assets/profiles/tiny-gguf-vulkan.json", "status": "added" }, { "path": "packages/llamadart_validation/assets/speech/README.md", "status": "added" }, { "path": "packages/llamadart_validation/assets/speech/jfk.wav", "status": "added" }, { "path": "packages/llamadart_validation/assets/speech/litert-asr.json", "status": "added" }, { "path": "packages/llamadart_validation/assets/speech/stt.json", "status": "added" }, { "path": "packages/llamadart_validation/assets/speech/tts.json", "status": "added" }, { "path": "packages/llamadart_validation/bin/dedicated_speech.dart", "status": "added" }, { "path": "packages/llamadart_validation/bin/report.dart", "status": "added" }, { "path": "packages/llamadart_validation/bin/run.dart", "status": "added" }, { "path": "packages/llamadart_validation/bin/speech.dart", "status": "added" }, { "path": "packages/llamadart_validation/bin/voice.dart", "status": "added" }, { "path": "packages/llamadart_validation/lib/io.dart", "status": "added" }, { "path": "packages/llamadart_validation/lib/llamadart_validation.dart", "status": "added" }, { "path": "packages/llamadart_validation/lib/npu_io.dart", "status": "added" }, { "path": "packages/llamadart_validation/lib/src/case_catalog.dart", "status": "added" }, { "path": "packages/llamadart_validation/lib/src/desktop_bundle.dart", "status": "added" }, { "path": "packages/llamadart_validation/lib/src/manifest.dart", "status": "added" }, { "path": "packages/llamadart_validation/lib/src/native_reference_request.dart", "status": "added" }, { "path": "packages/llamadart_validation/lib/src/npu_evidence.dart", "status": "added" }, { "path": "packages/llamadart_validation/lib/src/npu_monitor_io.dart", "status": "added" }, { "path": "packages/llamadart_validation/lib/src/npu_reference_io.dart", "status": "added" }, { "path": "packages/llamadart_validation/lib/src/placement.dart", "status": "added" }, { "path": "packages/llamadart_validation/lib/src/report.dart", "status": "added" }, { "path": "packages/llamadart_validation/lib/src/runner.dart", "status": "added" }, { "path": "packages/llamadart_validation/lib/src/runtime_environment.dart", "status": "added" }, { "path": "packages/llamadart_validation/lib/src/runtime_environment_io.dart", "status": "added" }, { "path": "packages/llamadart_validation/lib/src/runtime_environment_stub.dart", "status": "added" }, { "path": "packages/llamadart_validation/lib/src/speech_runner.dart", "status": "added" }, { "path": "packages/llamadart_validation/lib/src/voice_runner.dart", "status": "added" }, { "path": "packages/llamadart_validation/pubspec.lock", "status": "added" }, { "path": "packages/llamadart_validation/pubspec.yaml", "status": "added" }, { "path": "packages/llamadart_validation/schemas/event.schema.json", "status": "added" }, { "path": "packages/llamadart_validation/schemas/profile.schema.json", "status": "added" }, { "path": "packages/llamadart_validation/test/coverage_catalog_test.dart", "status": "added" }, { "path": "packages/llamadart_validation/test/dedicated_speech_test.dart", "status": "added" }, { "path": "packages/llamadart_validation/test/desktop_bundle_test.dart", "status": "added" }, { "path": "packages/llamadart_validation/test/model_io_test.dart", "status": "added" }, { "path": "packages/llamadart_validation/test/native_reference_request_test.dart", "status": "added" }, { "path": "packages/llamadart_validation/test/npu_evidence_test.dart", "status": "added" }, { "path": "packages/llamadart_validation/test/public_engine_batching_test.dart", "status": "added" }, { "path": "packages/llamadart_validation/test/speech_runner_test.dart", "status": "added" }, { "path": "packages/llamadart_validation/test/validation_test.dart", "status": "added" }, { "path": "packages/llamadart_validation/test/voice_runner_test.dart", "status": "added" }, { "path": "scripts/build_chat_app_web.sh", "status": "modified" }, { "path": "test/unit/backends/litert_lm/litert_lm_backend_web_test.dart", "status": "modified" }, { "path": "test/unit/backends/litert_lm/litert_lm_conversation_template_test.dart", "status": "modified" }, { "path": "test/unit/backends/litert_lm/litert_lm_service_test.dart", "status": "modified" }, { "path": "test/unit/backends/llama_cpp/llama_cpp_service_test.dart", "status": "modified" }, { "path": "test/unit/tooling/prepare_workspace_test.dart", "status": "modified" }, { "path": "test/unit/tooling/validation_npu_test.dart", "status": "added" }, { "path": "test/unit/tooling/validation_remote_test.dart", "status": "added" }, { "path": "tool/prepare_workspace.dart", "status": "modified" }, { "path": "tool/testing/run_local_e2e.dart", "status": "modified" }, { "path": "tool/testing/test_matrix.dart", "status": "modified" }, { "path": "tool/testing/validation.dart", "status": "added" }, { "path": "tool/testing/validation/bundle.dart", "status": "added" }, { "path": "tool/testing/validation/check_npu_apk.py", "status": "added" }, { "path": "tool/testing/validation/collect.dart", "status": "added" }, { "path": "tool/testing/validation/coverage_catalog.dart", "status": "added" }, { "path": "tool/testing/validation/firebase.example.json", "status": "added" }, { "path": "tool/testing/validation/firebase_blaze.example.json", "status": "added" }, { "path": "tool/testing/validation/gce.example.json", "status": "added" }, { "path": "tool/testing/validation/npu.dart", "status": "added" }, { "path": "tool/testing/validation/process.dart", "status": "added" }, { "path": "tool/testing/validation/process_host.py", "status": "added" }, { "path": "tool/testing/validation/remote.dart", "status": "added" }, { "path": "tool/testing/validation/run-remote.ps1", "status": "added" }, { "path": "tool/testing/validation/run-remote.sh", "status": "added" }, { "path": "tool/testing/validation/runtime_bundle.dart", "status": "added" }, { "path": "tool/testing/validation/test_npu_apk.py", "status": "added" }, { "path": "website/docs/changelog/recent-releases.md", "status": "modified" } ], "decision": "unverifiedPrerequisites", "failure_classification": "externalPrerequisitesUnavailable", "message": "Repository-local evidence is internally consistent, but auditor authentication, GitHub App publication, protected-environment provenance, and ruleset enforcement are not available. This is not operational merge readiness.", "external_prerequisites": { "app_installed": false, "protected_environment_configured": false, "independent_auditor_authenticated": false, "ruleset_enforced": false, "diagnostic_message": "No repository-local input can authenticate the dedicated GitHub App, protected environment, independent auditor, or conditional ruleset. See doc/high_risk_pre_merge_readiness.md." } } }