Skip to content

test(llup): replay automated v0.3 qualification evidence - #492

Draft
daniele21 wants to merge 12 commits into
agent/llup-v0-3-upgrade-rebasedfrom
agent/llup60-automated-replay
Draft

test(llup): replay automated v0.3 qualification evidence#492
daniele21 wants to merge 12 commits into
agent/llup-v0-3-upgrade-rebasedfrom
agent/llup60-automated-replay

Conversation

@daniele21

@daniele21 daniele21 commented Aug 30, 2026

Copy link
Copy Markdown
Owner

Scope

Stacked LLUP-60A evidence-only replay on top of frozen runtime candidate PR #490.

  • runtime source: 59af48313b450d9cff13c7f43458c2e5e6560374;
  • evidence HEAD: 2e213ad465bef56b724708e9b2b641bba0de3be5;
  • base: agent/llup-v0-3-upgrade-rebased@59af48313b450d9cff13c7f43458c2e5e6560374.

The diff is intentionally restricted to three evidence/governance files:

  • .github/workflows/llup60-automated-replay.yml;
  • docs/workstreams/llup-v0-3-automated-replay.md;
  • docs/workstreams/README.md for canonical active-workstream routing.

No runtime, backend, JNI, model-policy or product code is changed.

Refreshed qualification identities

The candidate was refreshed after dev gained the canonical Android AAB packaging command. Runtime behavior did not change, but exact build/package identity did, so pre-refresh LLUP-60 evidence is retained only as historical provenance.

Frozen pair:

  • control runtime: dev@80164329bbc41a00b75721e3d0524294c03fdb56, llama.cpp b9637 / aedb2a5e9ca3d4064148bbb919e0ddc0c1b70ab3;
  • candidate runtime: 59af48313b450d9cff13c7f43458c2e5e6560374, llama.cpp v0.3.0 / c1d0e7a004015f23bc0233470b747b596f29b264.

LLUP-60A automated replay — PASS

Authoritative exact-head evidence on 2e213ad465bef56b724708e9b2b641bba0de3be5:

  • LLUP-60 automated replay 33335263956: SUCCESS;
  • pull-request automatic Validate 33335266468: SUCCESS;
  • repository-owned requested /preflight strong run 33335275065: SUCCESS.

The repository selector escalated the requested preflight to FULL with reason unknown or repository-wide executable scope, because this evidence PR changes a GitHub Actions workflow. Effective selector:
profile=full android=true native=false packaging=false modules=all.

This escalation is accepted rather than manually downgraded.

Replay lanes passed:

  • evidence-only scope and exact runtime-parent provenance;
  • authoritative pin/submodule/backend/Q35 identity + regression fixture;
  • host-native backend ownership/API suite;
  • Q35 model-profile + runtime-core JVM contracts;
  • evaluation contracts across contracts/datasets/evaluators/engine/runtime-adapter/comparison/persistence;
  • observability contracts + benchmark-engine unit contracts;
  • exact Qwen3.5 0.8B and 2B GGUF download, digest/size verification and candidate llama-simple load/tokenize/generate smoke;
  • machine-readable replay manifest with runtime source SHA separated from evidence-harness SHA and promotion explicitly disabled while physical evidence is pending.

Published exact-head artifacts:

  • llup60-provenance;
  • llup60-qwen35-host-compatibility;
  • llup60-automated-replay-manifest.

Classify LLUP-60A = AUTOMATED_PREFLIGHT_CONFIRMED.

LLUP-50 deterministic preparation — PASS

Physical tooling owner PR #491 is refreshed at cefb893ae95cfd95339de4a24a955a652f0011e6 and its exact-head Validate run 33334618813 is green.

Frozen evidence refs:

  • control: evidence/llup50-control@fcbefc7cd9af84de570da96d039582175dd1700b;
  • candidate: evidence/llup50-candidate@a2a050d9551db541bb4c6b152cba8623c782164d.

Repository-owned package run 33334957429: SUCCESS. It produced exact-ref control, candidate and candidate-runtime app/test APK pairs plus source-revision/SHA-256 manifests.

Therefore there is no remaining deterministic package/tooling blocker before device execution.

Execution boundary

All remaining evidence is genuinely REAL_ENVIRONMENT:

  • LLUP-50 same-device control/candidate A/B;
  • physical LLRT KV-cache and evaluation-batch evidence;
  • device-only Q35 load/performance/PSS/memory/thermal/cancellation/recovery/lifecycle evidence;
  • representative-device resident-count, warm-idle and prepare-after-release recovery;
  • output-quality replay whose canonical execution path genuinely requires representative Android runtime evidence.

Global LLUP-60 is not complete until the required device-only LLUP-60B evidence passes.

Current workstream classification

All required deterministic automated gates for the frozen candidate/base and LLUP-50 package preparation are green on exact identities.

Classify the LLUP workstream boundary as WAITING_REAL_ENVIRONMENT.

This does not authorize promotion:

After LLUP-50 + LLUP-60B pass, re-check exact candidate/base freshness and FULL promotion requirements, then make the explicit LLUP-70 promote/reject decision.

Copy link
Copy Markdown
Owner Author

/preflight strong

Copy link
Copy Markdown
Owner Author

/preflight strong

Copy link
Copy Markdown
Owner Author

/preflight strong

Copy link
Copy Markdown
Owner Author

/preflight strong

Copy link
Copy Markdown
Owner Author

/preflight strong

Copy link
Copy Markdown
Owner Author

/preflight strong

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant