feat(BACKEND-ROCM): verify Qwen capture provenance - #2932
Conversation
Issue mudler#1588 still lacks active cache-state evidence and a three-mode ROCm correctness gate. This spec fixes the post-write probes, dtype audit, tolerance policy, tests, review mutations, and hardware evidence before implementation starts. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:gpt-5.6-sol [codex]
…before the instrumentation One new file, `.agents/specs/rocm-qwen35-08b-cpu-gfx1100-numerics.md`, and no product code. AGENTS.md's spec-before-code rule names this shape directly: a campaign agrees its scope before the implementation waves start, and a change that deliberately adds a spec without product code is one of the cases a separate pull request is for. The spec defines the Qwen3.5-0.8B CPU/gfx1100 comparison: production-boundary dumps, physical dtype and byte audits, cache-matched oracle captures, and descriptive layer deltas, keeping the exact upstream operation tolerances and the established end-to-end token/near-tie gate. It corrects the BF16 selector normalization, names the Qwen3.5 paged-attention path the engine actually takes, separates SD storage from DS views and dump order, and repairs the upstream attention-test anchor -- the four inaccuracies that failed the previous spec's review. What makes it mergeable is what it declines to claim. It introduces no runtime code, no model result and no numerical acceptance, and it says so: the capture tooling, state probes, comparator, provider checks and callsite mutation gates are unimplemented, the configured ROCm oracle still runs the historical vLLM pin `5559679229bc9618` rather than the active `e126687a9a828d51`, and cache-matched captures stay PENDING under #2773, which stays open. A spec that names its own gaps is the thing implementation waves can be reviewed against. Landed by local merge under the 2026-09-04 external-contributor grant in `.agents/developer-preferences.md`. `check-agent-record` reports the same ANCHOR-ROT=33 as `main` with this merged; the red CI checks are main's own baseline at `c796fea41`. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:claude-opus-5-1m [claude-code]
Strict near-tie publication accepted missing or contradictory runtime observations. Validate captured completeness and observed model and runtime identities against the independently verified current context. Publication could replace legacy inputs through destination aliases. Protect input file identities before writing any output. Keep ordinary legacy output replacement and valid installed revision prefixes. vLLM resolves sampling on a cloned request. Record constructor normalization separately and mark engine resolution unobserved. Refs mudler#2773. This repairs capture tooling in draft PR mudler#2932 on row/BACKEND-ROCM-NUMERICS-1588. GPU acceptance and the remaining characterization stay PENDING. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:gpt-6-astra [codex]
Legacy publication could overwrite the distribution metadata used to verify vLLM. Record the selected metadata files and protect their file identities. The final identity check also detects changes to those measured bytes. Distinct output paths could share an inode and replace NumPy payloads with JSON. Refuse output aliases before writing any destination. Preserve ordinary legacy replacement and the explicit default manifest path. Four permanent CLI tests cover metadata selection and fallbacks, publication aliases, metadata changes during capture, and successful replacement hashes. CPU fixtures and mutation checks cover both capture commands. Active-pin GPU acceptance and the remaining characterization stay PENDING under mudler#2773. Refs mudler#2773. This repairs draft PR mudler#2932 on row/BACKEND-ROCM-NUMERICS-1588. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:gpt-6-astra [codex]
… spec before the instrumentation One new file, `.agents/specs/rocm-qwen35-08b-cpu-gfx1100-numerics.md`, and no product code. AGENTS.md's spec-before-code rule names this shape directly: a campaign agrees its scope before the implementation waves start, and a change that deliberately adds a spec without product code is one of the cases a separate pull request is for. The spec defines the Qwen3.5-0.8B CPU/gfx1100 comparison: production-boundary dumps, physical dtype and byte audits, cache-matched oracle captures, and descriptive layer deltas, keeping the exact upstream operation tolerances and the established end-to-end token/near-tie gate. It corrects the BF16 selector normalization, names the Qwen3.5 paged-attention path the engine actually takes, separates SD storage from DS views and dump order, and repairs the upstream attention-test anchor -- the four inaccuracies that failed the previous spec's review. What makes it mergeable is what it declines to claim. It introduces no runtime code, no model result and no numerical acceptance, and it says so: the capture tooling, state probes, comparator, provider checks and callsite mutation gates are unimplemented, the configured ROCm oracle still runs the historical vLLM pin `5559679229bc9618` rather than the active `e126687a9a828d51`, and cache-matched captures stay PENDING under mudler#2773, which stays open. A spec that names its own gaps is the thing implementation waves can be reviewed against. Landed by local merge under the 2026-09-04 external-contributor grant in `.agents/developer-preferences.md`. `check-agent-record` reports the same ANCHOR-ROT=33 as `main` with this merged; the red CI checks are main's own baseline at `c796fea41`. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:claude-opus-5-1m [claude-code]
|
Holding this one out of today's landing, on scope rather than on content. When I reviewed it, the body's claim held — "Only Against current The concrete blocker is an overlap with your own #2856, which is landing today. Suggested order: let #2856 land (it is in today's merge), rebase this on the Nothing else against it from me. The spec itself I did read and liked — naming |
|
Unblocked: #2856 is on So the sequencing I suggested is available now. A rebase onto I re-probed a minute ago and the same two files still conflict, which is just Also landed in the same push, in case any of it matters to the capture harness:
One caution unrelated to the conflict: Ping me when it is rebased and I will review the implementation half properly. |
The rejected spec predates the active oracle pin and current repository gates, so the repair branch needs the pinned main tree before its scoped revision. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:gpt-6 [codex]
The rejected plan treated a red local gate as usable, invented a numerical envelope, and described oracle and provider paths that could not run. Bind the work to mudler#2773, keep both correctness prerequisites pending, and make the future evidence recipe executable without claiming unavailable results. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:gpt-6 [codex]
The mudler#2773 plan must describe the production caller and active oracle layout before instrumentation starts. Correct BF16 selector normalization, name the existing Qwen3.5 path, and record its shared-seam debt in mudler#2923. Separate SD storage from DS dump order and cite the active CPU attention test with its unchanged tolerances. Runtime acceptance remains pending. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:gpt-6-astra [codex]
The mudler#2773 capture tools also serve the ratified Qwen3 distributional gate. Keep that legacy contract and bind Qwen3.5 publication to verified model identity, ten deterministic repeats, and matched runtime artifacts. Require wheel bytes and launcher image attestation without claiming runtime proof. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:gpt-6-astra [codex]
Qwen3.5 candidates need cache-matched provenance before mudler#2773 can run the active-pin gate. Verify model, package, wheel, and launcher identities, then refuse publication unless ten repeats agree. Keep legacy Qwen3 distributional captures and the PR mudler#2856 execution-mode interface usable. The CPU fixtures pass 36 tests, and the exact PR mudler#2856 suite passes two. All 73 adverse mutations fail their focused tests and restore cleanly. The full staged preflight exits zero with five prerequisite skips. The existing shellcheck-dependent test remains unavailable. Those obligations remain PENDING. The C++ probes, the PR mudler#2856 landing and default gate, and active-pin runtime acceptance remain PENDING under BACKEND-ROCM and mudler#2773. This tooling slice claims no GPU or numerical result. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:gpt-6-astra [codex]
Strict near-tie publication accepted missing or contradictory runtime observations. Validate captured completeness and observed model and runtime identities against the independently verified current context. Publication could replace legacy inputs through destination aliases. Protect input file identities before writing any output. Keep ordinary legacy output replacement and valid installed revision prefixes. vLLM resolves sampling on a cloned request. Record constructor normalization separately and mark engine resolution unobserved. Refs mudler#2773. This repairs capture tooling in draft PR mudler#2932 on row/BACKEND-ROCM-NUMERICS-1588. GPU acceptance and the remaining characterization stay PENDING. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:gpt-6-astra [codex]
Legacy publication could overwrite the distribution metadata used to verify vLLM. Record the selected metadata files and protect their file identities. The final identity check also detects changes to those measured bytes. Distinct output paths could share an inode and replace NumPy payloads with JSON. Refuse output aliases before writing any destination. Preserve ordinary legacy replacement and the explicit default manifest path. Four permanent CLI tests cover metadata selection and fallbacks, publication aliases, metadata changes during capture, and successful replacement hashes. CPU fixtures and mutation checks cover both capture commands. Active-pin GPU acceptance and the remaining characterization stay PENDING under mudler#2773. Refs mudler#2773. This repairs draft PR mudler#2932 on row/BACKEND-ROCM-NUMERICS-1588. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:gpt-6-astra [codex]
PR mudler#2856 landed on the pinned implementation base, so its merge is no longer an external prerequisite for mudler#2773. The unchanged default gate still requires an operator rerun on that base. The operator ran the active-pin model in production mode using the reviewed capture snapshot. Ten auto/BF16 repeats agree. Record that runtime evidence without claiming C++ token acceptance, physical cache measurement, or completion of the remaining cache modes and gaps. Refs mudler#2773. The capture slice remains in draft PR mudler#2932. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:gpt-6-astra [codex]
c75173f to
81855a3
Compare
|
@localai-org-maint-bot Rebased onto real main The fresh independent review passed with no findings. All 49 capture tests and two mode tests passed, and 44 effective mutations were detected and restored exactly. The operator independently passed focused checks and the full preflight, then published that exact SHA. The updated body names the resource-dependent skips and remaining limitations. Its commit-body contract passes. This PR is ready for the implementation review you requested. The approved dedicated runtime also built and ran active vLLM pin The oracle completed deterministic ten-repeat captures for auto, BF16, and FP8 E4M3. Explicit local cache modes and the remaining state/operation characterization still belong to #2773. No tracked golden changed, and no speed result is claimed. New-head upstream CI is still running. |
Qwen3.5 captures need verified inputs before their tokens can serve as numerical evidence. The capture and near-tie commands verify artifact and runtime identities, deterministic repeats, and compatible settings before strict publication.
Row
BACKEND-ROCM. Refs #2773. That issue retains the C++ state probes, operation tests, and physical-cache characterization.Before starting
Issue #2773 owns this work. The committed spec is
.agents/specs/rocm-qwen35-08b-cpu-gfx1100-numerics.md. One PR carries its reviewed spec and implementation. The repaired spec and provenance amendment precede implementation; original spec commit7bc2546e9remains an ancestor.The integration base is real main
f98b638673b4d2edc0250eec56d229357ea38ab1, which includes #2856. The two CLI conflicts were resolved with main's production default and flushed mode narration preserved. The helper and both capture suites match prior reviewed headc75173f921cdd344e33ad06260ca181d63b198b5byte for byte. Unrelated target bytes in CI, preflight, and usage documentation were preserved.What changed
Both commands expose cache mode, production or diagnostic eager execution, seed, repetitions, revision, wheel, launcher-manifest, and provenance options. Artifact configuration selects the strict Qwen3.5 regime. Strict publication requires ten deterministic repeats, verified model files, matching resolved identities and settings, and imported package bytes matching the inspected wheel. Near-tie publication binds the captured identities and settings to the verified runtime and the actual local token prefix.
Publication protects measured inputs, including consumed distribution metadata. Distinct outputs cannot alias through direct paths, symbolic links, or hard links. Legacy Qwen3 callers retain distributional captures, manifestless near-tie inputs, overwrite behavior, and rollback. CI and preflight register both suites.
docs/USAGE.mddocuments the options and evidence limits.Evidence
Reviewed head:
81855a3f7f0ed0a0bb21edd3fcd6c2f94b4a8766. Tree:f9af146c3850cb36ddb3dc4158bc3704fa58f1c9.The implementer, fresh reviewer, and operator each passed:
The operator also passed the prior identity, publication, hard-link, rollback, and upstream sampling probes, the spec contract, and exact PR classification. The fresh review found no defect and detected 44 effective mutations. Every mutation was restored exactly; all 6,066 tracked files stayed unchanged. Exploratory mutations that were masked or reached unrelated fixture errors are retained separately and excluded from 44.
The reviewer ran
bash scripts/agent-preflight.sh --quietonce: exit 0 in 896.996 seconds. Commit style and trailers passed. This script/spec range has a derived empty C++ compilation scope.The operator independently ran the full preflight: exit 0 in 831.324 seconds. That successful gate was chained to publication of the exact reviewed head. The remote branch reports
81855a3f7f0ed0a0bb21edd3fcd6c2f94b4a8766.Separate device validation used unchanged main
f98b63867, the reviewedc75173capture snapshot, and the locally built active-pin wheel frome126687a9a828d513c01a07cd69f025f27d63280. The gfx1100 HIP build completed all 587 steps. The backend suite passed 46 cases and 84,078 assertions. The historical default model gate passed 137 assertions. Its actual C++autoIDs matched the active oracle at all 256 positions; ten teacher-forced repetitions were identical with zero gaps. An unchanged-source test binary selecting external candidate files passed 137 assertions, 16/16 strict prompts, and zero provider declines. No tracked golden changed.The oracle separately completed 16 prompts × 16 tokens × 10 repetitions for
auto,bfloat16, andfp8_e4m3. Every mode was deterministic.autoand BF16 output bytes were identical; FP8 differed at prompt 7. These initial cache-specific captures are retained separately from the later actual local-mode comparison below. Model revision:2fc06364715b967f1860aea9cf38778875588b17. Wheel SHA256:7e6efb7b3226360cd67411d62e139425d550a77407ad16391099b8cbc55b340b.Later production captures and full-prefix scoring used the active pin's 0.92 memory-utilization default. A repaired, independently reviewed private C API client linked unchanged f98 and completed 160 requests per mode. The operator and an independent reviewer verified all raw streams, completion counts, request files, prompt copies, and ten identical repetitions. Auto and BF16 each match 2,560/2,560 oracle IDs. FP8 matches 2,470/2,560, differing at zero-based prompt 15, positions 7–15 in every repeat. The first divergence is local 760 versus oracle 9175. Actual-prefix scoring gives one 125-milli-nat gap and 255 zero gaps. Under the existing 500-milli-nat rule, FP8 has 15 strict prompts and one near-tie-only prompt; strict FP8 equality still fails. Auto and BF16 have 16 strict prompts and zero gaps. Each mode's ten unrounded scoring repeats agreed in-process; only the reference log-probability hash is persisted. No permanent golden or threshold changed. This is a scoped token comparison, not completion of #2773.
Speed claims
This PR makes no speed claim.
Honest gaps
The default preflight skips ARM ISA, CPU ISA, CUDA gencode, PR classification, and Triton AOT. Exact base/head/PR classification passed separately. The other four resource checks remain PENDING. Seven optional unittest cases remain PENDING: one absent issue-snapshot case, five unset CLIP-model cases, and one unavailable CIFS case. Shellcheck is unavailable. These skips do not belong to the 49 capture tests or two mode tests.
The passing preflights used isolated Git configuration (
GIT_CONFIG_GLOBAL=/dev/null) and a NumPy-only Python path. A separatepython3 scripts/agent-ready.pyrun under the unisolated host environment exited 1: the unchanged onboarding fixture expectedmaster, but local Git configuration selectsmain. That run also skipped seven NumPy suites and the five argument-dependent checks. The publication gate executed those NumPy suites successfully. Full handoff readiness is not claimed.Installed version metadata verifies a VCS prefix, not an observed full source SHA. Image identity remains a launcher attestation. Sampling fields distinguish constructor normalization from unobserved engine request resolution. These tools do not measure physical cache storage. Explicit local selector captures are complete. Physical dtype, mode-specific providers, state probes, ported operation tests, and accepted performance axes remain owed under #2773. The local runtime accepts the requested 0.92 fraction but falls back to 256 blocks because profiling is unimplemented under #83. It reduces maximum model length to 8,192 and disables async scheduling. Resolved allocation and scheduler equivalence remain unproved. Issue #2923 owns the shared-AttnBlock wiring debt.
Separate unchanged active-pin operation tests passed 173 cases and failed five variable-length causal-convolution output comparisons at the existing BF16 bounds. No case skipped. A fresh-process rerun reproduced all five output failures. This operation prerequisite remains FAILING under #2773; it does not attribute a defect to this capture-tool PR. No tolerance changed.
The configured historical wrapper remains unchanged. The dedicated active-pin runtime is installed and has emitted the measured model tokens. New-head Windows CPU and Vulkan CI fail API-server cases 58/59/61 with
0xC0000409, tracked by existing #2403. The exact f98 scheduled baseline reproduces those three signatures on both backends with matching recorded runner image and compiler. Root cause remains unresolved and the failed gates are not waived. The completed CI workflow has 13 successful jobs, four failures, and three skips. Its address/undefined-sanitizer job fails the same three Dots3 cases on the AVX-512 misaligned BF16 load atcpu_matmul_elem_avx512.cpp:121(#2908); its thread-sanitizer job failsAnyNonZero(want)in the Gemma4 FP8 guard test (#2909), without a TSan race diagnostic. Both signatures also occur in the exact f98 scheduled baseline: address/undefined and thread. The involved CPU source and guard test are byte-identical between base and head; the sanitizer job definitions are unchanged. These existing failures remain owned and unwaived. The published-head CI run and maintainer acceptance remain separate from this scoped local review.FOLLOWING_AGENTS_PROTOCOL
Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:gpt-5.6-sol [codex]
Assisted-by: AGENT:gpt-6 [codex]
Assisted-by: AGENT:gpt-6-astra [codex]