test(msbuild): strengthen eval quality - #1118
Conversation
Expand MSBuild eval coverage with realistic fixtures, replayable references, and a changed-suite quality ratchet. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: eb11c85b-4046-4dbc-94b3-3bdbd4839a0d
Clear NuGet fallback folders so the intentional NU1101 case cannot resolve from machine state. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: eb11c85b-4046-4dbc-94b3-3bdbd4839a0d
|
@JanKrivanek @YuliiaKovalova, I would appreciate your thoughts and feedback on this draft, especially on:
For repository-size context, this PR adds 290 MSBuild artifact files totaling 176,115 bytes (172 KiB) uncompressed, or about 145,041 bytes (142 KiB) as a standalone ZIP: 83 KiB of fixtures/support files, 56 KiB of ATIF JSON, and 33 KiB of golden patches. The PR description now has a row-by-row review map for every changed skill eval and the TDD/Vally-guided design principles used to create them. |
|
/evaluate 0d63a77 |
|
❌ Evaluation did not complete successfully (the evaluate job reported 35 partial result file(s) were preserved for diagnosis but were not consolidated because the full matrix did not complete. |
|
/evaluate 0d63a77 |
📊 Skill Evaluation Results36 model/skill results across 18 skills and 2 models — ✅ 1 improved, ➖ 12 not proven improved, Measurement identity: evaluated commit Measurement health: 36 expected / 36 observed / 36 written; 0 missing, 0 unexpected, 0 invalid; 0 recovered comparison error slots and 0 unresolved comparison error slots. Objective completion gate: not enabled. Aggregate completion transitions are telemetry only, so this report does not claim that zero objective regressions were proven. A result passes only when preference-eligible distinct-stimulus votes have aggregate net win of at least 20% and an exact one-sided sign-test result of
ℹ️ How to read this report
⛔ Activation contract failed — binlog-failure-analysis (claude-sonnet-4.6)Why: Net win +33.3% (3W/2T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +35.0% across 8 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill) Next action: Narrow skill routing so the listed off-target scenarios stay dormant. State: Gate evidence: n=6; 3W/2T/1L; d=4; p=0.312; net +33.3%; 2 dormancy excluded Warnings: Dormancy contract: 1 unexpected activation(s) Overfit: Moderate (score 0.26) Repeated-run reliability (not used by the gate): 8 paired runs (5W/2T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ⛔ Activation contract failed — binlog-failure-analysis (gpt-5.6-luna)Why: Net win +16.7% (2W/3T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +5.0% across 8 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill) Next action: Narrow skill routing so the listed off-target scenarios stay dormant. State: Gate evidence: n=6; 2W/3T/1L; d=3; p=0.500; net +16.7%; 2 dormancy excluded Warnings: Dormancy contract: 1 unexpected activation(s) Overfit: Low (score 0.14) Repeated-run reliability (not used by the gate): 8 paired runs (2W/5T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ⛔ Activation contract failed — build-parallelism (claude-sonnet-4.6)Why: Net win +66.7% (4W/2T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +17.1% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill) Next action: Narrow skill routing so the listed off-target scenarios stay dormant. State: Gate evidence: n=6; 4W/2T/0L; d=4; p=0.063; net +66.7%; 1 dormancy excluded Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 5/6; plugin 5/6 Overfit: Moderate (score 0.34) Repeated-run reliability (not used by the gate): 7 paired runs (4W/2T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ⛔ Activation contract failed — build-perf-baseline (claude-sonnet-4.6)Why: Net win +50.0% (4W/1T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +40.0% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill) Next action: Narrow skill routing so the listed off-target scenarios stay dormant. State: Gate evidence: n=6; 4W/1T/1L; d=5; p=0.188; net +50.0%; 1 dormancy excluded Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 4/6; plugin 6/6 Overfit: High (score 0.52) Repeated-run reliability (not used by the gate): 7 paired runs (5W/1T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ⛔ Activation contract failed — build-perf-baseline (gpt-5.6-luna)Why: Net win +33.3% (3W/2T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +14.3% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill) Next action: Narrow skill routing so the listed off-target scenarios stay dormant. State: Gate evidence: n=6; 3W/2T/1L; d=4; p=0.312; net +33.3%; 1 dormancy excluded Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 5/6; plugin 6/6 Overfit: Moderate (score 0.23) Repeated-run reliability (not used by the gate): 7 paired runs (3W/2T/2L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ⛔ Activation contract failed — check-bin-obj-clash (claude-sonnet-4.6)Why: Net win +33.3% (3W/2T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +25.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill) Next action: Narrow skill routing so the listed off-target scenarios stay dormant. State: Gate evidence: n=6; 3W/2T/1L; d=4; p=0.312; net +33.3%; 1 dormancy excluded Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 5/6; plugin 5/6 Overfit: Moderate (score 0.22) Repeated-run reliability (not used by the gate): 7 paired runs (4W/2T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ⛔ Activation contract failed — check-bin-obj-clash (gpt-5.6-luna)Why: Net win -33.3% (0W/4T/2L over 6 preference-eligible stimulus vote(s), sign test p=0.250), mean preference -17.1% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill) Next action: Narrow skill routing so the listed off-target scenarios stay dormant. State: Gate evidence: n=6; 0W/4T/2L; d=2; p=0.250; net -33.3%; 1 dormancy excluded Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 6/6; plugin 5/6 Overfit: Low (score 0.11) Repeated-run reliability (not used by the gate): 7 paired runs (0W/4T/3L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ⛔ Activation contract failed — directory-build-organization (gpt-5.6-luna)Why: Net win +16.7% (1W/5T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +5.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill) Next action: Narrow skill routing so the listed off-target scenarios stay dormant. State: Gate evidence: n=6; 1W/5T/0L; d=1; p=0.500; net +16.7%; 1 dormancy excluded Warnings: Dormancy contract: 1 unexpected activation(s) Overfit: Low (score 0.07) Repeated-run reliability (not used by the gate): 7 paired runs (1W/6T/0L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ⛔ Activation contract failed — eval-performance (claude-sonnet-4.6)Why: Net win +50.0% (3W/3T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.125), mean preference +32.5% across 8 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — activation contract failed (2 explicit dormancy scenario(s) activated the isolated target skill) Next action: Narrow skill routing so the listed off-target scenarios stay dormant. State: Gate evidence: n=6; 3W/3T/0L; d=3; p=0.125; net +50.0%; 2 dormancy excluded Warnings: Dormancy contract: 2 unexpected activation(s); Activation: isolated 4/6; plugin 5/6 Overfit: Moderate (score 0.38) Repeated-run reliability (not used by the gate): 8 paired runs (5W/3T/0L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ⛔ Activation contract failed — eval-performance (gpt-5.6-luna)Why: Net win +33.3% (2W/4T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.250), mean preference +25.0% across 8 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — activation contract failed (2 explicit dormancy scenario(s) activated the isolated target skill) Next action: Narrow skill routing so the listed off-target scenarios stay dormant. State: Gate evidence: n=6; 2W/4T/0L; d=2; p=0.250; net +33.3%; 2 dormancy excluded Warnings: Dormancy contract: 2 unexpected activation(s); Activation: isolated 6/6; plugin 5/6 Overfit: Low (score 0.09) Repeated-run reliability (not used by the gate): 8 paired runs (3W/4T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ⛔ Activation contract failed — including-generated-files (claude-sonnet-4.6)Why: Net win +42.9% (4W/2T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +10.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill) Next action: Narrow skill routing so the listed off-target scenarios stay dormant. State: Gate evidence: n=7; 4W/2T/1L; d=5; p=0.188; net +42.9%; 1 dormancy excluded Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 4/7; plugin 6/7 Overfit: Moderate (score 0.41) Repeated-run reliability (not used by the gate): 8 paired runs (4W/2T/2L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ⛔ Activation contract failed — incremental-build (claude-sonnet-4.6)Why: Net win +12.5% (2W/5T/1L over 8 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +8.9% across 9 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill) Next action: Narrow skill routing so the listed off-target scenarios stay dormant. State: Gate evidence: n=8; 2W/5T/1L; d=3; p=0.500; net +12.5%; 1 dormancy excluded Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 4/8; plugin 4/8 Overfit: Moderate (score 0.36) Repeated-run reliability (not used by the gate): 9 paired runs (3W/5T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ⛔ Activation contract failed — incremental-build (gpt-5.6-luna)Why: Net win +50.0% (4W/4T/0L over 8 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +22.2% across 9 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill) Next action: Narrow skill routing so the listed off-target scenarios stay dormant. State: Gate evidence: n=8; 4W/4T/0L; d=4; p=0.063; net +50.0%; 1 dormancy excluded Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 6/8; plugin 8/8 Overfit: Low (score 0.10) Repeated-run reliability (not used by the gate): 9 paired runs (5W/4T/0L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ⛔ Activation contract failed — item-management (claude-sonnet-4.6)Why: Net win -16.7% (1W/3T/2L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -11.4% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill) Next action: Narrow skill routing so the listed off-target scenarios stay dormant. State: Gate evidence: n=6; 1W/3T/2L; d=3; p=0.500; net -16.7%; 1 dormancy excluded Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 5/6; plugin 4/6 Overfit: Moderate (score 0.38) Repeated-run reliability (not used by the gate): 7 paired runs (1W/3T/3L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ⛔ Activation contract failed — msbuild-antipatterns (claude-sonnet-4.6)Why: Net win +14.3% (1W/6T/0L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +10.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill) Next action: Narrow skill routing so the listed off-target scenarios stay dormant. State: Gate evidence: n=7; 1W/6T/0L; d=1; p=0.500; net +14.3%; 1 dormancy excluded Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 2/7; plugin 2/7 Overfit: Moderate (score 0.36) Repeated-run reliability (not used by the gate): 8 paired runs (2W/6T/0L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ⛔ Activation contract failed — msbuild-antipatterns (gpt-5.6-luna)Why: Net win -42.9% (0W/4T/3L over 7 preference-eligible stimulus vote(s), sign test p=0.125), mean preference -20.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill) Next action: Narrow skill routing so the listed off-target scenarios stay dormant. State: Gate evidence: n=7; 0W/4T/3L; d=3; p=0.125; net -42.9%; 1 dormancy excluded Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 5/7; plugin 4/7 Overfit: Low (score 0.09) Repeated-run reliability (not used by the gate): 8 paired runs (0W/4T/4L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ⛔ Activation contract failed — msbuild-server (claude-sonnet-4.6)Why: Net win +16.7% (3W/1T/2L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +20.0% across 8 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — activation contract failed (2 explicit dormancy scenario(s) activated the isolated target skill) Next action: Narrow skill routing so the listed off-target scenarios stay dormant. State: Gate evidence: n=6; 3W/1T/2L; d=5; p=0.500; net +16.7%; 2 dormancy excluded Warnings: Dormancy contract: 2 unexpected activation(s) Overfit: Moderate (score 0.47) Repeated-run reliability (not used by the gate): 8 paired runs (4W/1T/3L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ⛔ Activation contract failed — msbuild-server (gpt-5.6-luna)Why: Net win +50.0% (4W/1T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +35.0% across 8 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill) Next action: Narrow skill routing so the listed off-target scenarios stay dormant. State: Gate evidence: n=6; 4W/1T/1L; d=5; p=0.188; net +50.0%; 2 dormancy excluded Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 5/6; plugin 6/6 Overfit: Moderate (score 0.22) Repeated-run reliability (not used by the gate): 8 paired runs (5W/2T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ⛔ Activation contract failed — property-patterns (claude-sonnet-4.6)Why: Net win +50.0% (4W/1T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +31.4% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill) Next action: Narrow skill routing so the listed off-target scenarios stay dormant. State: Gate evidence: n=6; 4W/1T/1L; d=5; p=0.188; net +50.0%; 1 dormancy excluded Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 4/6; plugin 5/6 Overfit: Moderate (score 0.23) Repeated-run reliability (not used by the gate): 7 paired runs (5W/1T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ⛔ Activation contract failed — property-patterns (gpt-5.6-luna)Why: Net win +50.0% (3W/3T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.125), mean preference +17.1% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill) Next action: Narrow skill routing so the listed off-target scenarios stay dormant. State: Gate evidence: n=6; 3W/3T/0L; d=3; p=0.125; net +50.0%; 1 dormancy excluded Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 6/6; plugin 5/6 Overfit: Low (score 0.06) Repeated-run reliability (not used by the gate): 7 paired runs (3W/4T/0L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ⛔ Activation contract failed — resolve-project-references (claude-sonnet-4.6)Why: Net win +60.0% (4W/0T/1L over 5 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +27.5% across 8 paired run(s), 3 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill) Next action: Narrow skill routing so the listed off-target scenarios stay dormant. State: Gate evidence: n=5; 4W/0T/1L; d=5; p=0.188; net +60.0%; 3 dormancy excluded Warnings: Dormancy contract: 1 unexpected activation(s) Overfit: Moderate (score 0.32) Repeated-run reliability (not used by the gate): 8 paired runs (6W/0T/2L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ⛔ Activation contract failed — resolve-project-references (gpt-5.6-luna)Why: Net win +40.0% (2W/3T/0L over 5 preference-eligible stimulus vote(s), sign test p=0.250), mean preference +27.5% across 8 paired run(s), 3 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill) Next action: Narrow skill routing so the listed off-target scenarios stay dormant. State: Gate evidence: n=5; 2W/3T/0L; d=2; p=0.250; net +40.0%; 3 dormancy excluded Warnings: Dormancy contract: 1 unexpected activation(s) Overfit: Low (score 0.14) Repeated-run reliability (not used by the gate): 8 paired runs (4W/4T/0L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ⛔ Activation contract failed — target-authoring (claude-sonnet-4.6)Why: Net win +33.3% (3W/2T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +20.0% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill) Next action: Narrow skill routing so the listed off-target scenarios stay dormant. State: Gate evidence: n=6; 3W/2T/1L; d=4; p=0.312; net +33.3%; 1 dormancy excluded Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 6/6; plugin 4/6 Overfit: Moderate (score 0.29) Repeated-run reliability (not used by the gate): 7 paired runs (3W/3T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — binlog-generation (claude-sonnet-4.6)Why: Net win +28.6% (3W/3T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +20.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 3 of 7 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=7; 3W/3T/1L; d=4; p=0.312; net +28.6%; 1 dormancy excluded Warnings: Activation: isolated 5/7; plugin 5/7 Overfit: High (score 0.58) Repeated-run reliability (not used by the gate): 8 paired runs (3W/3T/2L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — binlog-generation (gpt-5.6-luna)Why: Net win +14.3% (3W/2T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +27.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.500 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=7; 3W/2T/2L; d=5; p=0.500; net +14.3%; 1 dormancy excluded Warnings: Activation: isolated 7/7; plugin 6/7 Overfit: Moderate (score 0.21) Repeated-run reliability (not used by the gate): 8 paired runs (3W/3T/2L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — build-parallelism (gpt-5.6-luna)Why: Net win +0.0% (3W/0T/3L over 6 preference-eligible stimulus vote(s), sign test p=0.656), mean preference -5.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=6; 3W/0T/3L; d=6; p=0.656; net +0.0%; 1 dormancy excluded Overfit: Moderate (score 0.22) Repeated-run reliability (not used by the gate): 7 paired runs (3W/0T/4L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — build-perf-diagnostics (claude-sonnet-4.6)Why: Net win -14.3% (1W/4T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -17.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=7; 1W/4T/2L; d=3; p=0.500; net -14.3%; 1 dormancy excluded Overfit: Moderate (score 0.39) Repeated-run reliability (not used by the gate): 8 paired runs (1W/4T/3L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. Routine passing details for 1 result are in Full Results. Details for 8 results were omitted to keep this comment under GitHub's limit. Open the workflow summary or Full Results for the complete breakdown. 🔍 Full Results - all metrics and investigation details
▶ Sessions Visualisation -- interactive replay of all evaluation sessions |
JanKrivanek
left a comment
There was a problem hiding this comment.
Do we want to split the infra changes and msbuild eval changes? Both are quite loaded by themselves
Restore dormancy constraint protection, correct in-scope activation expectations, and strengthen MSBuild skill routing from CI evidence. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 2d0a0a20-fd83-4343-8434-4763f8120503
|
❌ Evaluation did not complete successfully (the evaluate job reported 36 partial result file(s) were preserved for diagnosis but were not consolidated because the full matrix did not complete. |
Reclassify in-scope suitability checks, sharpen true routing boundaries, and isolate binlog fixture builds from shared build-server state. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 2d0a0a20-fd83-4343-8434-4763f8120503
📊 Skill Evaluation Results36 model/skill results across 18 skills and 2 models — ✅ 3 improved, ➖ 27 not proven improved, Measurement identity: evaluated commit Measurement health: 36 expected / 36 observed / 36 written; 0 missing, 0 unexpected, 0 invalid; 0 recovered comparison error slots and 0 unresolved comparison error slots. Objective completion gate: not enabled. Aggregate completion transitions are telemetry only, so this report does not claim that zero objective regressions were proven. A result passes only when preference-eligible distinct-stimulus votes have aggregate net win of at least 20% and an exact one-sided sign-test result of
ℹ️ How to read this report
⛔ Activation contract failed — build-parallelism (claude-sonnet-4.6)Why: Net win +50.0% (3W/3T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.125), mean preference +22.9% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill) Next action: Narrow skill routing so the listed off-target scenarios stay dormant. State: Gate evidence: n=6; 3W/3T/0L; d=3; p=0.125; net +50.0%; 1 dormancy excluded Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 3/6; plugin 5/6 Overfit: Moderate (score 0.43) Repeated-run reliability (not used by the gate): 7 paired runs (4W/3T/0L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ⛔ Activation contract failed — build-perf-baseline (claude-sonnet-4.6)Why: Net win +33.3% (4W/0T/2L over 6 preference-eligible stimulus vote(s), sign test p=0.344), mean preference +5.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill) Next action: Narrow skill routing so the listed off-target scenarios stay dormant. State: Gate evidence: n=6; 4W/0T/2L; d=6; p=0.344; net +33.3%; 1 dormancy excluded Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 5/6; plugin 5/6 Overfit: Moderate (score 0.43) Repeated-run reliability (not used by the gate): 7 paired runs (4W/0T/3L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ⛔ Activation contract failed — directory-build-organization (gpt-5.6-luna)Why: Net win +33.3% (3W/2T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +11.4% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill) Next action: Narrow skill routing so the listed off-target scenarios stay dormant. State: Gate evidence: n=6; 3W/2T/1L; d=4; p=0.312; net +33.3%; 1 dormancy excluded Warnings: Dormancy contract: 1 unexpected activation(s) Overfit: Low (score 0.10) Repeated-run reliability (not used by the gate): 7 paired runs (3W/3T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ⛔ Activation contract failed — including-generated-files (claude-sonnet-4.6)Why: Net win -14.3% (3W/0T/4L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -12.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill) Next action: Narrow skill routing so the listed off-target scenarios stay dormant. State: Gate evidence: n=7; 3W/0T/4L; d=7; p=0.500; net -14.3%; 1 dormancy excluded Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 5/7; plugin 4/7 Overfit: Moderate (score 0.37) Repeated-run reliability (not used by the gate): 8 paired runs (3W/1T/4L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ⛔ Activation contract failed — item-management (claude-sonnet-4.6)Why: Net win +66.7% (4W/2T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +37.1% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill) Next action: Narrow skill routing so the listed off-target scenarios stay dormant. State: Gate evidence: n=6; 4W/2T/0L; d=4; p=0.063; net +66.7%; 1 dormancy excluded Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 6/6; plugin 3/6 Overfit: Moderate (score 0.34) Repeated-run reliability (not used by the gate): 7 paired runs (5W/2T/0L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ⛔ Activation contract failed — property-patterns (claude-sonnet-4.6)Why: Net win +33.3% (3W/2T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +14.3% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill) Next action: Narrow skill routing so the listed off-target scenarios stay dormant. State: Gate evidence: n=6; 3W/2T/1L; d=4; p=0.312; net +33.3%; 1 dormancy excluded Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 5/6; plugin 4/6 Overfit: Moderate (score 0.26) Repeated-run reliability (not used by the gate): 7 paired runs (3W/2T/2L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — binlog-failure-analysis (claude-sonnet-4.6)Why: Net win +57.1% (5W/1T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.109), mean preference +42.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.109 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=7; 5W/1T/1L; d=6; p=0.109; net +57.1%; 1 dormancy excluded Overfit: Moderate (score 0.24) Repeated-run reliability (not used by the gate): 8 paired runs (5W/2T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — binlog-failure-analysis (gpt-5.6-luna)Why: Net win +28.6% (4W/1T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.344), mean preference +2.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.344 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=7; 4W/1T/2L; d=6; p=0.344; net +28.6%; 1 dormancy excluded Overfit: Low (score 0.10) Repeated-run reliability (not used by the gate): 8 paired runs (4W/2T/2L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — binlog-generation (gpt-5.6-luna)Why: Net win +57.1% (4W/3T/0L over 7 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +37.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 3 of 7 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=7; 4W/3T/0L; d=4; p=0.063; net +57.1%; 1 dormancy excluded Overfit: Moderate (score 0.28) Repeated-run reliability (not used by the gate): 8 paired runs (4W/3T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — build-parallelism (gpt-5.6-luna)Why: Net win -50.0% (0W/3T/3L over 6 preference-eligible stimulus vote(s), sign test p=0.125), mean preference -25.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=6; 0W/3T/3L; d=3; p=0.125; net -50.0%; 1 dormancy excluded Warnings: Activation: isolated 5/6; plugin 6/6 Overfit: Low (score 0.11) Repeated-run reliability (not used by the gate): 7 paired runs (0W/4T/3L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — build-perf-baseline (gpt-5.6-luna)Why: Net win +33.3% (4W/0T/2L over 6 preference-eligible stimulus vote(s), sign test p=0.344), mean preference +25.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.344 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=6; 4W/0T/2L; d=6; p=0.344; net +33.3%; 1 dormancy excluded Warnings: Activation: isolated 6/6; plugin 5/6 Overfit: Moderate (score 0.23) Repeated-run reliability (not used by the gate): 7 paired runs (5W/0T/2L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — build-perf-diagnostics (claude-sonnet-4.6)Why: Net win +0.0% (3W/1T/3L over 7 preference-eligible stimulus vote(s), sign test p=0.656), mean preference -20.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=7; 3W/1T/3L; d=6; p=0.656; net +0.0%; 1 dormancy excluded Warnings: Activation: isolated 6/7; plugin 7/7 Overfit: Moderate (score 0.45) Repeated-run reliability (not used by the gate): 8 paired runs (3W/1T/4L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — build-perf-diagnostics (gpt-5.6-luna)Why: Net win +28.6% (3W/3T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +10.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 3 of 7 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=7; 3W/3T/1L; d=4; p=0.312; net +28.6%; 1 dormancy excluded Overfit: Low (score 0.14) Repeated-run reliability (not used by the gate): 8 paired runs (3W/4T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — check-bin-obj-clash (claude-sonnet-4.6)Why: Net win +28.6% (4W/1T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.344), mean preference +20.0% across 7 paired run(s) — not credible (sign test p=0.344 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=7; 4W/1T/2L; d=6; p=0.344; net +28.6% Warnings: Activation: isolated 6/7; plugin 6/7 Overfit: Low (score 0.19) Repeated-run reliability (not used by the gate): 7 paired runs (4W/1T/2L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — check-bin-obj-clash (gpt-5.6-luna)Why: Net win -14.3% (0W/6T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -5.7% across 7 paired run(s) — no improvement Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=7; 0W/6T/1L; d=1; p=0.500; net -14.3% Overfit: Low (score 0.13) Repeated-run reliability (not used by the gate): 7 paired runs (0W/6T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — directory-build-organization (claude-sonnet-4.6)Why: Net win +33.3% (3W/2T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +5.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 2 of 6 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=6; 3W/2T/1L; d=4; p=0.312; net +33.3%; 1 dormancy excluded Warnings: Activation: isolated 5/6; plugin 5/6 Overfit: Moderate (score 0.33) Repeated-run reliability (not used by the gate): 7 paired runs (3W/2T/2L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — eval-performance (claude-sonnet-4.6)Why: Net win +62.5% (6W/1T/1L over 8 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +40.0% across 8 paired run(s) — not credible (sign test p=0.063 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=8; 6W/1T/1L; d=7; p=0.063; net +62.5% Warnings: Activation: isolated 8/8; plugin 7/8 Overfit: Moderate (score 0.27) Repeated-run reliability (not used by the gate): 8 paired runs (6W/1T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — eval-performance (gpt-5.6-luna)Why: Net win +50.0% (5W/2T/1L over 8 preference-eligible stimulus vote(s), sign test p=0.109), mean preference +35.0% across 8 paired run(s) — not credible (sign test p=0.109 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=8; 5W/2T/1L; d=6; p=0.109; net +50.0% Overfit: Low (score 0.09) Repeated-run reliability (not used by the gate): 8 paired runs (5W/2T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — extension-points (claude-sonnet-4.6)Why: Net win +14.3% (3W/2T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +5.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.500 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=7; 3W/2T/2L; d=5; p=0.500; net +14.3%; 1 dormancy excluded Warnings: Activation: isolated 4/7; plugin 5/7 Overfit: Moderate (score 0.26) Repeated-run reliability (not used by the gate): 8 paired runs (3W/3T/2L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — extension-points (gpt-5.6-luna)Why: Net win +57.1% (4W/3T/0L over 7 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +22.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 3 of 7 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=7; 4W/3T/0L; d=4; p=0.063; net +57.1%; 1 dormancy excluded Overfit: Moderate (score 0.24) Repeated-run reliability (not used by the gate): 8 paired runs (4W/3T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — including-generated-files (gpt-5.6-luna)Why: Net win +42.9% (4W/2T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +15.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.188 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=7; 4W/2T/1L; d=5; p=0.188; net +42.9%; 1 dormancy excluded Overfit: Low (score 0.17) Repeated-run reliability (not used by the gate): 8 paired runs (4W/3T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — incremental-build (claude-sonnet-4.6)Why: Net win +33.3% (3W/6T/0L over 9 preference-eligible stimulus vote(s), sign test p=0.125), mean preference +13.3% across 9 paired run(s) — not credible — 6 of 9 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=9; 3W/6T/0L; d=3; p=0.125; net +33.3% Warnings: Activation: isolated 8/9; plugin 2/9 Overfit: Moderate (score 0.33) Repeated-run reliability (not used by the gate): 9 paired runs (3W/6T/0L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — incremental-build (gpt-5.6-luna)Why: Net win +33.3% (4W/4T/1L over 9 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +13.3% across 9 paired run(s) — not credible (sign test p=0.188 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=9; 4W/4T/1L; d=5; p=0.188; net +33.3% Warnings: Activation: isolated 9/9; plugin 8/9 Overfit: Low (score 0.18) Repeated-run reliability (not used by the gate): 9 paired runs (4W/4T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — item-management (gpt-5.6-luna)Why: Net win -16.7% (0W/5T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -5.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=6; 0W/5T/1L; d=1; p=0.500; net -16.7%; 1 dormancy excluded Overfit: Low (score 0.11) Repeated-run reliability (not used by the gate): 7 paired runs (0W/6T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — msbuild-antipatterns (claude-sonnet-4.6)Why: Net win -14.3% (1W/4T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -17.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=7; 1W/4T/2L; d=3; p=0.500; net -14.3%; 1 dormancy excluded Warnings: Activation: isolated 3/7; plugin 2/7 Overfit: Moderate (score 0.29) Repeated-run reliability (not used by the gate): 8 paired runs (1W/4T/3L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — msbuild-antipatterns (gpt-5.6-luna)Why: Net win -14.3% (2W/2T/3L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -2.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=7; 2W/2T/3L; d=5; p=0.500; net -14.3%; 1 dormancy excluded Warnings: Activation: isolated 6/7; plugin 4/7 Overfit: Low (score 0.09) Repeated-run reliability (not used by the gate): 8 paired runs (2W/2T/4L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — msbuild-modernization (claude-sonnet-4.6)Why: Net win -16.7% (1W/3T/2L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +2.9% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=6; 1W/3T/2L; d=3; p=0.500; net -16.7%; 1 dormancy excluded Warnings: Activation: isolated 6/6; plugin 5/6 Overfit: Moderate (score 0.32) Repeated-run reliability (not used by the gate): 7 paired runs (1W/4T/2L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — msbuild-modernization (gpt-5.6-luna)Why: Net win +16.7% (1W/5T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +0.0% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 5 of 6 preference-eligible stimulus vote(s) tied, leaving only 1 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=6; 1W/5T/0L; d=1; p=0.500; net +16.7%; 1 dormancy excluded Overfit: Low (score 0.05) Repeated-run reliability (not used by the gate): 7 paired runs (1W/5T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. Details for 8 results were omitted to keep this comment under GitHub's limit. Open the workflow summary or Full Results for the complete breakdown. 🔍 Full Results - all metrics and investigation details
|
Put required input constraints before overlapping performance, generated-file, item, and property vocabulary identified by the second CI evaluation. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 2d0a0a20-fd83-4343-8434-4763f8120503
|
❌ Evaluation did not complete successfully (the evaluate job reported 36 partial result file(s) were preserved for diagnosis but were not consolidated because the full matrix did not complete. |
Preserve the off-target contracts with leading exclusion rules and retry transient binlog setup failures that otherwise drop one comparison arm. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 2d0a0a20-fd83-4343-8434-4763f8120503
|
❌ Evaluation did not complete successfully (the evaluate job reported 36 partial result file(s) were preserved for diagnosis but were not consolidated because the full matrix did not complete. |
Describe the positive applicability test for the three remaining Sonnet dormancy misses without leading with overlapping off-target vocabulary. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 2d0a0a20-fd83-4343-8434-4763f8120503
📊 Skill Evaluation Results36 model/skill results across 18 skills and 2 models — ✅ 6 improved, ➖ 27 not proven improved, Measurement identity: evaluated commit Measurement health: 36 expected / 36 observed / 36 written; 0 missing, 0 unexpected, 0 invalid; 0 recovered comparison error slots and 0 unresolved comparison error slots. Objective completion gate: not enabled. Aggregate completion transitions are telemetry only, so this report does not claim that zero objective regressions were proven. A result passes only when preference-eligible distinct-stimulus votes have aggregate net win of at least 20% and an exact one-sided sign-test result of
ℹ️ How to read this report
⛔ Activation contract failed — build-parallelism (claude-sonnet-4.6)Why: Net win +16.7% (2W/3T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +20.0% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill) Next action: Narrow skill routing so the listed off-target scenarios stay dormant. State: Gate evidence: n=6; 2W/3T/1L; d=3; p=0.500; net +16.7%; 1 dormancy excluded Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 6/6; plugin 5/6 Overfit: Moderate (score 0.32) Repeated-run reliability (not used by the gate): 7 paired runs (3W/3T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ⛔ Activation contract failed — including-generated-files (claude-sonnet-4.6)Why: Net win -14.3% (1W/4T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -17.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill) Next action: Narrow skill routing so the listed off-target scenarios stay dormant. State: Gate evidence: n=7; 1W/4T/2L; d=3; p=0.500; net -14.3%; 1 dormancy excluded Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 5/7; plugin 6/7 Overfit: Moderate (score 0.46) Repeated-run reliability (not used by the gate): 8 paired runs (1W/4T/3L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ⛔ Activation contract failed — property-patterns (claude-sonnet-4.6)Why: Net win +66.7% (4W/2T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +45.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill) Next action: Narrow skill routing so the listed off-target scenarios stay dormant. State: Gate evidence: n=6; 4W/2T/0L; d=4; p=0.063; net +66.7%; 1 dormancy excluded Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 5/6; plugin 4/6 Overfit: Moderate (score 0.24) Repeated-run reliability (not used by the gate): 7 paired runs (5W/2T/0L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — build-parallelism (gpt-5.6-luna)Why: Net win +16.7% (2W/3T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +14.3% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 3 of 6 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=6; 2W/3T/1L; d=3; p=0.500; net +16.7%; 1 dormancy excluded Overfit: Low (score 0.11) Repeated-run reliability (not used by the gate): 7 paired runs (2W/4T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — build-perf-baseline (claude-sonnet-4.6)Why: Net win -33.3% (2W/0T/4L over 6 preference-eligible stimulus vote(s), sign test p=0.344), mean preference -25.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=6; 2W/0T/4L; d=6; p=0.344; net -33.3%; 1 dormancy excluded Warnings: Activation: isolated 5/6; plugin 5/6 Overfit: Moderate (score 0.35) Repeated-run reliability (not used by the gate): 7 paired runs (2W/0T/5L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — build-perf-baseline (gpt-5.6-luna)Why: Net win +50.0% (4W/1T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +31.4% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.188 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=6; 4W/1T/1L; d=5; p=0.188; net +50.0%; 1 dormancy excluded Warnings: Activation: isolated 6/6; plugin 5/6 Overfit: Moderate (score 0.25) Repeated-run reliability (not used by the gate): 7 paired runs (5W/1T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — build-perf-diagnostics (claude-sonnet-4.6)Why: Net win +28.6% (4W/1T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.344), mean preference +5.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.344 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=7; 4W/1T/2L; d=6; p=0.344; net +28.6%; 1 dormancy excluded Overfit: High (score 0.50) Repeated-run reliability (not used by the gate): 8 paired runs (4W/1T/3L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — build-perf-diagnostics (gpt-5.6-luna)Why: Net win +57.1% (5W/1T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.109), mean preference +20.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.109 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=7; 5W/1T/1L; d=6; p=0.109; net +57.1%; 1 dormancy excluded Overfit: Low (score 0.16) Repeated-run reliability (not used by the gate): 8 paired runs (5W/2T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — check-bin-obj-clash (claude-sonnet-4.6)Why: Net win +28.6% (3W/3T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +20.0% across 7 paired run(s) — not credible — 3 of 7 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=7; 3W/3T/1L; d=4; p=0.312; net +28.6% Warnings: Activation: isolated 6/7; plugin 6/7 Overfit: Low (score 0.18) Repeated-run reliability (not used by the gate): 7 paired runs (3W/3T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — check-bin-obj-clash (gpt-5.6-luna)Why: Net win +57.1% (4W/3T/0L over 7 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +22.9% across 7 paired run(s) — not credible — 3 of 7 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=7; 4W/3T/0L; d=4; p=0.063; net +57.1% Overfit: Low (score 0.11) Repeated-run reliability (not used by the gate): 7 paired runs (4W/3T/0L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — directory-build-organization (claude-sonnet-4.6)Why: Net win +33.3% (2W/4T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.250), mean preference +20.0% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 4 of 6 preference-eligible stimulus vote(s) tied, leaving only 2 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=6; 2W/4T/0L; d=2; p=0.250; net +33.3%; 1 dormancy excluded Warnings: Activation: isolated 4/6; plugin 5/6 Overfit: Moderate (score 0.36) Repeated-run reliability (not used by the gate): 7 paired runs (2W/5T/0L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — directory-build-organization (gpt-5.6-luna)Why: Net win +16.7% (1W/5T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +5.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 5 of 6 preference-eligible stimulus vote(s) tied, leaving only 1 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=6; 1W/5T/0L; d=1; p=0.500; net +16.7%; 1 dormancy excluded Warnings: Activation: isolated 6/6; plugin 5/6 Overfit: Low (score 0.07) Repeated-run reliability (not used by the gate): 7 paired runs (1W/6T/0L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — eval-performance (claude-sonnet-4.6)Why: Net win +37.5% (5W/1T/2L over 8 preference-eligible stimulus vote(s), sign test p=0.227), mean preference +30.0% across 8 paired run(s) — not credible (sign test p=0.227 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=8; 5W/1T/2L; d=7; p=0.227; net +37.5% Warnings: Activation: isolated 8/8; plugin 7/8 Overfit: Moderate (score 0.36) Repeated-run reliability (not used by the gate): 8 paired runs (5W/1T/2L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — extension-points (claude-sonnet-4.6)Why: Net win +0.0% (2W/3T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.687), mean preference -12.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=7; 2W/3T/2L; d=4; p=0.687; net +0.0%; 1 dormancy excluded Warnings: Activation: isolated 4/7; plugin 4/7 Overfit: Moderate (score 0.41) Repeated-run reliability (not used by the gate): 8 paired runs (2W/3T/3L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — extension-points (gpt-5.6-luna)Why: Net win +28.6% (3W/3T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +22.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 3 of 7 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=7; 3W/3T/1L; d=4; p=0.312; net +28.6%; 1 dormancy excluded Overfit: Moderate (score 0.29) Repeated-run reliability (not used by the gate): 8 paired runs (4W/3T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — including-generated-files (gpt-5.6-luna)Why: Net win +0.0% (1W/5T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.750), mean preference +0.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=7; 1W/5T/1L; d=2; p=0.750; net +0.0%; 1 dormancy excluded Overfit: Low (score 0.17) Repeated-run reliability (not used by the gate): 8 paired runs (1W/6T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — incremental-build (claude-sonnet-4.6)Why: Net win +44.4% (4W/5T/0L over 9 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +17.8% across 9 paired run(s) — not credible — 5 of 9 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=9; 4W/5T/0L; d=4; p=0.063; net +44.4% Warnings: Activation: isolated 8/9; plugin 4/9 Overfit: Moderate (score 0.45) Repeated-run reliability (not used by the gate): 9 paired runs (4W/5T/0L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — incremental-build (gpt-5.6-luna)Why: Net win +22.2% (3W/5T/1L over 9 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +15.6% across 9 paired run(s) — not credible — 5 of 9 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=9; 3W/5T/1L; d=4; p=0.312; net +22.2% Warnings: Activation: isolated 9/9; plugin 7/9 Overfit: Low (score 0.14) Repeated-run reliability (not used by the gate): 9 paired runs (3W/5T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — item-management (claude-sonnet-4.6)Why: Net win +0.0% (1W/4T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.750), mean preference -8.6% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=6; 1W/4T/1L; d=2; p=0.750; net +0.0%; 1 dormancy excluded Warnings: Activation: isolated 4/6; plugin 3/6 Overfit: Moderate (score 0.32) Repeated-run reliability (not used by the gate): 7 paired runs (1W/5T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — item-management (gpt-5.6-luna)Why: Net win +16.7% (2W/3T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +5.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 3 of 6 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=6; 2W/3T/1L; d=3; p=0.500; net +16.7%; 1 dormancy excluded Overfit: Low (score 0.11) Repeated-run reliability (not used by the gate): 7 paired runs (2W/4T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — msbuild-antipatterns (claude-sonnet-4.6)Why: Net win +0.0% (1W/5T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.750), mean preference +0.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=7; 1W/5T/1L; d=2; p=0.750; net +0.0%; 1 dormancy excluded Warnings: Activation: isolated 2/7; plugin 2/7 Overfit: Moderate (score 0.32) Repeated-run reliability (not used by the gate): 8 paired runs (1W/6T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — msbuild-antipatterns (gpt-5.6-luna)Why: Net win -14.3% (0W/6T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -10.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=7; 0W/6T/1L; d=1; p=0.500; net -14.3%; 1 dormancy excluded Warnings: Activation: isolated 6/7; plugin 5/7 Overfit: Low (score 0.08) Repeated-run reliability (not used by the gate): 8 paired runs (0W/6T/2L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — msbuild-modernization (claude-sonnet-4.6)Why: Net win +0.0% (1W/4T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.750), mean preference +22.9% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=6; 1W/4T/1L; d=2; p=0.750; net +0.0%; 1 dormancy excluded Overfit: Moderate (score 0.29) Repeated-run reliability (not used by the gate): 7 paired runs (2W/4T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — msbuild-modernization (gpt-5.6-luna)Why: Net win +33.3% (2W/4T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.250), mean preference +11.4% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 4 of 6 preference-eligible stimulus vote(s) tied, leaving only 2 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=6; 2W/4T/0L; d=2; p=0.250; net +33.3%; 1 dormancy excluded Overfit: Low (score 0.06) Repeated-run reliability (not used by the gate): 7 paired runs (2W/5T/0L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — msbuild-server (gpt-5.6-luna)Why: Net win +62.5% (6W/1T/1L over 8 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +40.0% across 8 paired run(s) — not credible (sign test p=0.063 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=8; 6W/1T/1L; d=7; p=0.063; net +62.5% Overfit: Low (score 0.19) Repeated-run reliability (not used by the gate): 8 paired runs (6W/1T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — property-patterns (gpt-5.6-luna)Why: Net win +33.3% (3W/2T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +11.4% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 2 of 6 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=6; 3W/2T/1L; d=4; p=0.312; net +33.3%; 1 dormancy excluded Overfit: Low (score 0.08) Repeated-run reliability (not used by the gate): 7 paired runs (3W/3T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — resolve-project-references (claude-sonnet-4.6)Why: Net win +66.7% (5W/0T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.109), mean preference +37.5% across 8 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.109 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=6; 5W/0T/1L; d=6; p=0.109; net +66.7%; 2 dormancy excluded Overfit: Moderate (score 0.36) Repeated-run reliability (not used by the gate): 8 paired runs (7W/0T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — resolve-project-references (gpt-5.6-luna)Why: Net win +16.7% (2W/3T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +17.5% across 8 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — not credible — 3 of 6 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=6; 2W/3T/1L; d=3; p=0.500; net +16.7%; 2 dormancy excluded Warnings: Activation: isolated 4/6; plugin 6/6 Overfit: Low (score 0.15) Repeated-run reliability (not used by the gate): 8 paired runs (3W/4T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. Routine passing details for 3 results are in Full Results. Details for 5 results were omitted to keep this comment under GitHub's limit. Open the workflow summary or Full Results for the complete breakdown. 🔍 Full Results - all metrics and investigation details
|
Remove off-target vocabulary from the three descriptions that Sonnet selected before any workspace inspection, while retaining full boundary guidance in the skill bodies. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 2d0a0a20-fd83-4343-8434-4763f8120503
📊 Skill Evaluation Results36 model/skill results across 18 skills and 2 models — ✅ 5 improved, ➖ 26 not proven improved, Measurement identity: evaluated commit Measurement health: 36 expected / 36 observed / 36 written; 0 missing, 0 unexpected, 0 invalid; 0 recovered comparison error slots and 0 unresolved comparison error slots. Objective completion gate: not enabled. Aggregate completion transitions are telemetry only, so this report does not claim that zero objective regressions were proven. A result passes only when preference-eligible distinct-stimulus votes have aggregate net win of at least 20% and an exact one-sided sign-test result of
ℹ️ How to read this report
⛔ Activation contract failed — build-parallelism (claude-sonnet-4.6)Why: Net win +16.7% (3W/1T/2L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +8.6% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill) Next action: Narrow skill routing so the listed off-target scenarios stay dormant. State: Gate evidence: n=6; 3W/1T/2L; d=5; p=0.500; net +16.7%; 1 dormancy excluded Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 6/6; plugin 5/6 Overfit: Low (score 0.17) Repeated-run reliability (not used by the gate): 7 paired runs (3W/1T/3L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ⛔ Activation contract failed — including-generated-files (claude-sonnet-4.6)Why: Net win -14.3% (2W/2T/3L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -10.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill) Next action: Narrow skill routing so the listed off-target scenarios stay dormant. State: Gate evidence: n=7; 2W/2T/3L; d=5; p=0.500; net -14.3%; 1 dormancy excluded Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 5/7; plugin 3/7 Overfit: Moderate (score 0.37) Repeated-run reliability (not used by the gate): 8 paired runs (2W/2T/4L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ⛔ Activation contract failed — item-management (claude-sonnet-4.6)Why: Net win +16.7% (2W/3T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +5.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill) Next action: Narrow skill routing so the listed off-target scenarios stay dormant. State: Gate evidence: n=6; 2W/3T/1L; d=3; p=0.500; net +16.7%; 1 dormancy excluded Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 4/6; plugin 2/6 Overfit: Moderate (score 0.33) Repeated-run reliability (not used by the gate): 7 paired runs (2W/4T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ⛔ Activation contract failed — property-patterns (claude-sonnet-4.6)Why: Net win +0.0% (1W/4T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.750), mean preference +0.0% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill) Next action: Narrow skill routing so the listed off-target scenarios stay dormant. State: Gate evidence: n=6; 1W/4T/1L; d=2; p=0.750; net +0.0%; 1 dormancy excluded Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 5/6; plugin 4/6 Overfit: Moderate (score 0.27) Repeated-run reliability (not used by the gate): 7 paired runs (1W/5T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ⛔ Activation contract failed — property-patterns (gpt-5.6-luna)Why: Net win +33.3% (3W/2T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +14.3% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill) Next action: Narrow skill routing so the listed off-target scenarios stay dormant. State: Gate evidence: n=6; 3W/2T/1L; d=4; p=0.312; net +33.3%; 1 dormancy excluded Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 6/6; plugin 5/6 Overfit: Low (score 0.05) Repeated-run reliability (not used by the gate): 7 paired runs (3W/2T/2L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — binlog-failure-analysis (claude-sonnet-4.6)Why: Net win +14.3% (3W/2T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +12.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.500 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=7; 3W/2T/2L; d=5; p=0.500; net +14.3%; 1 dormancy excluded Overfit: Moderate (score 0.24) Repeated-run reliability (not used by the gate): 8 paired runs (3W/3T/2L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — binlog-generation (gpt-5.6-luna)Why: Net win +57.1% (5W/1T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.109), mean preference +55.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.109 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=7; 5W/1T/1L; d=6; p=0.109; net +57.1%; 1 dormancy excluded Overfit: Moderate (score 0.25) Repeated-run reliability (not used by the gate): 8 paired runs (6W/1T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — build-parallelism (gpt-5.6-luna)Why: Net win +33.3% (3W/2T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +11.4% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 2 of 6 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=6; 3W/2T/1L; d=4; p=0.312; net +33.3%; 1 dormancy excluded Overfit: Low (score 0.14) Repeated-run reliability (not used by the gate): 7 paired runs (3W/3T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — build-perf-baseline (claude-sonnet-4.6)Why: Net win +50.0% (4W/1T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +2.9% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.188 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=6; 4W/1T/1L; d=5; p=0.188; net +50.0%; 1 dormancy excluded Warnings: Activation: isolated 5/6; plugin 5/6 Overfit: Moderate (score 0.38) Repeated-run reliability (not used by the gate): 7 paired runs (4W/1T/2L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — build-perf-baseline (gpt-5.6-luna)Why: Net win +33.3% (4W/0T/2L over 6 preference-eligible stimulus vote(s), sign test p=0.344), mean preference +14.3% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.344 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=6; 4W/0T/2L; d=6; p=0.344; net +33.3%; 1 dormancy excluded Overfit: Moderate (score 0.24) Repeated-run reliability (not used by the gate): 7 paired runs (4W/0T/3L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — build-perf-diagnostics (claude-sonnet-4.6)Why: Net win +0.0% (3W/1T/3L over 7 preference-eligible stimulus vote(s), sign test p=0.656), mean preference -5.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=7; 3W/1T/3L; d=6; p=0.656; net +0.0%; 1 dormancy excluded Overfit: Moderate (score 0.49) Repeated-run reliability (not used by the gate): 8 paired runs (3W/1T/4L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — build-perf-diagnostics (gpt-5.6-luna)Why: Net win +28.6% (3W/3T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +10.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 3 of 7 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=7; 3W/3T/1L; d=4; p=0.312; net +28.6%; 1 dormancy excluded Overfit: Low (score 0.17) Repeated-run reliability (not used by the gate): 8 paired runs (3W/4T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — check-bin-obj-clash (claude-sonnet-4.6)Why: Net win +28.6% (3W/3T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +20.0% across 7 paired run(s) — not credible — 3 of 7 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=7; 3W/3T/1L; d=4; p=0.312; net +28.6% Warnings: Activation: isolated 6/7; plugin 5/7 Overfit: Moderate (score 0.20) Repeated-run reliability (not used by the gate): 7 paired runs (3W/3T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — check-bin-obj-clash (gpt-5.6-luna)Why: Net win -14.3% (1W/4T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -5.7% across 7 paired run(s) — no improvement Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=7; 1W/4T/2L; d=3; p=0.500; net -14.3% Overfit: Low (score 0.10) Repeated-run reliability (not used by the gate): 7 paired runs (1W/4T/2L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — directory-build-organization (claude-sonnet-4.6)Why: Net win +16.7% (2W/3T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +14.3% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 3 of 6 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=6; 2W/3T/1L; d=3; p=0.500; net +16.7%; 1 dormancy excluded Warnings: Activation: isolated 4/6; plugin 4/6 Overfit: Moderate (score 0.23) Repeated-run reliability (not used by the gate): 7 paired runs (2W/4T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — directory-build-organization (gpt-5.6-luna)Why: Net win +50.0% (3W/3T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.125), mean preference +17.1% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 3 of 6 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=6; 3W/3T/0L; d=3; p=0.125; net +50.0%; 1 dormancy excluded Warnings: Activation: isolated 5/6; plugin 6/6 Overfit: Low (score 0.07) Repeated-run reliability (not used by the gate): 7 paired runs (3W/4T/0L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — eval-performance (claude-sonnet-4.6)Why: Net win +62.5% (6W/1T/1L over 8 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +25.0% across 8 paired run(s) — not credible (sign test p=0.063 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=8; 6W/1T/1L; d=7; p=0.063; net +62.5% Warnings: Activation: isolated 8/8; plugin 7/8 Overfit: Moderate (score 0.37) Repeated-run reliability (not used by the gate): 8 paired runs (6W/1T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — eval-performance (gpt-5.6-luna)Why: Net win +37.5% (3W/5T/0L over 8 preference-eligible stimulus vote(s), sign test p=0.125), mean preference +30.0% across 8 paired run(s) — not credible — 5 of 8 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=8; 3W/5T/0L; d=3; p=0.125; net +37.5% Overfit: Low (score 0.10) Repeated-run reliability (not used by the gate): 8 paired runs (3W/5T/0L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — extension-points (claude-sonnet-4.6)Why: Net win -14.3% (1W/4T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -25.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=7; 1W/4T/2L; d=3; p=0.500; net -14.3%; 1 dormancy excluded Warnings: Activation: isolated 4/7; plugin 5/7 Overfit: Moderate (score 0.39) Repeated-run reliability (not used by the gate): 8 paired runs (1W/4T/3L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — extension-points (gpt-5.6-luna)Why: Net win +42.9% (3W/4T/0L over 7 preference-eligible stimulus vote(s), sign test p=0.125), mean preference +15.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 4 of 7 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=7; 3W/4T/0L; d=3; p=0.125; net +42.9%; 1 dormancy excluded Overfit: Moderate (score 0.25) Repeated-run reliability (not used by the gate): 8 paired runs (3W/5T/0L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — including-generated-files (gpt-5.6-luna)Why: Net win +14.3% (3W/2T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +5.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.500 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=7; 3W/2T/2L; d=5; p=0.500; net +14.3%; 1 dormancy excluded Warnings: Activation: isolated 6/7; plugin 6/7 Overfit: Low (score 0.17) Repeated-run reliability (not used by the gate): 8 paired runs (3W/3T/2L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — item-management (gpt-5.6-luna)Why: Net win +16.7% (2W/3T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +5.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 3 of 6 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=6; 2W/3T/1L; d=3; p=0.500; net +16.7%; 1 dormancy excluded Overfit: Low (score 0.11) Repeated-run reliability (not used by the gate): 7 paired runs (2W/4T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — msbuild-antipatterns (claude-sonnet-4.6)Why: Net win +28.6% (2W/5T/0L over 7 preference-eligible stimulus vote(s), sign test p=0.250), mean preference +10.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 5 of 7 preference-eligible stimulus vote(s) tied, leaving only 2 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=7; 2W/5T/0L; d=2; p=0.250; net +28.6%; 1 dormancy excluded Warnings: Activation: isolated 2/7; plugin 2/7 Overfit: Moderate (score 0.35) Repeated-run reliability (not used by the gate): 8 paired runs (2W/6T/0L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — msbuild-antipatterns (gpt-5.6-luna)Why: Net win +0.0% (1W/5T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.750), mean preference -5.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=7; 1W/5T/1L; d=2; p=0.750; net +0.0%; 1 dormancy excluded Warnings: Activation: isolated 6/7; plugin 6/7 Overfit: Low (score 0.09) Repeated-run reliability (not used by the gate): 8 paired runs (1W/5T/2L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — msbuild-modernization (claude-sonnet-4.6)Why: Net win +33.3% (3W/2T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +14.3% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 2 of 6 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=6; 3W/2T/1L; d=4; p=0.312; net +33.3%; 1 dormancy excluded Overfit: Moderate (score 0.35) Repeated-run reliability (not used by the gate): 7 paired runs (3W/2T/2L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — msbuild-modernization (gpt-5.6-luna)Why: Net win +50.0% (3W/3T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.125), mean preference +22.9% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 3 of 6 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=6; 3W/3T/0L; d=3; p=0.125; net +50.0%; 1 dormancy excluded Overfit: Low (score 0.05) Repeated-run reliability (not used by the gate): 7 paired runs (4W/3T/0L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — msbuild-server (claude-sonnet-4.6)Why: Net win +62.5% (6W/1T/1L over 8 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +40.0% across 8 paired run(s) — not credible (sign test p=0.063 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=8; 6W/1T/1L; d=7; p=0.063; net +62.5% Warnings: Activation: isolated 8/8; plugin 7/8 Overfit: Moderate (score 0.50) Repeated-run reliability (not used by the gate): 8 paired runs (6W/1T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — msbuild-server (gpt-5.6-luna)Why: Net win +50.0% (4W/4T/0L over 8 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +35.0% across 8 paired run(s) — not credible — 4 of 8 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=8; 4W/4T/0L; d=4; p=0.063; net +50.0% Overfit: Low (score 0.18) Repeated-run reliability (not used by the gate): 8 paired runs (4W/4T/0L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. Routine passing details for 1 result are in Full Results. Details for 7 results were omitted to keep this comment under GitHub's limit. Open the workflow summary or Full Results for the complete breakdown. 🔍 Full Results - all metrics and investigation details
|
Remove ambiguity from the four off-target prompts while preserving expect_activation false, and restore the prior routing descriptions after positive-only metadata regressed both model families. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 2d0a0a20-fd83-4343-8434-4763f8120503
|
❌ Evaluation did not complete successfully (the evaluate job reported 36 partial result file(s) were preserved for diagnosis but were not consolidated because the full matrix did not complete. |
Retry transient warm-up failures, narrow the two remaining routing descriptions to concrete inputs, and reclassify property placement after repeated cross-family activation and quality evidence. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 2d0a0a20-fd83-4343-8434-4763f8120503
📊 Skill Evaluation Results36 model/skill results across 18 skills and 2 models — ✅ 6 improved, ➖ 27 not proven improved, Measurement identity: evaluated commit Measurement health: 36 expected / 36 observed / 36 written; 0 missing, 0 unexpected, 0 invalid; 0 recovered comparison error slots and 0 unresolved comparison error slots. Objective completion gate: not enabled. Aggregate completion transitions are telemetry only, so this report does not claim that zero objective regressions were proven. A result passes only when preference-eligible distinct-stimulus votes have aggregate net win of at least 20% and an exact one-sided sign-test result of
ℹ️ How to read this report
⛔ Activation contract failed — build-parallelism (claude-sonnet-4.6)Why: Net win -33.3% (1W/2T/3L over 6 preference-eligible stimulus vote(s), sign test p=0.312), mean preference -5.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill) Next action: Narrow skill routing so the listed off-target scenarios stay dormant. State: Gate evidence: n=6; 1W/2T/3L; d=4; p=0.312; net -33.3%; 1 dormancy excluded Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 6/6; plugin 5/6 Overfit: Moderate (score 0.35) Repeated-run reliability (not used by the gate): 7 paired runs (2W/2T/3L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ⛔ Activation contract failed — including-generated-files (claude-sonnet-4.6)Why: Net win -14.3% (1W/4T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -10.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill) Next action: Narrow skill routing so the listed off-target scenarios stay dormant. State: Gate evidence: n=7; 1W/4T/2L; d=3; p=0.500; net -14.3%; 1 dormancy excluded Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 5/7; plugin 3/7 Overfit: Moderate (score 0.40) Repeated-run reliability (not used by the gate): 8 paired runs (1W/4T/3L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ⛔ Activation contract failed — item-management (claude-sonnet-4.6)Why: Net win +33.3% (2W/4T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.250), mean preference +11.4% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill) Next action: Narrow skill routing so the listed off-target scenarios stay dormant. State: Gate evidence: n=6; 2W/4T/0L; d=2; p=0.250; net +33.3%; 1 dormancy excluded Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 5/6; plugin 2/6 Overfit: Moderate (score 0.34) Repeated-run reliability (not used by the gate): 7 paired runs (2W/5T/0L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — binlog-failure-analysis (claude-sonnet-4.6)Why: Net win +14.3% (4W/0T/3L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +12.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.500 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=7; 4W/0T/3L; d=7; p=0.500; net +14.3%; 1 dormancy excluded Overfit: Moderate (score 0.26) Repeated-run reliability (not used by the gate): 8 paired runs (4W/1T/3L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — build-parallelism (gpt-5.6-luna)Why: Net win +50.0% (3W/3T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.125), mean preference +22.9% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 3 of 6 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=6; 3W/3T/0L; d=3; p=0.125; net +50.0%; 1 dormancy excluded Overfit: Low (score 0.12) Repeated-run reliability (not used by the gate): 7 paired runs (4W/3T/0L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — build-perf-baseline (claude-sonnet-4.6)Why: Net win +66.7% (4W/2T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +25.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 2 of 6 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=6; 4W/2T/0L; d=4; p=0.063; net +66.7%; 1 dormancy excluded Warnings: Activation: isolated 5/6; plugin 5/6 Overfit: Moderate (score 0.28) Repeated-run reliability (not used by the gate): 7 paired runs (4W/2T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — build-perf-baseline (gpt-5.6-luna)Why: Net win +50.0% (4W/1T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +14.3% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.188 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=6; 4W/1T/1L; d=5; p=0.188; net +50.0%; 1 dormancy excluded Overfit: Moderate (score 0.23) Repeated-run reliability (not used by the gate): 7 paired runs (5W/1T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — build-perf-diagnostics (claude-sonnet-4.6)Why: Net win -42.9% (1W/2T/4L over 7 preference-eligible stimulus vote(s), sign test p=0.188), mean preference -17.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=7; 1W/2T/4L; d=5; p=0.188; net -42.9%; 1 dormancy excluded Warnings: Activation: isolated 7/7; plugin 6/7 Overfit: Moderate (score 0.42) Repeated-run reliability (not used by the gate): 8 paired runs (2W/2T/4L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — build-perf-diagnostics (gpt-5.6-luna)Why: Net win +28.6% (2W/5T/0L over 7 preference-eligible stimulus vote(s), sign test p=0.250), mean preference +10.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 5 of 7 preference-eligible stimulus vote(s) tied, leaving only 2 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=7; 2W/5T/0L; d=2; p=0.250; net +28.6%; 1 dormancy excluded Overfit: Low (score 0.17) Repeated-run reliability (not used by the gate): 8 paired runs (2W/6T/0L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — check-bin-obj-clash (claude-sonnet-4.6)Why: Net win +0.0% (3W/1T/3L over 7 preference-eligible stimulus vote(s), sign test p=0.656), mean preference -8.6% across 7 paired run(s) — no improvement Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=7; 3W/1T/3L; d=6; p=0.656; net +0.0% Warnings: Activation: isolated 6/7; plugin 6/7 Overfit: Low (score 0.18) Repeated-run reliability (not used by the gate): 7 paired runs (3W/1T/3L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — check-bin-obj-clash (gpt-5.6-luna)Why: Net win +0.0% (1W/5T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.750), mean preference +0.0% across 7 paired run(s) — no improvement Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=7; 1W/5T/1L; d=2; p=0.750; net +0.0% Warnings: Activation: isolated 7/7; plugin 6/7 Overfit: Low (score 0.11) Repeated-run reliability (not used by the gate): 7 paired runs (1W/5T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — directory-build-organization (claude-sonnet-4.6)Why: Net win +50.0% (3W/3T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.125), mean preference +25.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 3 of 6 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=6; 3W/3T/0L; d=3; p=0.125; net +50.0%; 1 dormancy excluded Warnings: Activation: isolated 4/6; plugin 4/6 Overfit: Moderate (score 0.32) Repeated-run reliability (not used by the gate): 7 paired runs (3W/4T/0L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — directory-build-organization (gpt-5.6-luna)Why: Net win +16.7% (1W/5T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +5.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 5 of 6 preference-eligible stimulus vote(s) tied, leaving only 1 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=6; 1W/5T/0L; d=1; p=0.500; net +16.7%; 1 dormancy excluded Overfit: Low (score 0.08) Repeated-run reliability (not used by the gate): 7 paired runs (1W/6T/0L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — eval-performance (gpt-5.6-luna)Why: Net win +25.0% (4W/2T/2L over 8 preference-eligible stimulus vote(s), sign test p=0.344), mean preference +25.0% across 8 paired run(s) — not credible (sign test p=0.344 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=8; 4W/2T/2L; d=6; p=0.344; net +25.0% Overfit: Low (score 0.09) Repeated-run reliability (not used by the gate): 8 paired runs (4W/2T/2L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — extension-points (claude-sonnet-4.6)Why: Net win +0.0% (1W/5T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.750), mean preference -7.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=7; 1W/5T/1L; d=2; p=0.750; net +0.0%; 1 dormancy excluded Warnings: Activation: isolated 4/7; plugin 4/7 Overfit: Moderate (score 0.26) Repeated-run reliability (not used by the gate): 8 paired runs (1W/6T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — extension-points (gpt-5.6-luna)Why: Net win +28.6% (3W/3T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +22.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 3 of 7 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=7; 3W/3T/1L; d=4; p=0.312; net +28.6%; 1 dormancy excluded Overfit: Moderate (score 0.26) Repeated-run reliability (not used by the gate): 8 paired runs (4W/3T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — including-generated-files (gpt-5.6-luna)Why: Net win +14.3% (1W/6T/0L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +10.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 6 of 7 preference-eligible stimulus vote(s) tied, leaving only 1 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=7; 1W/6T/0L; d=1; p=0.500; net +14.3%; 1 dormancy excluded Warnings: Activation: isolated 7/7; plugin 6/7 Overfit: Low (score 0.16) Repeated-run reliability (not used by the gate): 8 paired runs (2W/6T/0L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — incremental-build (gpt-5.6-luna)Why: Net win +33.3% (4W/4T/1L over 9 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +20.0% across 9 paired run(s) — not credible (sign test p=0.188 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=9; 4W/4T/1L; d=5; p=0.188; net +33.3% Overfit: Low (score 0.09) Repeated-run reliability (not used by the gate): 9 paired runs (4W/4T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — item-management (gpt-5.6-luna)Why: Net win +0.0% (1W/4T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.750), mean preference +2.9% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=6; 1W/4T/1L; d=2; p=0.750; net +0.0%; 1 dormancy excluded Overfit: Low (score 0.11) Repeated-run reliability (not used by the gate): 7 paired runs (1W/4T/2L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — msbuild-antipatterns (claude-sonnet-4.6)Why: Net win +28.6% (3W/3T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +5.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 3 of 7 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=7; 3W/3T/1L; d=4; p=0.312; net +28.6%; 1 dormancy excluded Warnings: Activation: isolated 2/7; plugin 2/7 Overfit: Moderate (score 0.29) Repeated-run reliability (not used by the gate): 8 paired runs (3W/3T/2L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — msbuild-antipatterns (gpt-5.6-luna)Why: Net win -14.3% (1W/4T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +7.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=7; 1W/4T/2L; d=3; p=0.500; net -14.3%; 1 dormancy excluded Warnings: Activation: isolated 6/7; plugin 4/7 Overfit: Low (score 0.09) Repeated-run reliability (not used by the gate): 8 paired runs (2W/4T/2L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — msbuild-modernization (claude-sonnet-4.6)Why: Net win +16.7% (2W/3T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +28.6% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 3 of 6 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=6; 2W/3T/1L; d=3; p=0.500; net +16.7%; 1 dormancy excluded Overfit: Moderate (score 0.29) Repeated-run reliability (not used by the gate): 7 paired runs (3W/3T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — msbuild-modernization (gpt-5.6-luna)Why: Net win +16.7% (2W/3T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +0.0% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 3 of 6 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=6; 2W/3T/1L; d=3; p=0.500; net +16.7%; 1 dormancy excluded Overfit: Low (score 0.08) Repeated-run reliability (not used by the gate): 7 paired runs (2W/3T/2L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — msbuild-server (claude-sonnet-4.6)Why: Net win +62.5% (6W/1T/1L over 8 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +55.0% across 8 paired run(s) — not credible (sign test p=0.063 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=8; 6W/1T/1L; d=7; p=0.063; net +62.5% Warnings: Activation: isolated 8/8; plugin 7/8 Overfit: High (score 0.53) Repeated-run reliability (not used by the gate): 8 paired runs (6W/1T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — property-patterns (claude-sonnet-4.6)Why: Net win +28.6% (3W/3T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +20.0% across 7 paired run(s) — not credible — 3 of 7 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=7; 3W/3T/1L; d=4; p=0.312; net +28.6% Warnings: Activation: isolated 6/7; plugin 6/7 Overfit: Moderate (score 0.37) Repeated-run reliability (not used by the gate): 7 paired runs (3W/3T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — property-patterns (gpt-5.6-luna)Why: Net win +14.3% (3W/2T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +5.7% across 7 paired run(s) — not credible (sign test p=0.500 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=7; 3W/2T/2L; d=5; p=0.500; net +14.3% Warnings: Activation: isolated 6/7; plugin 6/7 Overfit: Low (score 0.09) Repeated-run reliability (not used by the gate): 7 paired runs (3W/2T/2L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — resolve-project-references (claude-sonnet-4.6)Why: Net win +50.0% (4W/1T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +27.5% across 8 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.188 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=6; 4W/1T/1L; d=5; p=0.188; net +50.0%; 2 dormancy excluded Warnings: Activation: isolated 6/6; plugin 5/6 Overfit: Moderate (score 0.38) Repeated-run reliability (not used by the gate): 8 paired runs (5W/2T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — resolve-project-references (gpt-5.6-luna)Why: Net win -33.3% (1W/2T/3L over 6 preference-eligible stimulus vote(s), sign test p=0.312), mean preference -10.0% across 8 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — no improvement Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=6; 1W/2T/3L; d=4; p=0.312; net -33.3%; 2 dormancy excluded Warnings: Activation: isolated 4/6; plugin 6/6 Overfit: Low (score 0.15) Repeated-run reliability (not used by the gate): 8 paired runs (1W/4T/3L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. Routine passing details for 1 result are in Full Results. Details for 7 results were omitted to keep this comment under GitHub's limit. Open the workflow summary or Full Results for the complete breakdown. 🔍 Full Results - all metrics and investigation details
|
Keep the three off-target contracts while removing negative lexical matches from prompts and requiring concrete positive artifacts in routing metadata. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 2d0a0a20-fd83-4343-8434-4763f8120503
📊 Skill Evaluation Results36 model/skill results across 18 skills and 2 models — ✅ 7 improved, ➖ 25 not proven improved, Measurement identity: evaluated commit Measurement health: 36 expected / 36 observed / 36 written; 0 missing, 0 unexpected, 0 invalid; 0 recovered comparison error slots and 0 unresolved comparison error slots. Objective completion gate: not enabled. Aggregate completion transitions are telemetry only, so this report does not claim that zero objective regressions were proven. A result passes only when preference-eligible distinct-stimulus votes have aggregate net win of at least 20% and an exact one-sided sign-test result of
ℹ️ How to read this report
⛔ Activation contract failed — build-parallelism (claude-sonnet-4.6)Why: Net win +66.7% (4W/2T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +25.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill) Next action: Narrow skill routing so the listed off-target scenarios stay dormant. State: Gate evidence: n=6; 4W/2T/0L; d=4; p=0.063; net +66.7%; 1 dormancy excluded Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 6/6; plugin 5/6 Overfit: Moderate (score 0.44) Repeated-run reliability (not used by the gate): 7 paired runs (4W/2T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ⛔ Activation contract failed — including-generated-files (claude-sonnet-4.6)Why: Net win +14.3% (3W/2T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +5.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill) Next action: Narrow skill routing so the listed off-target scenarios stay dormant. State: Gate evidence: n=7; 3W/2T/2L; d=5; p=0.500; net +14.3%; 1 dormancy excluded Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 5/7; plugin 5/7 Overfit: Moderate (score 0.41) Repeated-run reliability (not used by the gate): 8 paired runs (3W/3T/2L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ⛔ Activation contract failed — item-management (claude-sonnet-4.6)Why: Net win -16.7% (1W/3T/2L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -14.3% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill) Next action: Narrow skill routing so the listed off-target scenarios stay dormant. State: Gate evidence: n=6; 1W/3T/2L; d=3; p=0.500; net -16.7%; 1 dormancy excluded Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 6/6; plugin 2/6 Overfit: Moderate (score 0.28) Repeated-run reliability (not used by the gate): 7 paired runs (1W/4T/2L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ⛔ Activation contract failed — target-authoring (claude-sonnet-4.6)Why: Net win -33.3% (1W/2T/3L over 6 preference-eligible stimulus vote(s), sign test p=0.312), mean preference -14.3% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill) Next action: Narrow skill routing so the listed off-target scenarios stay dormant. State: Gate evidence: n=6; 1W/2T/3L; d=4; p=0.312; net -33.3%; 1 dormancy excluded Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 6/6; plugin 4/6 Overfit: Moderate (score 0.38) Repeated-run reliability (not used by the gate): 7 paired runs (2W/2T/3L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — binlog-failure-analysis (claude-sonnet-4.6)Why: Net win +57.1% (5W/1T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.109), mean preference +22.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.109 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=7; 5W/1T/1L; d=6; p=0.109; net +57.1%; 1 dormancy excluded Overfit: Low (score 0.18) Repeated-run reliability (not used by the gate): 8 paired runs (5W/1T/2L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — build-parallelism (gpt-5.6-luna)Why: Net win +0.0% (2W/2T/2L over 6 preference-eligible stimulus vote(s), sign test p=0.687), mean preference +0.0% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=6; 2W/2T/2L; d=4; p=0.687; net +0.0%; 1 dormancy excluded Overfit: Low (score 0.11) Repeated-run reliability (not used by the gate): 7 paired runs (2W/3T/2L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — build-perf-baseline (claude-sonnet-4.6)Why: Net win +16.7% (3W/1T/2L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -17.1% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.500 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=6; 3W/1T/2L; d=5; p=0.500; net +16.7%; 1 dormancy excluded Warnings: Activation: isolated 5/6; plugin 5/6 Overfit: Moderate (score 0.32) Repeated-run reliability (not used by the gate): 7 paired runs (3W/1T/3L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — build-perf-baseline (gpt-5.6-luna)Why: Net win +0.0% (2W/2T/2L over 6 preference-eligible stimulus vote(s), sign test p=0.687), mean preference -5.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=6; 2W/2T/2L; d=4; p=0.687; net +0.0%; 1 dormancy excluded Warnings: Activation: isolated 5/6; plugin 6/6 Overfit: Moderate (score 0.23) Repeated-run reliability (not used by the gate): 7 paired runs (2W/2T/3L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — build-perf-diagnostics (claude-sonnet-4.6)Why: Net win +14.3% (2W/4T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +5.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 4 of 7 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=7; 2W/4T/1L; d=3; p=0.500; net +14.3%; 1 dormancy excluded Overfit: Moderate (score 0.40) Repeated-run reliability (not used by the gate): 8 paired runs (2W/5T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — build-perf-diagnostics (gpt-5.6-luna)Why: Net win +42.9% (3W/4T/0L over 7 preference-eligible stimulus vote(s), sign test p=0.125), mean preference +15.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 4 of 7 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=7; 3W/4T/0L; d=3; p=0.125; net +42.9%; 1 dormancy excluded Overfit: Low (score 0.15) Repeated-run reliability (not used by the gate): 8 paired runs (3W/5T/0L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — check-bin-obj-clash (gpt-5.6-luna)Why: Net win +14.3% (2W/4T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +5.7% across 7 paired run(s) — not credible — 4 of 7 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=7; 2W/4T/1L; d=3; p=0.500; net +14.3% Warnings: Activation: isolated 7/7; plugin 6/7 Overfit: Low (score 0.10) Repeated-run reliability (not used by the gate): 7 paired runs (2W/4T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — directory-build-organization (claude-sonnet-4.6)Why: Net win +50.0% (3W/3T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.125), mean preference +25.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 3 of 6 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=6; 3W/3T/0L; d=3; p=0.125; net +50.0%; 1 dormancy excluded Warnings: Activation: isolated 4/6; plugin 4/6 Overfit: Moderate (score 0.37) Repeated-run reliability (not used by the gate): 7 paired runs (3W/4T/0L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — directory-build-organization (gpt-5.6-luna)Why: Net win +0.0% (1W/4T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.750), mean preference +0.0% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=6; 1W/4T/1L; d=2; p=0.750; net +0.0%; 1 dormancy excluded Warnings: Activation: isolated 5/6; plugin 6/6 Overfit: Low (score 0.09) Repeated-run reliability (not used by the gate): 7 paired runs (1W/5T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — eval-performance (claude-sonnet-4.6)Why: Net win +62.5% (6W/1T/1L over 8 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +47.5% across 8 paired run(s) — not credible (sign test p=0.063 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=8; 6W/1T/1L; d=7; p=0.063; net +62.5% Warnings: Activation: isolated 8/8; plugin 7/8 Overfit: Moderate (score 0.41) Repeated-run reliability (not used by the gate): 8 paired runs (6W/1T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — eval-performance (gpt-5.6-luna)Why: Net win +25.0% (3W/4T/1L over 8 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +25.0% across 8 paired run(s) — not credible — 4 of 8 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=8; 3W/4T/1L; d=4; p=0.312; net +25.0% Overfit: Low (score 0.10) Repeated-run reliability (not used by the gate): 8 paired runs (3W/4T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — extension-points (claude-sonnet-4.6)Why: Net win +14.3% (3W/2T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +2.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.500 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=7; 3W/2T/2L; d=5; p=0.500; net +14.3%; 1 dormancy excluded Warnings: Activation: isolated 4/7; plugin 5/7 Overfit: Low (score 0.19) Repeated-run reliability (not used by the gate): 8 paired runs (4W/2T/2L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — extension-points (gpt-5.6-luna)Why: Net win +14.3% (2W/4T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +10.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 4 of 7 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=7; 2W/4T/1L; d=3; p=0.500; net +14.3%; 1 dormancy excluded Warnings: Activation: isolated 7/7; plugin 6/7 Overfit: Moderate (score 0.24) Repeated-run reliability (not used by the gate): 8 paired runs (3W/4T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — including-generated-files (gpt-5.6-luna)Why: Net win +42.9% (4W/2T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +20.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.188 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=7; 4W/2T/1L; d=5; p=0.188; net +42.9%; 1 dormancy excluded Overfit: Moderate (score 0.20) Repeated-run reliability (not used by the gate): 8 paired runs (5W/2T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — incremental-build (claude-sonnet-4.6)Why: Net win +33.3% (4W/4T/1L over 9 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +13.3% across 9 paired run(s) — not credible (sign test p=0.188 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=9; 4W/4T/1L; d=5; p=0.188; net +33.3% Warnings: Activation: isolated 7/9; plugin 5/9 Overfit: Moderate (score 0.31) Repeated-run reliability (not used by the gate): 9 paired runs (4W/4T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — incremental-build (gpt-5.6-luna)Why: Net win +44.4% (4W/5T/0L over 9 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +17.8% across 9 paired run(s) — not credible — 5 of 9 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=9; 4W/5T/0L; d=4; p=0.063; net +44.4% Overfit: Low (score 0.16) Repeated-run reliability (not used by the gate): 9 paired runs (4W/5T/0L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — item-management (gpt-5.6-luna)Why: Net win +16.7% (1W/5T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +11.4% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 5 of 6 preference-eligible stimulus vote(s) tied, leaving only 1 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=6; 1W/5T/0L; d=1; p=0.500; net +16.7%; 1 dormancy excluded Overfit: Low (score 0.11) Repeated-run reliability (not used by the gate): 7 paired runs (2W/5T/0L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — msbuild-antipatterns (claude-sonnet-4.6)Why: Net win +0.0% (1W/5T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.750), mean preference -5.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=7; 1W/5T/1L; d=2; p=0.750; net +0.0%; 1 dormancy excluded Warnings: Activation: isolated 2/7; plugin 2/7 Overfit: Moderate (score 0.24) Repeated-run reliability (not used by the gate): 8 paired runs (1W/5T/2L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — msbuild-antipatterns (gpt-5.6-luna)Why: Net win +14.3% (3W/2T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +25.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.500 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=7; 3W/2T/2L; d=5; p=0.500; net +14.3%; 1 dormancy excluded Warnings: Activation: isolated 6/7; plugin 5/7 Overfit: Low (score 0.09) Repeated-run reliability (not used by the gate): 8 paired runs (4W/2T/2L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — msbuild-modernization (claude-sonnet-4.6)Why: Net win +50.0% (3W/3T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.125), mean preference +25.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 3 of 6 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=6; 3W/3T/0L; d=3; p=0.125; net +50.0%; 1 dormancy excluded Warnings: Activation: isolated 6/6; plugin 5/6 Overfit: Low (score 0.19) Repeated-run reliability (not used by the gate): 7 paired runs (3W/4T/0L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — msbuild-modernization (gpt-5.6-luna)Why: Net win +33.3% (2W/4T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.250), mean preference +5.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 4 of 6 preference-eligible stimulus vote(s) tied, leaving only 2 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=6; 2W/4T/0L; d=2; p=0.250; net +33.3%; 1 dormancy excluded Overfit: Low (score 0.06) Repeated-run reliability (not used by the gate): 7 paired runs (2W/4T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — property-patterns (claude-sonnet-4.6)Why: Net win +57.1% (4W/3T/0L over 7 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +31.4% across 7 paired run(s) — not credible — 3 of 7 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=7; 4W/3T/0L; d=4; p=0.063; net +57.1% Warnings: Activation: isolated 7/7; plugin 6/7 Overfit: Moderate (score 0.20) Repeated-run reliability (not used by the gate): 7 paired runs (4W/3T/0L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — property-patterns (gpt-5.6-luna)Why: Net win -14.3% (0W/6T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -5.7% across 7 paired run(s) — no improvement Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=7; 0W/6T/1L; d=1; p=0.500; net -14.3% Warnings: Activation: isolated 6/7; plugin 7/7 Overfit: Low (score 0.07) Repeated-run reliability (not used by the gate): 7 paired runs (0W/6T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — resolve-project-references (gpt-5.6-luna)Why: Net win +0.0% (1W/4T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.750), mean preference +0.0% across 8 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — no improvement Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=6; 1W/4T/1L; d=2; p=0.750; net +0.0%; 2 dormancy excluded Warnings: Activation: isolated 4/6; plugin 6/6 Overfit: Low (score 0.14) Repeated-run reliability (not used by the gate): 8 paired runs (1W/6T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. Routine passing details for 2 results are in Full Results. Details for 6 results were omitted to keep this comment under GitHub's limit. Open the workflow summary or Full Results for the complete breakdown. 🔍 Full Results - all metrics and investigation details
|
Use off-target boundaries that remain meaningful when the target is the only loaded skill, while retaining anti-hijack coverage for compilation, T4, property, and compiler-performance workflows. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 2d0a0a20-fd83-4343-8434-4763f8120503
|
❌ Evaluation did not complete successfully (the evaluate job reported |
1 similar comment
|
❌ Evaluation did not complete successfully (the evaluate job reported |
|
❌ Evaluation did not complete successfully (the evaluate job reported 36 partial result file(s) were preserved for diagnosis but were not consolidated because the full matrix did not complete. |
|
❌ Evaluation did not complete successfully (the evaluate job reported 36 partial result file(s) were preserved for diagnosis but were not consolidated because the full matrix did not complete. |
|
❌ Evaluation did not complete successfully (the evaluate job reported 36 partial result file(s) were preserved for diagnosis but were not consolidated because the full matrix did not complete. |
Summary
dotnet-msbuildeval suites from 33 to 139 distinct stimuli.create-skill-testandimprove-skill-qualityauthoring guidance.Design principles
These evals use repository eval-review practices, TDD, and Vally guidance as two complementary quality bars.
runsmeasures reliability; it does not add scenario breadth.Review map: what changed in each eval
agent.msbuildNU1101diagnosis and project-organization advice. The package case is diagnosis-only and cannot resolve from a remote feed or machine fallback cache.binlog-failure-analysisbinlog-generationbuild-parallelismBuildInParallel, rejection of unsafe graph mode for run-time project discovery, solution filters, and an incremental-build routing boundary. The original solution fixture was repaired so it builds real projects.build-perf-baselineUseArtifactsOutput, deterministic cache-safe CI, already-optimized no-op behavior, and a non-MSBuild boundary.build-perf-diagnosticscheck-bin-obj-clashProjectReferencemetadata, safe shared roots, full mixed-solution repair, audit-only classification, and default-layout no-op behavior.directory-build-organizationeval-performanceTreatAsLocalProperty.extension-pointsincluding-generated-filesobjpaths, an executable full repair, and a Roslyn source-generator boundary.incremental-buildInputs/Outputs,FileWritesand Clean behavior, volatile outputs, stale-input log evidence, the limits ofOutputsalone, compiler-cache explanations, Visual Studio up-to-date behavior, cold-build restraint, and a complete executable repair.item-managementCompile Remove, already-correct no-op behavior, and a target-ordering boundary.msbuild-antipatternsExecpath bugs, clean-project restraint, and a modernization routing boundary.msbuild-modernizationmsbuild-serverproperty-patternsTargetFrameworkconditions in props, cross-platform OS detection, preservation of existing constants, no-op behavior, and a props-versus-targets boundary.resolve-project-referencestarget-authoringReturnsversusOutputs, executable anti-pattern repair, already-correct no-op behavior, and an incremental-build boundary.Shared quality tooling
check_eval_quality.pyapplies 22 fixture, reference, grader, security, and quality checks.HEAD^;--base-refsupports PR-base checks;--allaudits the repository.Repository-space impact
The PR adds 290 MSBuild artifact files totaling 176,115 bytes (172 KiB) in the working tree. A standalone ZIP of those files is about 145,041 bytes (142 KiB).
The largest per-suite additions are
binlog-failure-analysis(23 KiB),eval-performance(22 KiB),binlog-generation(14 KiB), andmsbuild-server(14 KiB). These figures measure this PR's added files, not Git pack-level deduplication.Reliability and safety
Contoso.IntentionallyMissing, an empty local feed, a project-local package cache, and cleared fallback folders.sign.exesentinel with a harmlessdotnet --versionmarker and removed answer-bearing fixture comments.plugins/dotnet-msbuildskill or agent behavior changes are included.Validation
NU1101and path-marker fixture checks passed.Scope
This draft contains only:
tests/dotnet-msbuild/**eng/eval-quality/check_eval_quality.py, its self-tests, README, and the MSBuild allowlist removals.agents/skills/create-skill-test/**.agents/skills/improve-skill-quality/**Other plugin eval repairs and unrelated working-tree changes are excluded.