Skip to content

knowledge: improve review precision from BCApps PR 10990-neg-3926331607-source-faithful-20261005 feedback - #217

Closed
Wenjie Fan (gggdttt) wants to merge 1 commit into
mainfrom
self-improvement/bcapps-10990-neg-3926331607-source-faithful-20261005-quality
Closed

Wenjie Fan (gggdttt) wants to merge 1 commit into
mainfrom
self-improvement/bcapps-10990-neg-3926331607-source-faithful-20261005-quality

Conversation

@gggdttt

@gggdttt Wenjie Fan (gggdttt) commented Oct 5, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Improves BCQuality knowledge based on maintainer thumbs-down feedback from explicit BCApps PR review runs.

Source feedback

Validation

  • Frontmatter and article structure validation completed.

Generated by the BC-ALAgentsInternal self-improvement workflow.

Offline evaluation: regression

Candidate correctness failed: unexpected or missed findings remain. Existing ignored gold comments retain their neutral scoring semantics.

Preparation attempt 1: validated.

Selection: new. New: synthetic__upgrade-contoso-demo-data-no-upgrade-wiring-01. Reused (complete payloads): ``.

  • synthetic__upgrade-contoso-demo-data-no-upgrade-wiring-01 / false_positive_guard / Add-Expense-VAT-settings-to-Contoso-demo-tool BCApps#10990 (comment): Exact pipeline-validated negative snapshot (base_commit 1e3872168af58518ac2a02e8cf5934cbc28eee32) reproducing the reviewed hunk: Codeunit.Run(Codeunit::"Create Expense VAT Rates") added inside CreateMasterData() of a codeunit implementing "Contoso Demo Data Module". The maintainer rejected the request for an Upgrade-subtype entry point because this is Preview demo data invoked only through the explicit Contoso Demo Data tool, never automatically on install or version upgrade -- no subscriber or Subtype=Install/Upgrade exists anywhere in this module for any of its demo-data procedures, so the alleged upgrade-coverage gap is not a defect introduced by this change. Limitation: this guard covers only Contoso demo-data-module seeding helpers invoked exclusively through the demo-data tool; it does not generalize to codeunits that genuinely participate in the install/upgrade pipeline, nor does it certify every Preview-labelled feature as upgrade-exempt.
    Coverage references and outcome shapes are validated mechanically. Semantic equivalence, severity calibration and recommendation quality are NOT proved by these checks or a matching F1.

Common engine code: ecf8e31759d6ddd6d78e3a0b7836b40134368009; dataset SHA256: 39B118C1CDE2A4BA879BA544288C777E80D99B14F7882870045CC993BBE830B5.
Model: gpt-5.6-luna; judge: gpt-5.3-codex. Exact D: synthetic__upgrade-contoso-demo-data-no-upgrade-wiring-01.
Dataset base: 6fe2d36319d9198f64696f027eaaef495916531d; candidate: f74ae0f8e65da48e1c6aece9233c0e9d1cbf9efd.

Arm Evaluation commit Engine pin Knowledge pin
baseline 226ccd7565c8966564abeb04599679e069e46744 0df2f2c5eaf3e51f289571c68c6d0cb55212dd09 ac249ba4c95eaef14edbe9818fcdd65561bf6a2f
candidate 78b9b0fa1864bef3f9c2dc8404e1765603bf88c0 55248fe87e4adfca5a6e276f32e95e1d7b89f587 0d728298efa0dc112bae858aab792e7c6fc43d46

Baseline

Run: https://github.com/microsoft/BC-Bench/actions/runs/37281893633; conclusion: success; wall clock: 6.7 minutes.
Raw findings and runtime pin proofs are retained in the self-improvement-feedback artifact and the linked evaluation run.

Entry Expected Generated Missed Unexpected F1 AI credits Agent seconds
synthetic__upgrade-contoso-demo-data-no-upgrade-wiring-01 0 1 0 1 0 unavailable 255.4883047

Observed AI credits: unavailable; coverage 0/1 entries. Missing entries are not extrapolated into a total; reported amounts do not attest full credit coverage.

Aggregate metric Value
total 1
expected_comment_count 0
generated_comment_count 1
matched_comment_count 0
missed_comment_count 0
incorrect_comment_count 1
precision 0
recall 1
f1 0

Candidate

Run: https://github.com/microsoft/BC-Bench/actions/runs/37282634446; conclusion: success; wall clock: 6.4 minutes.
Raw findings and runtime pin proofs are retained in the self-improvement-feedback artifact and the linked evaluation run.

Entry Expected Generated Missed Unexpected F1 AI credits Agent seconds
synthetic__upgrade-contoso-demo-data-no-upgrade-wiring-01 0 1 0 1 0 unavailable 216.3225404

Observed AI credits: unavailable; coverage 0/1 entries. Missing entries are not extrapolated into a total; reported amounts do not attest full credit coverage.

Aggregate metric Value
total 1
expected_comment_count 0
generated_comment_count 1
matched_comment_count 0
missed_comment_count 0
incorrect_comment_count 1
precision 0
recall 1
f1 0

Missing telemetry is unavailable, not zero. Evaluation-only metrics exclude candidate generation and are not the full-cycle cost.
Human review must verify source-patch fidelity, gold correctness, and target attribution. No automatic merge or branch-protection claim is made.

Cycle usage and elapsed time (running, as of 2026-10-05T08:23:34.0216801+00:00)

Orchestrator work: 1395.519 seconds. This is not the completed GitHub workflow duration.

Stage / attempt Status / outcome Elapsed seconds
preflight / 1 completed / 8.175
quality_preparation / 1 completed / 76.325
gold_preparation / 1 completed / selected_new 437.724
evaluation_preparation / 1 completed / 30.169
baseline / 1 completed / success 427.238
candidate / 1 completed / success 414.947
gate_verdict / 1 completed / regression 0.013
publication / 1 running / 0.036
receipts / 1 not_started / unavailable
quality_validation / 1 completed / 5.142
gold_attempt / 1 completed / validated 421.643
Measure Known subtotal Whole-cycle total Known / required components
ai_credit: aiCredits 114.39858 unavailable 2 / 6
premium_request: premiumRequests 73 unavailable 2 / 6
token: inputTokens 12690837 unavailable 4 / 6
token: outputTokens 126487 unavailable 4 / 6
token: totalTokens 12817324 unavailable 4 / 6

Subtotals are CLI/Bench-reported usage, not invoices. Missing values are not zero. Emitted-span coverage does not prove complete child/background billing; judge usage is not exported. Premium requests are not AI credits or currency.
Nested stage, agent and API durations are not added into cycle wall clock. Root queue/setup, telemetry export, artifact upload and later dashboard publication are excluded.
PR evidence is an as-of snapshot before publication completes. Final-for-orchestrator metrics are emitted to self-improvement-cycle.json in the root self-improvement-feedback artifact and its Actions summary when persistence succeeds; a running checkpoint is not final evidence.
Root run: https://github.com/microsoft/BC-ALAgentsInternal/actions/runs/37280881649/attempts/1

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed draft head 0d728298efa0dc112bae858aab792e7c6fc43d46. The narrower Contoso-specific scope is an improvement over the earlier general demo/preview proposal, but the regeneration rationale remains factually wrong.

At contoso-demo-data-seeding-does-not-require-upgrade-codeunit.md:16-21, the article claims modules are only invoked explicitly and users rerun/refresh the demo tool to receive every new seeding change. ContosoDemoTool does the opposite for an already-generated level: CreateDemoData skips IsModuleGenerated modules (lines 63-70), GenerateDemoData checks the same guard (181), and IsModuleGenerated uses the recorded Data Level (213-215). When all selected modules were generated it raises All the selected modules have already been generated. The tool also has assisted-new-company and build-script entry points (35-54, 248-277, 334-340), so it is not exclusively an explicit manual invocation. Verified against BCApps commit ec65cce27e0eed89f885aad0c190b8af725b29ca.

A negative rule can still say that new seeding logic for intentionally disposable/unsupported demo datasets does not by itself require a production upgrade codeunit. It must not justify that by claiming already-generated demo companies can simply rerun the module, or guarantee compatibility solely from implementing this interface. State how changed seeding applies to newly generated data and explicitly scope any absence of migration to the intended demo-data compatibility policy. The PR's own offline candidate evaluation also still reports the original unexpected finding, so it does not establish that the suppression is effective yet.

@gggdttt

Copy link
Copy Markdown
Collaborator Author

Closing this generated candidate because the source-faithful baseline did not reproduce the selected Upgrade finding, so the article change cannot be attributed to an evaluation improvement. The regression run, branch, commit, and artifacts remain preserved as evidence.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants