Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 2 additions & 2 deletions .claude-plugin/marketplace.json
Original file line number Diff line number Diff line change
Expand Up @@ -8,8 +8,8 @@
{
"name": "bcquality",
"source": "./",
"description": "Business Central AL quality knowledge base and skills, packaged as an installable plugin. Exposes read-only plan enrichment and review adapters through BCQuality's Entry protocol.",
"version": "0.3.0",
"description": "Business Central AL quality knowledge base and skills, packaged as an installable plugin. Exposes read-only plan, implementation-guidance, and review adapters through BCQuality's Entry protocol.",
"version": "0.4.0",
"skills": [
"./skills/"
]
Expand Down
6 changes: 4 additions & 2 deletions .github/scripts/validate_frontmatter.py
Original file line number Diff line number Diff line change
Expand Up @@ -46,9 +46,11 @@

STANDARD_INPUTS = {
"pr-diff", "object-list", "file-path", "folder-path", "repository", "telemetry-query",
"development-plan",
"development-plan", "implementation-diff", "decision-context", "consumed-guidance",
}
ALLOWED_OUTPUTS = {
"findings-report", "development-guidance-report", "implementation-guidance-report",
}
ALLOWED_OUTPUTS = {"findings-report", "development-guidance-report"}
VALID_SAMPLE_KINDS = {"good", "bad"}

ACTION_SKILL_SECTIONS = ["Source", "Relevance", "Worklist", "Action", "Output"]
Expand Down
31 changes: 21 additions & 10 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -29,11 +29,13 @@ copilot plugin list
The list should include `bcquality`. The plugin exposes
[`al-code-review`](skills/al-code-review/SKILL.md) and the read-only
[`al-development-plan`](skills/al-development-plan/SKILL.md) plan-enrichment
skill. Installation and skill discovery are the general pattern; reviewing an
app is one example of using it.
and [`al-implementation-guidance`](skills/al-implementation-guidance/SKILL.md)
consultation skills. Installation and skill discovery are the general pattern;
reviewing an app is one example of using it.

Plugin version `0.3.0` adds `al-development-plan`, a read-only adapter for
enriching an **existing** plan. It does not generate a plan or implement code.
Plugin version `0.4.0` adds `al-implementation-guidance`, a read-only,
just-in-time consultation over the current diff and decision context. It does
not edit code or run an implementation workflow.

The adapters are intentionally not second implementations:

Expand All @@ -47,6 +49,11 @@ standalone host skill: skills/al-development-plan/SKILL.md
-> routing contract: skills/entry.md
-> enrichment skill: microsoft/skills/development/al-development-plan.md
-> referenced constraints for the consumer's existing workflow (read-only)

standalone host skill: skills/al-implementation-guidance/SKILL.md
-> routing contract: skills/entry.md
-> consultation skill: microsoft/skills/development/al-implementation-guidance.md
-> focused constraints for the consumer's next decision (read-only)
```

Only the files under `skills/*/SKILL.md` follow the host's packaging format.
Expand Down Expand Up @@ -113,13 +120,16 @@ intentionally left to those deterministic tools rather than duplicated here.

The read-only `al-development-plan` interface selects relevant constraints
before the consumer implements its own existing plan.
The distinct `al-implementation-guidance` interface consults the same corpus
after implementation begins and selects only constraints capable of changing
the next bounded implementation or validation decision.

Repository-specific orchestrators retain planning, implementation, approvals,
tests, environment, propagation, and delivery ownership. The intended flow is
consumer analysis and normalized plan -> read-only BCQuality guidance ->
existing implementation phases -> independent final BCQuality review ->
delivery. Consumer uptake and a real runtime pilot are follow-up work, not
implemented integrations or demonstrated authoring improvements.
consumer analysis and normalized plan -> read-only plan guidance -> consumer
implementation -> explicit read-only implementation checkpoints -> consumer
edits and tests -> independent final BCQuality review -> delivery. BCQuality
does not own checkpoint state or invoke itself automatically.

`no-knowledge` means no additional applicable BCQuality constraints, with empty
`knowledge`; it does not make a plan unsafe or prevent the consumer from using
Expand Down Expand Up @@ -151,8 +161,9 @@ and apply the relevant knowledge. Both live in three layers:
| [Community](community/) | Community-owned skills and their knowledge. |
| [Custom](custom/) | Organization-specific additions and overrides in your own fork. |

Review skills emit a `findings-report`; plan enrichment emits a read-only
`development-guidance-report`. Both contracts are defined in
Review skills emit a `findings-report`; plan enrichment emits
`development-guidance-report`; implementation consultation emits
`implementation-guidance-report`. All contracts are defined in
[`skills/do.md`](skills/do.md). See
[how agents consume BCQuality](docs/agent-consumption.md) for the integration
flow.
Expand Down
120 changes: 114 additions & 6 deletions docs/agent-consumption.md
Original file line number Diff line number Diff line change
Expand Up @@ -79,7 +79,8 @@ Add scheduling, retries, and rendering only when needed, using the

When BCQuality is installed as a standalone plugin, it additionally exposes
`skills/al-code-review/SKILL.md` and
`skills/al-development-plan/SKILL.md`. These are host-format adapters, not
`skills/al-development-plan/SKILL.md`, and
`skills/al-implementation-guidance/SKILL.md`. These are host-format adapters, not
additional action skills: each creates the task context and enters the same
flow at Entry.

Expand Down Expand Up @@ -175,6 +176,7 @@ The output contracts are defined in the DO meta-skill:

- A **findings report** carries review findings, domain labels, references, confidence, and suppressions.
- A **development guidance report** carries read-only knowledge constraints and validation considerations for an existing plan.
- An **implementation guidance report** carries focused additional constraints for a current diff, phase, and next implementation or validation decision.

The orchestrator parses this **without skill-specific logic**. This is the point of the contract: orchestrators and action skills evolve independently.

Expand All @@ -183,6 +185,11 @@ selects applicable knowledge, and returns constraints without changing the
target. It does not generate a replacement plan, run tests, implement code, or
drive a review/fix loop. Implementation stays in the consuming workflow.

For implementation consultation, the skill additionally reads the current
implementation diff and explicit decision context. It does not broaden the
plan again: it returns only knowledge that can change the next bounded
implementation or validation decision.

### 7. Orchestrator integrates
The orchestrator turns findings into PR comments, build gates, or IDE diagnostics. It can feed read-only guidance into its own implementation phases, preserving all existing approvals and delivery gates.

Expand All @@ -205,11 +212,21 @@ integration boundary, not a shipped consumer integration:
3. Execute the dispatched `al-development-plan` skill. Persist the unchanged
report and provenance **after** any consumer state initialization or cleanup
that could erase them. BCQuality does not own the state directory or lifecycle.
4. Inject relevant constraints and validation considerations into the existing
Baseline, Implement, propagation (such as MiApp), and Critique phases, or
equivalents. Re-enrich on material plan or applicability changes; preserve
the relationship between plan version, guidance, and implementation attempt.
5. Run an independent final BCQuality review against the completed diff using
4. Implement in the consumer workflow. At explicit checkpoints, invoke
`al-implementation-guidance` with the existing plan, readable repository,
current implementation diff, and decision context. Useful checkpoints
include schema/data upgrade, public APIs/events/interfaces, permissions,
external effects/job queues/`HttpClient`, telemetry/privacy, UI/page
background tasks, and tests. The coding agent may also initiate a bounded
consultation. This hybrid recommendation is visible orchestration, not
automatic hidden behavior.
5. Apply returned constraints, edit, compile, test, and retry in the consumer.
Persist consumed article path, decision key, and evidence fingerprint in
consumer-owned state. Requery when the diff, affected files/symbols,
changed AL tokens, next decision, acceptance criteria, validation result, or
applicability context materially changes. Exact unchanged triples are
omitted deterministically; BCQuality stores no state.
6. Run an independent final BCQuality review against the completed diff using
the **same recorded immutable checkout** used for enrichment. Review the
actual changes, not the guidance report as proof of correctness, then apply
the consumer's ordinary delivery gates.
Expand All @@ -230,6 +247,90 @@ outcomes and recorded unknowns: seek missing context, re-enrich, escalate, or
apply their existing risk controls without relabeling the report as successful.
Do not fill gaps with generic articles simply to unlock implementation.

For implementation consultation, `no-knowledge` means no additional applicable
constraints for the exact current decision and evidence fingerprint. It does
not establish functional correctness, safe deployment, adequate tests, or
release readiness.

### Minimal implementation consultation

The consumer supplies the actual values separately from Entry's type list:

```text
BCQuality root: C:\Knowledge\BCQuality
development-plan: <existing normalized plan>
repository: C:\Repos\MyBusinessCentralApp
implementation-diff: <current patch with exact changed files>
decision-context:
phase: public-contract-checkpoint
decision: Choose a compatible shape for the new provider capability.
decision-key: provider-capability-contract
evidence-fingerprint: sha256:<consumer-computed-current-evidence>
affected-files: [src/Provider/INotificationProvider.Interface.al]
affected-symbols: [interface Notification Provider]
changed-tokens: [interface, procedure]
tests: [Compile existing and new provider implementations.]
acceptance-criteria: [Existing implementations remain compatible.]
development-plan-pin: plan:v1
implementation-evidence-pin: diff:v2
consumed-guidance: []

Invoke skills/entry.md with inputs-available:
[development-plan, repository, implementation-diff, decision-context,
consumed-guidance]. Execute only the dispatched implementation-guidance skill
and return its implementation-guidance-report unchanged.
```

An abbreviated successful report is:

```json
{
"skill": { "id": "al-implementation-guidance", "version": 1 },
"outcome": "completed",
"summary": {
"request": "Expose delivery status through the provider contract.",
"phase": "public-contract-checkpoint",
"decision": "Choose a compatible shape for the new provider capability.",
"decision-key": "provider-capability-contract",
"evidence-fingerprint": "sha256:example",
"candidates": 1,
"selected": 1,
"omitted-consumed": 0
},
"pins": {
"knowledge-checkout": "0123456789abcdef0123456789abcdef01234567",
"development-plan": "plan:v1",
"implementation-evidence": "diff:v2"
},
"context": {
"bc-version": "28",
"technologies": ["al"],
"countries": ["w1"],
"application-area": ["all"],
"affected-files": ["src/Provider/INotificationProvider.Interface.al"],
"affected-symbols": ["interface Notification Provider"],
"changed-tokens": ["interface", "procedure"],
"unknown": []
},
"knowledge": [
{
"path": "microsoft/knowledge/interfaces/extend-published-interfaces-dont-edit-them.md",
"used-for": "Choose a compatible public contract shape.",
"constraints": ["Preserve the shipped interface and use the article's compatible extension shape."],
"sample-paths": ["microsoft/knowledge/interfaces/extend-published-interfaces-dont-edit-them.good.al"]
}
],
"validation-considerations": [],
"deduplication": { "strategy": "omit-exact-consumed-match", "omitted": [] },
"suppressed": [],
"unresolved": []
}
```

The example illustrates shape and provenance, not a functional-correctness
claim. Consumers must still open the cited knowledge through the skill,
compile, test, and independently review the resulting diff.

### Pinning and pilot evidence

A configured tag or ref alone does not establish runtime pinning. Record the
Expand All @@ -248,6 +349,13 @@ final reviews, including failures and unresolved results. Credential-free
fixture preparation and scorer regressions do not demonstrate improved repair
quality, compilation, runtime success, or production integration.

A focused authoring experiment should use four matched arms: baseline without
guidance, plan guidance only, implementation guidance only, and combined plan
plus implementation guidance. Hold starting task/code, model, tools, runtime,
checkpoints, and delivery gates constant; evaluate resulting diffs, compile and
test evidence, independent final review, guidance precision, duplication, and
cost. This is a recommended design, not a reported experiment or result.

## Knowledge-backed and agent findings

BCQuality is an **additive** knowledge layer. The agent surfaces two kinds of findings, both shaped to the same DO output contract:
Expand Down
38 changes: 38 additions & 0 deletions evaluation/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -168,3 +168,41 @@ article-read traces, resulting diffs, compile/test outcomes, and independent
review evidence, including failures, no-knowledge, partial, and unresolved
results. No compile/run or authoring-quality claim follows from the
credential-free checks above.

## Read-only implementation guidance

`implementation-guidance-fixtures.json` covers a distinct just-in-time
consultation contract. Its production-shaped synthetic AL contexts exercise a
new privacy constraint after the implementation surface expands beyond the
plan, exact consumed-guidance omission, an honest no-additional-guidance
control, missing current diff/decision context, and a public-interface
checkpoint. They are neutral fixtures, not copies of knowledge samples or
evidence of a production integration.

Validate and prepare the implementation manifest with the shared scorer:

```powershell
$run = Join-Path ([IO.Path]::GetTempPath()) 'bcquality-implementation-guidance-run'
pwsh ./tools/Test-DevelopmentGuidanceFixtures.ps1 -Root . `
-ManifestPath ./evaluation/implementation-guidance-fixtures.json `
-PrepareDirectory $run
pwsh ./tools/Test-ImplementationGuidanceEvaluator.ps1
```

The same runner-owned baseline and workspace-map flow used for plan guidance
proves that evaluated target repositories remain byte-for-byte and
Git-identity stable across scoring. The implementation regression adds strict
phase/decision/pin preservation and exact path + decision key + evidence
fingerprint deduplication checks. This remains before/after evidence, not an OS
sandbox; it cannot detect a transient reverted write.

BCQuality does not keep consumed state or run checkpoints automatically. The
consumer chooses explicit checkpoints, persists consumed triples, recomputes
evidence fingerprints, owns edits/build/tests/retries/delivery, and runs an
independent final review. `no-knowledge` is only an additive retrieval outcome,
never a functional-correctness or release-readiness claim.

A future experiment should compare four matched arms: baseline, plan-only,
implementation-only, and combined. Hold task, starting code, model, tools,
runtime, checkpoints, and gates constant. This is a recommended evaluation
design, not a claimed result.
Loading