Skip to content

Add WinDbg diagnosis plugin - #209

Merged
Nikola Metulev (nmetulev) merged 6 commits into
mainfrom
rtrimceski-ms-add-debug-diagnostician
Oct 7, 2026
Merged

Nikola Metulev (nmetulev) merged 6 commits into
mainfrom
rtrimceski-ms-add-debug-diagnostician

Conversation

@rtrimceski-MS

@rtrimceski-MS rtrimceski-MS commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

Description

Adds the windbg 1.0.0 local plugin with curated user-mode and kernel-mode diagnostic skills, the shared windbg-diagnostic-method, deterministic report validation, and Copilot diagnostician/contrarian agents.

Registers the plugin in the Copilot, Claude Code, and Codex catalogs; adds repository ownership, documentation, changelog, and CI validation; and provides consistent windbg-user-* and windbg-kernel-* skill IDs.

Related Issue

N/A

Type of Change

  • 📦 Catalog entry added, moved to a new pin, or removed
  • 📝 Documentation
  • 🔧 Config / CI

Checklist

  • All three catalogs (.github/plugin, .claude-plugin, .agents/plugins) list the same plugins at the same pin
  • Installed the plugin from this branch on each host and confirmed its skills load
  • If user-facing: added a bullet to ## [Unreleased] in CHANGELOG.md

Additional Notes

The runtime package is limited to the public inventory documented in skills.json and the plugin README. Package validation restricts documentation links to approved public hosts and rejects email addresses.

The bug-family skills originated as copies of the versions in Agency Marketplace (windows-engineering), then received public packaging, naming, routing, and validation updates in this PR.

Manual host validation:

  • @rtrimceski-MS verified the plugin in Claude Code using a real user-mode memory dump.
  • Dragos verified the plugin in GitHub Copilot CLI and VS Code.

Before publication, reviewers should confirm redistribution/license approval, representative WinDbg/dump/TTD/driver testing, long-term ownership, and WinDbg-Feedback triage. CODEOWNERS includes the existing catalog maintainers plus @rtrimceski-MS.

Validation:

  • node scripts/check-catalogs.mjs — passed; all three plugins resolve and match across the three catalogs.
  • npm ci --prefix scripts/vally --no-audit --no-fund — passed.
  • node scripts/validate-windbg.mjs — passed.
  • node scripts/tests/test-windbg-validator.mjs — passed.
  • node scripts/vally/lint-skills.mjs — passed; 11/11 skills.
  • node scripts/tests/test-catalog-tools.mjs — passed; all five catalog-tool cases.
  • git diff --check — passed.

AI Description

This section is auto-generated by AI when the PR is opened or updated. To opt out, delete this entire section including the marker comments.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@rtrimceski-MS
rtrimceski-MS force-pushed the rtrimceski-ms-add-debug-diagnostician branch from 020b8bb to bee99ec Compare October 7, 2026 09:13
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

@nmetulev Nikola Metulev (nmetulev) left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for putting this together. I merged main into the branch (README/CONTRIBUTING conflicts) and then tested it on Copilot CLI, Claude Code, and Codex. I also ran routing benchmarks with Claude Sonnet 5 and GPT-5.5, with this plugin alone and alongside winui, winappcli, and superpowers (48 skills total).

Overall it's in good shape. Install and skill discovery work on all three hosts. Skills routed correctly on 68/68 debugging prompts, and the responses followed each skill's workflow. Nothing clashed with the other plugins. Requested changes below.

Must fix

1. contrarian agent doesn't load in Copilot. Copilot only loads agents directly under com.github.copilot/agents/, so agents/fleet/contrarian.agent.md is skipped. --agent debugging-diagnostician:contrarian fails with "No such agent", and only diagnostician is listed. In an end-to-end run, the diagnostician only got a review by searching for the file and passing its text to a general-purpose sub-agent. That depends on the model choosing to do it, doesn't enforce tools: [], and doesn't control the model. Moving the file to agents/contrarian.agent.md makes it load (verified). The diagnostician should then invoke it by name rather than by the fleet/contrarian.agent.md path. The name contrarian is fine as-is since Copilot namespaces plugin agents.

2. com.github.copilot/instructions/*.instructions.md never reach the model. instructions/ isn't a component location Copilot loads for Agent Plugins manifests, and the files have no applyTo front matter either. I confirmed neither file is in the agent's context.

  • Most of diagnostic-reasoning duplicates the agent body. The unique parts are the fix-confidence calibration table, the trigger/fix gates, and the output contract. The contrarian references that table by file name but never sees it.
  • root-cause-analysis has content that exists nowhere else: the evidence ladder, first-pass normalization commands, the quality bar, and the safety/privacy rules.
  • In the end-to-end run the agent omitted fix_confidence and reported contrarian_loopback: 1 (not a boolean), which is consistent with these files not loading.

Suggestion: move the unique content into a shared method skill (e.g. windbg-diagnostic-method). The agent is Copilot-only, so Claude Code and Codex users currently get none of this methodology, and a skill would load on all three hosts. Folding it into diagnostician.agent.md is the simpler option, but it only helps Copilot.

3. validate-diagnosis-output reference script has logic bugs.

  • In lines like $results.ReasoningChainHeading = HasHeading '^##\s+Reasoning Chain\b' -and ($content -match ...), PowerShell passes -and (...) to HasHeading as extra arguments ($args), which are then ignored. As a result, the Reasoning Chain, Trigger Verification, and Contrarian Verdict checks only check that the heading exists. Verified: function HasHeading($p){$true}; HasHeading 'x' -and $false returns True. Wrap the call: (HasHeading '...') -and (...).
  • Check 5 (≥2 alternatives) isn't implemented; the script only checks the heading.
  • The content checks run against the whole document rather than the section body.
  • The Mermaid check rejects valid sequence diagrams with implicit participants (A->>B: msg).
  • The skill describes itself as a "Deterministic, no-LLM structural validator", but no script ships; the agent re-implements it inline. It also refers to final completion, a "Pre-Completion Checklist", and "the task tool", none of which exist in this agent.

Should fix

4. um-exception-triage over-triggers on managed .NET, WinUI/XAML, and MSIX crashes with GPT models. With gpt-5.5, the WinUI XAML parse exception, C# NullReferenceException, and "crashes after MSIX install" prompts routed here in 8 of 9 repeated runs. Claude had none. The plugin is native-only by design (README: "not a managed .NET diagnostics package"), and no skill covers SOS or 0xE0434352. Scoping the description fixed it: 0 of 9 false positives, still 9/9 on real cases. Suggested wording (I tested a stricter variant, not this exact text, so worth a quick re-run):

Use when a native (C/C++) app, service, or user-mode driver host (including UMDF) crashes with a structured exception in a dump or WinDbg session, including native faults inside managed processes; establish context and classify it. Not for managed .NET exceptions (use SOS/dotnet-dump), WinUI/XAML app errors, or kernel bugchecks.

5. virtual-memory-exhaustion needs a similar exclusion. A ".NET 8 service with high memory, analyze with dotnet-dump" prompt triggered it 1 time in 3. Something like "Not for managed .NET heap growth" should cover it.

6. validate-diagnosis-output on Claude Code and Codex. Claude Code and Codex load it (it adds to the always-on context), but the agent and report flow it validates are Copilot-only. Consider moving it into the agent or making it a real script.

Naming and structure (for discussion)

  • Plugin name: debugging-diagnostician is redundant and doesn't say Windows. Since the other entries in this catalog use product names (winui, winappcli), I'd suggest windbg. It also shortens the namespaced skill IDs on Claude and Codex (debugging-diagnostician:um-exception-triage → windbg:...).
  • Skill names: prefixes are inconsistent (km-* next to kernel-bugcheck-triage, and um- on only one skill). Copilot invokes skills by bare name, so a common prefix like windbg-* would follow the winui-*/winapp-* convention and avoid collisions with other plugins. Agents don't need it.
  • mutex-held-across-co-await is a single bug pattern rather than a workflow. It might fit better as a hypothesis or section inside wait-chain-analysis or um-exception-triage.

Cleanup

  • The hard-coded "ten" in the plugin and catalog descriptions, skills.json, validate-package.mjs, and the README, and the per-skill 1.0.0 in every Feedback section and version: front matter, will drift on the next change. I'd describe the value rather than the count, and drop per-skill versions.
  • scripts/validate-package.mjs is minified, hard-codes counts, and ships inside the plugin to users. Consider moving it to the repo's scripts/ (or dropping it) since vally lint and check-catalogs.mjs already cover most of it.
  • CHANGELOG.md: the edited "This repo now hosts..." bullet got rewrapped with a dangling "Copilot and Claude Code" line.

What I ran

  • node scripts/check-catalogs.mjs, node plugins/debugging-diagnostician/scripts/validate-package.mjs, node scripts/vally/lint-skills.mjs (11/11), node scripts/tests/test-catalog-tools.mjs: all pass.
  • claude plugin validate (plugin and marketplace) passes; installs with 11 skills and 0 agents.
  • Codex 0.159.2: installs, and codex debug prompt-input shows all 11 skills.
  • Copilot CLI: installs with 11 skills and the diagnostician agent; contrarian and the instruction files don't load (above).
  • Routing benchmark: 25 prompts (17 positive, 8 negative) × 2 models × 2 configs, isolated profiles with shell and write tools denied, plus 3× repeats for the description A/B.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@rtrimceski-MS

rtrimceski-MS commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor Author

Thanks Nikola Metulev (@nmetulev) — I pushed fixes for the following items from your review:

  • Must fix This repo is missing a LICENSE file #1: moved contrarian.agent.md directly under com.github.copilot/agents/, updated the diagnostician to invoke contrarian by name, and verified debugging-diagnostician:contrarian loads in Copilot CLI.
  • Must fix This repo is missing important files #3: replaced the prose-only validator with a bundled PowerShell script; fixed section-scoped checks, alternative counting, Boolean handling, and implicit Mermaid participants; removed references to nonexistent agent concepts; and added pass/fail regression tests.
  • Should fix initial #4/Tweaks #5: tightened um-exception-triage and virtual-memory-exhaustion descriptions to exclude managed .NET and WinUI/XAML scenarios.
  • Cleanup: removed hard-coded skill counts and per-skill versions, moved package validation to repository scripts/, reformatted it, and fixed the changelog wrapping.

The methodology packaging (must fix #2), validator host scope (should fix #6), and naming/skill-structure suggestions remain open for the design-decision pass.

Current catalog, package, Vally, validator regression, and catalog-tool tests pass.

rtrimceski-MS and others added 2 commits October 7, 2026 11:20
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@rtrimceski-MS rtrimceski-MS changed the title Add debugging diagnostician plugin Add WinDbg diagnosis plugin Oct 7, 2026
@rtrimceski-MS

Copy link
Copy Markdown
Contributor Author

Follow-up on the remaining design feedback:

  • Renamed the plugin to windbg and updated all catalogs/install docs.
  • Normalized the bug-family skills to windbg-user-* and windbg-kernel-* IDs.
  • Replaced the unsupported Copilot instruction files with the cross-host windbg-diagnostic-method skill, which now contains the evidence ladder, five-phase method, routing, confidence calibration, trigger/fix gates, report contract, safety guidance, and deterministic validator.
  • Folded report validation into that method skill and removed the separate validator skill/context entry.
  • Kept windbg-user-mutex-held-across-co-await as a standalone skill after review.
  • Updated package validation, CI, documentation, CODEOWNERS, and regression tests for the new structure.

The renamed package resolves across all three catalogs, all 11 skills pass Vally, and the validator pass/fail regression tests pass. Copilot CLI loads the local windbg plugin and accepts the windbg:contrarian agent namespace.

- Shorten windbg-user-exception-triage description to 271 chars (under the 300-char bar) while keeping the managed .NET / WinUI exclusions.

- Ask each bug-family skill to load windbg-diagnostic-method first, so Claude-family models apply the method without the Copilot agent (Sonnet co-load 1/17 -> 10/17 in routing benchmark; no routing or false-positive regressions).

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@nmetulev
Nikola Metulev (nmetulev) merged commit 5ce74fa into main Oct 7, 2026
6 checks passed
@nmetulev
Nikola Metulev (nmetulev) deleted the rtrimceski-ms-add-debug-diagnostician branch October 7, 2026 20:54
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants