Repository navigation
Add WinDbg diagnosis plugin - #209
Conversation
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
020b8bb to
bee99ec
Compare
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Nikola Metulev (nmetulev)
left a comment
There was a problem hiding this comment.
Thanks for putting this together. I merged main into the branch (README/CONTRIBUTING conflicts) and then tested it on Copilot CLI, Claude Code, and Codex. I also ran routing benchmarks with Claude Sonnet 5 and GPT-5.5, with this plugin alone and alongside winui, winappcli, and superpowers (48 skills total).
Overall it's in good shape. Install and skill discovery work on all three hosts. Skills routed correctly on 68/68 debugging prompts, and the responses followed each skill's workflow. Nothing clashed with the other plugins. Requested changes below.
Must fix
1. contrarian agent doesn't load in Copilot. Copilot only loads agents directly under com.github.copilot/agents/, so agents/fleet/contrarian.agent.md is skipped. --agent debugging-diagnostician:contrarian fails with "No such agent", and only diagnostician is listed. In an end-to-end run, the diagnostician only got a review by searching for the file and passing its text to a general-purpose sub-agent. That depends on the model choosing to do it, doesn't enforce tools: [], and doesn't control the model. Moving the file to agents/contrarian.agent.md makes it load (verified). The diagnostician should then invoke it by name rather than by the fleet/contrarian.agent.md path. The name contrarian is fine as-is since Copilot namespaces plugin agents.
2. com.github.copilot/instructions/*.instructions.md never reach the model. instructions/ isn't a component location Copilot loads for Agent Plugins manifests, and the files have no applyTo front matter either. I confirmed neither file is in the agent's context.
- Most of
diagnostic-reasoningduplicates the agent body. The unique parts are the fix-confidence calibration table, the trigger/fix gates, and the output contract. The contrarian references that table by file name but never sees it. root-cause-analysishas content that exists nowhere else: the evidence ladder, first-pass normalization commands, the quality bar, and the safety/privacy rules.- In the end-to-end run the agent omitted
fix_confidenceand reportedcontrarian_loopback: 1(not a boolean), which is consistent with these files not loading.
Suggestion: move the unique content into a shared method skill (e.g. windbg-diagnostic-method). The agent is Copilot-only, so Claude Code and Codex users currently get none of this methodology, and a skill would load on all three hosts. Folding it into diagnostician.agent.md is the simpler option, but it only helps Copilot.
3. validate-diagnosis-output reference script has logic bugs.
- In lines like
$results.ReasoningChainHeading = HasHeading '^##\s+Reasoning Chain\b' -and ($content -match ...), PowerShell passes-and (...)toHasHeadingas extra arguments ($args), which are then ignored. As a result, the Reasoning Chain, Trigger Verification, and Contrarian Verdict checks only check that the heading exists. Verified:function HasHeading($p){$true}; HasHeading 'x' -and $falsereturnsTrue. Wrap the call:(HasHeading '...') -and (...). - Check 5 (≥2 alternatives) isn't implemented; the script only checks the heading.
- The content checks run against the whole document rather than the section body.
- The Mermaid check rejects valid sequence diagrams with implicit participants (
A->>B: msg). - The skill describes itself as a "Deterministic, no-LLM structural validator", but no script ships; the agent re-implements it inline. It also refers to
final completion, a "Pre-Completion Checklist", and "the task tool", none of which exist in this agent.
Should fix
4. um-exception-triage over-triggers on managed .NET, WinUI/XAML, and MSIX crashes with GPT models. With gpt-5.5, the WinUI XAML parse exception, C# NullReferenceException, and "crashes after MSIX install" prompts routed here in 8 of 9 repeated runs. Claude had none. The plugin is native-only by design (README: "not a managed .NET diagnostics package"), and no skill covers SOS or 0xE0434352. Scoping the description fixed it: 0 of 9 false positives, still 9/9 on real cases. Suggested wording (I tested a stricter variant, not this exact text, so worth a quick re-run):
Use when a native (C/C++) app, service, or user-mode driver host (including UMDF) crashes with a structured exception in a dump or WinDbg session, including native faults inside managed processes; establish context and classify it. Not for managed .NET exceptions (use SOS/dotnet-dump), WinUI/XAML app errors, or kernel bugchecks.
5. virtual-memory-exhaustion needs a similar exclusion. A ".NET 8 service with high memory, analyze with dotnet-dump" prompt triggered it 1 time in 3. Something like "Not for managed .NET heap growth" should cover it.
6. validate-diagnosis-output on Claude Code and Codex. Claude Code and Codex load it (it adds to the always-on context), but the agent and report flow it validates are Copilot-only. Consider moving it into the agent or making it a real script.
Naming and structure (for discussion)
- Plugin name:
debugging-diagnosticianis redundant and doesn't say Windows. Since the other entries in this catalog use product names (winui,winappcli), I'd suggestwindbg. It also shortens the namespaced skill IDs on Claude and Codex (debugging-diagnostician:um-exception-triage→windbg:...). - Skill names: prefixes are inconsistent (
km-*next tokernel-bugcheck-triage, andum-on only one skill). Copilot invokes skills by bare name, so a common prefix likewindbg-*would follow thewinui-*/winapp-*convention and avoid collisions with other plugins. Agents don't need it. mutex-held-across-co-awaitis a single bug pattern rather than a workflow. It might fit better as a hypothesis or section insidewait-chain-analysisorum-exception-triage.
Cleanup
- The hard-coded "ten" in the plugin and catalog descriptions,
skills.json,validate-package.mjs, and the README, and the per-skill1.0.0in every Feedback section andversion:front matter, will drift on the next change. I'd describe the value rather than the count, and drop per-skill versions. scripts/validate-package.mjsis minified, hard-codes counts, and ships inside the plugin to users. Consider moving it to the repo'sscripts/(or dropping it) since vally lint andcheck-catalogs.mjsalready cover most of it.CHANGELOG.md: the edited "This repo now hosts..." bullet got rewrapped with a dangling "Copilot and Claude Code" line.
What I ran
node scripts/check-catalogs.mjs,node plugins/debugging-diagnostician/scripts/validate-package.mjs,node scripts/vally/lint-skills.mjs(11/11),node scripts/tests/test-catalog-tools.mjs: all pass.claude plugin validate(plugin and marketplace) passes; installs with 11 skills and 0 agents.- Codex 0.159.2: installs, and
codex debug prompt-inputshows all 11 skills. - Copilot CLI: installs with 11 skills and the
diagnosticianagent;contrarianand the instruction files don't load (above). - Routing benchmark: 25 prompts (17 positive, 8 negative) × 2 models × 2 configs, isolated profiles with shell and write tools denied, plus 3× repeats for the description A/B.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
|
Thanks Nikola Metulev (@nmetulev) — I pushed fixes for the following items from your review:
The methodology packaging (must fix #2), validator host scope (should fix #6), and naming/skill-structure suggestions remain open for the design-decision pass. Current catalog, package, Vally, validator regression, and catalog-tool tests pass. |
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
|
Follow-up on the remaining design feedback:
The renamed package resolves across all three catalogs, all 11 skills pass Vally, and the validator pass/fail regression tests pass. Copilot CLI loads the local |
- Shorten windbg-user-exception-triage description to 271 chars (under the 300-char bar) while keeping the managed .NET / WinUI exclusions. - Ask each bug-family skill to load windbg-diagnostic-method first, so Claude-family models apply the method without the Copilot agent (Sonnet co-load 1/17 -> 10/17 in routing benchmark; no routing or false-positive regressions). Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Description
Adds the
windbg1.0.0 local plugin with curated user-mode and kernel-mode diagnostic skills, the sharedwindbg-diagnostic-method, deterministic report validation, and Copilot diagnostician/contrarian agents.Registers the plugin in the Copilot, Claude Code, and Codex catalogs; adds repository ownership, documentation, changelog, and CI validation; and provides consistent
windbg-user-*andwindbg-kernel-*skill IDs.Related Issue
N/A
Type of Change
Checklist
.github/plugin,.claude-plugin,.agents/plugins) list the same plugins at the same pin## [Unreleased]inCHANGELOG.mdAdditional Notes
The runtime package is limited to the public inventory documented in
skills.jsonand the plugin README. Package validation restricts documentation links to approved public hosts and rejects email addresses.The bug-family skills originated as copies of the versions in Agency Marketplace (windows-engineering), then received public packaging, naming, routing, and validation updates in this PR.
Manual host validation:
@rtrimceski-MSverified the plugin in Claude Code using a real user-mode memory dump.Before publication, reviewers should confirm redistribution/license approval, representative WinDbg/dump/TTD/driver testing, long-term ownership, and WinDbg-Feedback triage. CODEOWNERS includes the existing catalog maintainers plus
@rtrimceski-MS.Validation:
node scripts/check-catalogs.mjs— passed; all three plugins resolve and match across the three catalogs.npm ci --prefix scripts/vally --no-audit --no-fund— passed.node scripts/validate-windbg.mjs— passed.node scripts/tests/test-windbg-validator.mjs— passed.node scripts/vally/lint-skills.mjs— passed; 11/11 skills.node scripts/tests/test-catalog-tools.mjs— passed; all five catalog-tool cases.git diff --check— passed.AI Description
This section is auto-generated by AI when the PR is opened or updated. To opt out, delete this entire section including the marker comments.