The best first contribution here is the one this format exists for: take an agent plugin, hook,
skill or AGENTS.md you already use, produce a report for it, and open the result.
You do not need to be right about it. A report saying "this plugin's hook never fires on my setup"
is worth more than one confirming it works, because the second is what everybody assumes and the
first is what nobody checks.
What to do: follow the README's 30-second quickstart against your chosen subject, then open an
issue with the statement attached. If the statement surprises you, say what you expected — that
contrast is the useful part.
What the report can and cannot say. Rows carry a basis. A row recomputed from the artifact
is evidence; a row taken from the author's claim is a claim, and the schema keeps them apart on
purpose. Two things follow:
- Do not describe a claimed row as a measured one.
- Do not read a
NotAvailable or NotApplicable row as a failure. Absence of evidence is its own
category here, not a low score — an artifact that declares no hooks is not a broken artifact.
One rule, and it is not optional: measure the artifact as published. Do not modify it to
make it measurable, and do not contact its author about your findings as though this were a bug
report. Everything in paper/measurements/ was collected under that rule and this should be too.
No code required. Sized to an afternoon, most of it reading your own output.
The best first contribution here is the one this format exists for: take an agent plugin, hook,
skill or
AGENTS.mdyou already use, produce a report for it, and open the result.You do not need to be right about it. A report saying "this plugin's hook never fires on my setup"
is worth more than one confirming it works, because the second is what everybody assumes and the
first is what nobody checks.
What to do: follow the README's 30-second quickstart against your chosen subject, then open an
issue with the statement attached. If the statement surprises you, say what you expected — that
contrast is the useful part.
What the report can and cannot say. Rows carry a basis. A row recomputed from the artifact
is evidence; a row taken from the author's claim is a claim, and the schema keeps them apart on
purpose. Two things follow:
NotAvailableorNotApplicablerow as a failure. Absence of evidence is its owncategory here, not a low score — an artifact that declares no hooks is not a broken artifact.
One rule, and it is not optional: measure the artifact as published. Do not modify it to
make it measurable, and do not contact its author about your findings as though this were a bug
report. Everything in
paper/measurements/was collected under that rule and this should be too.No code required. Sized to an afternoon, most of it reading your own output.