Skip to content

A gaming benchmark: deliberately inflate every figure and measure what resists #17

Description

@Bubblegunn

The README already publishes the tool's worst case honestly: an unmarked generated file took one author's surviving-lines share from 51.6% to 99.3%, with the defence named beside it. That is one attack, found by thinking of it.

The elevation is turning that paragraph into a measured benchmark. A repository built to be gamed, a list of attacks, and for each figure the number before and after. Then the tool's claim stops being "we thought about gaming" and becomes "here is what each figure does when someone tries".

Attacks worth including, and the list should grow:

  • Splitting one change across many commits, to inflate commit counts and cadence.
  • Committing generated or vendored code without marking it.
  • Reformatting a file to take its blame.
  • Committing at spread-out times to fake a cadence.
  • Rewriting history to reassign authorship, which is trivially possible and the report should say so.
  • Long single-line files, since a line count is not a work count.
  • A repository with one enormous initial import.

Done when

  • The gamed repository is generated by a committed script, so anybody can rebuild it and check the numbers rather than trust a table.
  • Every figure is reported before and after each attack, and figures that do not resist are published exactly as loudly as the ones that do. A benchmark that only lists successes is marketing.
  • Each attack names the defence if there is one, and says plainly if there is not.
  • The README's adversarial section is replaced by the generated table, so it cannot drift from what the code does.

Out of scope. Any attempt to detect gaming automatically. This tool's answer to a determined liar is transparency about what each figure can be made to say, not an arms race.

Why this matters more here than anywhere else. The report is meant to be shown to someone deciding about a person's career. A figure that inflates easily and is presented as evidence is not a bug in a tool, it is a way to mislead an employer. Publishing the attack surface is the honest thing and, incidentally, the thing that makes this tool citable.


This one is a study, not a feature. The output is a result and a write-up, and it is fine if the result is boring.

How studies work in these repositories. The design goes in the repository before the data is collected, including what a null result would mean, so the conclusion cannot drift with the numbers. Every figure comes from a run somebody else can repeat, with the command and the version committed beside it. A null or negative result is published in the same place a positive one would go, including when it undercuts the tool. Nothing invented: a number is measured, cited, or it does not appear.

Taking this on: comment and it is yours. A pull request containing only the design document is a complete contribution and often the most useful one, so do not feel you have to finish the whole thing to start.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    help wantedExtra attention is neededresearchA study or measurement, not a feature; the output is a result and a write-up

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions