Skip to content

[Validation] Evaluate MVP evidence and choose RunSift's next phase #36

Description

@Dyu20705

Problem statement

RunSift must not move from an offline rule-based MVP into GitHub integration, machine learning, metrics, or platform infrastructure based only on technical interest. The next phase must be selected from observed maintainer behavior and reviewed failures.

Experiment

Use the released MVP against historical failures from real maintainers while requiring no GitHub App installation.

Collect:

  • multiple reviewed failures from at least three maintainers or repositories;
  • the action the maintainer actually took;
  • whether the selected evidence was useful;
  • whether the category and temporal assessment were understandable;
  • whether abstention was more useful than an incorrect guess;
  • willingness to repeat the workflow or test a read-only beta.

Metrics and analysis

  • evidence usefulness rate;
  • rule coverage and abstention rate;
  • reviewed category agreement;
  • correct-next-action rate;
  • false-positive analysis for transient/flaky-like recommendations;
  • repeated use or beta intent;
  • cases rules cannot solve safely;
  • privacy and collection friction.

Decision options

Choose and document exactly one primary next direction:

  1. continue improving the rule-based CLI;
  2. proceed with read-only GitHub collection;
  3. simplify the taxonomy;
  4. pivot to evidence ranking or similar-incident retrieval;
  5. start a reproducible ML baseline using reviewed natural data;
  6. stop or archive the project.

Metrics, GitHub App work, ML, LLMs, dashboards, and broader infrastructure must not be activated automatically.

Checklist

  • Define review protocol and consent/data-handling notes
  • Recruit at least three maintainers or repositories
  • Review multiple failures per participating repository where possible
  • Record labels, corrections, evidence usefulness, and actual next actions
  • Measure rule coverage, abstention, and disagreement
  • Analyze privacy and workflow friction
  • Publish a decision record with supporting evidence
  • Update roadmap, milestones, and backlog to match the selected direction
  • Close or downgrade work that is not supported by evidence

Acceptance criteria

  • The decision is based primarily on reviewed natural failures, not synthetic fixtures
  • Strong observations are separated from weak heuristics
  • The report includes negative results and failure cases
  • One primary next direction is selected with explicit go/no-go conditions
  • [Post-MVP Candidate] CI Reliability Metrics and Trend Exports #19 and any future ML/service epics remain inactive until this gate supports them

Dependencies

Metadata

Metadata

Assignees

No one assigned

    Labels

    area:documentationREADME, examples, demo assets, and case studiesarea:productProduct scope, requirements, roadmap, and non-goalspriority:highImportant work planned for the current milestonesize:lLarge estimate; consider splitting before implementationtype:researchResearch, spike, or investigation

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions