Hi,
The article “GPT-5.6 Luna vs GPT-6 Astra: Is a $1.20 Model Good Enough for Code Review?” says that the prompts, raw model outputs, judge verdicts, bug-class labels, repeat runs, and scoring scripts are publicly available.
However, I cannot find a link to the repository or commit containing these materials. The article links to the AI-Code-Review-Evals organization, but I only see the benchmark PR repositories there.
Could you please provide the exact URL to the reproducibility materials?
Article: https://entelligence.ai/blogs/gpt-5.6-luna-vs-gpt-6-astra-is-a-1.20-model-good-enough-for-code-review
Thank you!
Hi,
The article “GPT-5.6 Luna vs GPT-6 Astra: Is a $1.20 Model Good Enough for Code Review?” says that the prompts, raw model outputs, judge verdicts, bug-class labels, repeat runs, and scoring scripts are publicly available.
However, I cannot find a link to the repository or commit containing these materials. The article links to the AI-Code-Review-Evals organization, but I only see the benchmark PR repositories there.
Could you please provide the exact URL to the reproducibility materials?
Article: https://entelligence.ai/blogs/gpt-5.6-luna-vs-gpt-6-astra-is-a-1.20-model-good-enough-for-code-review
Thank you!