Hi, thanks for the great work!
I noticed that the paper/project uses GPT-based evaluation for assessing generation quality (e.g., realism, consistency, interaction quality, etc.).
Would it be possible to release the evaluation pipeline/scripts used in the paper?
In particular, it would be very helpful to have:
- GPT prompts/templates
- evaluation scripts
- score parsing logic
- API settings (model version, temperature, etc.)
- batch evaluation pipeline
This would greatly improve reproducibility and help the community perform fair comparisons with future methods.
Thanks!
Hi, thanks for the great work!
I noticed that the paper/project uses GPT-based evaluation for assessing generation quality (e.g., realism, consistency, interaction quality, etc.).
Would it be possible to release the evaluation pipeline/scripts used in the paper?
In particular, it would be very helpful to have:
This would greatly improve reproducibility and help the community perform fair comparisons with future methods.
Thanks!