We'd still have a single judgement that would be the mean of results all paraphrases.
Purpose: suppose you want to evaluate e.g. sentiment of some text, on 0-100 scale. You can ask the question multiple ways. It just seems that judge with various paraphrases will give more reliable results that a single paraphrase only.
We'd still have a single judgement that would be the mean of results all paraphrases.
Purpose: suppose you want to evaluate e.g. sentiment of some text, on 0-100 scale. You can ask the question multiple ways. It just seems that judge with various paraphrases will give more reliable results that a single paraphrase only.