Knowledge
EVOUNA
EVOUNA (QA-Eval): per-(model, question) human correctness judgments for five Open-QA systems (FiD, GPT-3.5/text-davinci-003, ChatGPT, GPT-4, New Bing) answering the same Open-domain questions from Natural Questions (3,610) and TriviaQA (2,000). Each model answer was manually judged correct/incorrect by a human annotator.
5,162items
5subjects
Apache-2.0license
knowledgedomain
textmodality
item-level responses released
Saturation status: No
Response matrix
Loading response matrix…
Correct (1)Incorrect (0)Unobserved
Scale: 1 = correct · 0 = incorrect