General
EVOUNA
EVOUNA (QA-Eval): per-(model, question) human correctness judgments for five Open-QA systems (FiD, GPT-3.5/text-davinci-003, ChatGPT, GPT-4, New Bing) answering the same Open-domain questions from Natural Questions (3,610) and TriviaQA (2,000). Each model answer was manually judged correct/incorrect by a human annotator.
5,176items
5subjects
Apache-2.0license
knowledgedomain
textmodality
item-level responses released
Saturation status: Unknown
Response matrix
Fit to width. Hover for subject & item; click a cell for details.

Correct (1)Incorrect (0)Unobserved
Scale: 1 = correct · 0 = incorrect
Subjects
- 1New Bing0.8297
- 2GPT-40.8294
- 3ChatGPT (gpt-3.5-turbo)0.7583
- 4FiD0.7337
- 5GPT-3.5 (text-davinci-003)0.7005