General

EVOUNA

EVOUNA (QA-Eval): per-(model, question) human correctness judgments for five Open-QA systems (FiD, GPT-3.5/text-davinci-003, ChatGPT, GPT-4, New Bing) answering the same Open-domain questions from Natural Questions (3,610) and TriviaQA (2,000). Each model answer was manually judged correct/incorrect by a human annotator.

5,176items
5subjects
Apache-2.0license
knowledgedomain
textmodality
item-level responses released
Saturation status: Unknown

Response matrix

Fit to width. Hover for subject & item; click a cell for details.

EVOUNA response matrix: AI models (rows) against items (columns)
Correct (1)Incorrect (0)Unobserved

Scale: 1 = correct · 0 = incorrect

Subjects

  1. 1New Bing0.8297
  2. 2GPT-40.8294
  3. 3ChatGPT (gpt-3.5-turbo)0.7583
  4. 4FiD0.7337
  5. 5GPT-3.5 (text-davinci-003)0.7005