Preference

Tulu Human Evaluation

Tulu 1 human evaluation: 332 prompt instances, each with completions from open instruction-tuned models (tulu_65B, tulu_7B, tulu_human_only_65B) and ChatGPT, graded by human annotators in pairwise comparisons. Subjects are the AI models; the response is the per-(model, item) human acceptability label (yes/no -> 1/0). Only the released human-eval slice is per-item public; the paper's automatic benchmark predictions (MMLU, GSM8k, BBH, etc.) are not released and are excluded.

332items
4subjects
Apache-2.0license
preferencedomain
textmodality
item-level responses released
Saturation status: Unknown

Response matrix

Fit to width. Hover for subject & item; click a cell for details.

Tulu Human Evaluation response matrix: AI models (rows) against items (columns)
Correct (1)Incorrect (0)Unobserved

Scale: 1 = correct · 0 = incorrect

Subjects

  1. 1chatgpt0.9086
  2. 2tulu_65B0.8035
  3. 3tulu_human_only_65B0.7265
  4. 4tulu_7B0.6889