Preference
Tulu Human Evaluation
Tulu 1 human evaluation: 332 prompt instances, each with completions from open instruction-tuned models (tulu_65B, tulu_7B, tulu_human_only_65B) and ChatGPT, graded by human annotators in pairwise comparisons. Subjects are the AI models; the response is the per-(model, item) human acceptability label (yes/no -> 1/0). Only the released human-eval slice is per-item public; the paper's automatic benchmark predictions (MMLU, GSM8k, BBH, etc.) are not released and are excluded.
332items
4subjects
Apache-2.0license
preferencedomain
textmodality
item-level responses released
Saturation status: Unknown
Response matrix
Fit to width. Hover for subject & item; click a cell for details.

Correct (1)Incorrect (0)Unobserved
Scale: 1 = correct · 0 = incorrect
Subjects
- 1chatgpt0.9086
- 2tulu_65B0.8035
- 3tulu_human_only_65B0.7265
- 4tulu_7B0.6889