Skip to main content

Search AIMS

Find pages, publications, events, course projects, and software. Results update as you type; press Enter to open the first result.

General

Tulu Human Evaluation

Released Tulu 1 human acceptability judgments for four language models on 332 instruction-following prompts.

1,028items
4subjects
Apache-2.0 for the original Open Instruct release; source prompts and model outputs retain their upstream terms.license
instruction_followingdomain
textmodality
item-level responses released
Saturation status: Yes

Response matrix

Loading response matrix…

Correct (1)Incorrect (0)Unobserved

Scale: 1 = correct · 0 = incorrect