Preference
JudgeBench
JudgeBench tests whether a model acting as a judge picks the correct response out of a pair, on comparisons drawn from knowledge, reasoning, math and coding, with each pair shown in both orders.
350items
19subjects
MITlicense
reward_modelingdomain
preferencedomain
textmodality
item-level responses released
Saturation status: No
Response matrix
Rasch analysis p = σ(θ − z + c)
2,897 responses, 80/20 split over cells · 19 subjects · 350 items · 4 conditions
1 of 19 subjects answered every item alike, so its θ is unbounded and sits pegged to the column edge.
AUC train
0.858
AUC test
0.747
Loading session strips…
Correct (1)Incorrect (0)Unobserved
Scale: 1 = correct · 0 = incorrect