Preference

JudgeBench

JudgeBench tests whether a model acting as a judge picks the correct response out of a pair, on comparisons drawn from knowledge, reasoning, math and coding, with each pair shown in both orders.

350items
19subjects
MITlicense
reward_modelingdomain
preferencedomain
textmodality
item-level responses released
Saturation status: No

Response matrix

Rasch analysis p = σ(θ − z + c)

2,897 responses, 80/20 split over cells · 19 subjects · 350 items · 4 conditions

1 of 19 subjects answered every item alike, so its θ is unbounded and sits pegged to the column edge.

AUC train
0.858
AUC test
0.747

Loading session strips…

Correct (1)Incorrect (0)Unobserved

Scale: 1 = correct · 0 = incorrect