Preference

WildBench

WildBench rates how well models handle challenging tasks taken from real user conversations, with a strong model acting as the judge and rating each answer on a numeric scale.

1,024items
71subjects
CC-BY-4.0license
preferencedomain
textmodality
item-level responses released
Saturation status: No

Response matrix

Loading session strips…

lowhighUnobserved

Scale: {1, 2, ..., 10}