Preference

Arena-Hard-Auto

Arena-Hard-Auto rates how well models answer hard user requests, many of them programming and technical questions, by having a strong model judge each answer against a fixed baseline model's answer to the same request.

500items
72subjects
Apache-2.0license
preferencedomain
textmodality
item-level responses released
Saturation status: No

Response matrix

Loading session strips…

lowhighUnobserved

Scale: {0, 0.125, 0.25, ..., 1} (8-level judge scale)