Preference
Arena-Hard-Auto
Arena-Hard-Auto rates how well models answer hard user requests, many of them programming and technical questions, by having a strong model judge each answer against a fixed baseline model's answer to the same request.
500items
72subjects
Apache-2.0license
preferencedomain
textmodality
item-level responses released
Saturation status: No
Response matrix
Loading session strips…
lowhighUnobserved
Scale: {0, 0.125, 0.25, ..., 1} (8-level judge scale)