General

BiGGen-Bench

BiGGen-Bench rates model answers to prompts spanning reasoning, instruction following, safety, planning and other capabilities, with strong models and in some cases human annotators scoring each answer against a rubric.

764items
103subjects
CC-BY-SA-4.0license
generaldomain
textmodality
item-level responses released
Saturation status: No

Response matrix

Loading session strips…

lowhighUnobserved

Scale: {-1, 1, 2, 3, 4, 5} (-1 = N/A)