General
BiGGen-Bench
BiGGen-Bench rates model answers to prompts spanning reasoning, instruction following, safety, planning and other capabilities, with strong models and in some cases human annotators scoring each answer against a rubric.
764items
103subjects
CC-BY-SA-4.0license
generaldomain
textmodality
item-level responses released
Saturation status: No
Response matrix
Loading session strips…
lowhighUnobserved
Scale: {-1, 1, 2, 3, 4, 5} (-1 = N/A)