General
SugarCrepe
SugarCrepe: a non-gameable vision-language compositionality benchmark of ~7.5k (image, positive caption, hard-negative caption) triples across 7 hard-negative types (add/replace/swap of attributes, objects, relations). Models must pick the caption that matches the COCO-2017 image. This build ingests the released per-item GPT-4V responses from the official repo (the only model with per-instance outputs released; the 17 CLIP models in the paper have only aggregate accuracy), in two caption-presentation orders (positive-first / negative-first).
7,512items
1subjects
MITlicense
reasoningdomain
imagemodality
textmodality
item-level responses released
Saturation status: Unknown
Response matrix
Fit to width. Hover for subject & item; click a cell for details.

Correct (1)Incorrect (0)Unobserved
Scale: 1 = correct · 0 = incorrect