General

SugarCrepe

SugarCrepe: a non-gameable vision-language compositionality benchmark of ~7.5k (image, positive caption, hard-negative caption) triples across 7 hard-negative types (add/replace/swap of attributes, objects, relations). Models must pick the caption that matches the COCO-2017 image. This build ingests the released per-item GPT-4V responses from the official repo (the only model with per-instance outputs released; the 17 CLIP models in the paper have only aggregate accuracy), in two caption-presentation orders (positive-first / negative-first).

7,512items
1subjects
MITlicense
reasoningdomain
imagemodality
textmodality
item-level responses released
Saturation status: Unknown

Response matrix

Fit to width. Hover for subject & item; click a cell for details.

SugarCrepe response matrix: AI models (rows) against items (columns)
Correct (1)Incorrect (0)Unobserved

Scale: 1 = correct · 0 = incorrect