Skip to main content

Search AIMS

Find pages, publications, events, course projects, and software. Results update as you type; press Enter to open the first result.

Reasoning

SugarCrepe

SugarCrepe: a non-gameable vision-language compositionality benchmark of ~7.5k (image, positive caption, hard-negative caption) triples across 7 hard-negative types (add/replace/swap of attributes, objects, relations). Models must pick the caption that matches the COCO-2017 image. This build ingests the released per-item GPT-4V responses from the official repo (the only model with per-instance outputs released; the 17 CLIP models in the paper have only aggregate accuracy), in two caption-presentation orders (positive-first / negative-first).

15,024items
1subjects
MIT for the SugarCrepe release; CC-BY-4.0 for COCO annotations and website. COCO images retain their individual Flickr/Creative Commons licenses, including noncommercial conditions; see each trace and the original COCO annotations.license
reasoningdomain
imagemodality
textmodality
item-level responses released
Saturation status: Yes

Response matrix

Loading response matrix…

Correct (1)Incorrect (0)Unobserved

Scale: 1 = correct · 0 = incorrect