General
FineGRAIN (T2I)
FineGRAIN T2I failure-mode benchmark: ~17 text-to-image models x 760 prompts, each prompt tagged with one of 27 fine-grained failure modes (counting, colour/shape/texture attribute binding, spatial relations, physics, text rendering, negation, perspective, ...) across 11 categories. Each generated image carries a human label for whether the prompt's failure mode is present; response is the human success verdict (1 = no failure / prompt rendered correctly, 0 = failure).
760items
17subjects
MITlicense
generaldomain
textmodality
imagemodality
item-level responses released
Saturation status: Unknown
Response matrix
Fit to width. Hover for subject & item; click a cell for details.

Correct (1)Incorrect (0)Unobserved
Scale: 1 = correct · 0 = incorrect
Subjects
- 1qwen-image1
- 2gemini_image1
- 3seeDream31
- 4flux2_pro1
- 5wan221
- 6flux_kontext1
- 7gpt_image11
- 8gpt_image151
- 9nano_banana21
- 10hidream1
- 11sdv1.51
- 12sd2.11
- 13sd3.5_medium0.516
- 14flux0.516
- 15sd3_m0.516
- 16sd3_xl0.516
- 17sd3.5_large0.516