General

FineGRAIN (T2I)

FineGRAIN T2I failure-mode benchmark: ~17 text-to-image models x 760 prompts, each prompt tagged with one of 27 fine-grained failure modes (counting, colour/shape/texture attribute binding, spatial relations, physics, text rendering, negation, perspective, ...) across 11 categories. Each generated image carries a human label for whether the prompt's failure mode is present; response is the human success verdict (1 = no failure / prompt rendered correctly, 0 = failure).

760items
17subjects
MITlicense
generaldomain
textmodality
imagemodality
item-level responses released
Saturation status: Unknown

Response matrix

Fit to width. Hover for subject & item; click a cell for details.

FineGRAIN (T2I) response matrix: AI models (rows) against items (columns)
Correct (1)Incorrect (0)Unobserved

Scale: 1 = correct · 0 = incorrect

Subjects

  1. 1qwen-image1
  2. 2gemini_image1
  3. 3seeDream31
  4. 4flux2_pro1
  5. 5wan221
  6. 6flux_kontext1
  7. 7gpt_image11
  8. 8gpt_image151
  9. 9nano_banana21
  10. 10hidream1
  11. 11sdv1.51
  12. 12sd2.11
  13. 13sd3.5_medium0.516
  14. 14flux0.516
  15. 15sd3_m0.516
  16. 16sd3_xl0.516
  17. 17sd3.5_large0.516