Preference

WildVision-Bench

WildVision-Bench: 500 in-the-wild image+text instructions answered by ~20 vision-language models. A GPT-4o judge scores each model response head-to-head against a fixed claude-3-sonnet reference on a five-level preference scale (Better++/Better/Tie/Worse/Worse++), mapped to a win-score in [0, 1] per (model, item).

425items
20subjects
CC-BY-4.0license
preferencedomain
imagemodality
textmodality
item-level responses released
Saturation status: Unknown

Response matrix

Fit to width. Hover for subject & item; click a cell for details.

WildVision-Bench response matrix: AI models (rows) against items (columns)
lowhighUnobserved

Scale: {0.0, 0.25, 0.5, 0.75, 1.0}

Subjects

  1. 1gpt-4o0.782
  2. 2gpt-4-vision-preview0.697
  3. 3Reka-Flash0.5945
  4. 4claude-3-opus-202402290.5675
  5. 5yi-vl-plus0.536
  6. 6liuhaotian/llava-v1.6-34b0.5125
  7. 7claude-3-sonnet-202402290.5005
  8. 8claude-3-haiku-202403070.4175
  9. 9gemini-pro-vision0.395
  10. 10deepseek-ai/deepseek-vl-7b-chat0.394
  11. 11liuhaotian/llava-v1.6-vicuna-13b0.393
  12. 12THUDM/cogvlm-chat-hf0.368
  13. 13liuhaotian/llava-v1.6-vicuna-7b0.343
  14. 14idefics2-8b-chatty0.321
  15. 15Qwen/Qwen-VL-Chat0.2605
  16. 16liuhaotian/llava-v1.5-13b0.2375
  17. 17BAAI/Bunny-v1_0-3B0.228
  18. 18openbmb/MiniCPM-V0.2125
  19. 19bczhou/tiny-llava-v1-hf0.169
  20. 20unum-cloud/uform-gen2-qwen-500m0.1575