Preference
WildVision-Bench
WildVision-Bench: 500 in-the-wild image+text instructions answered by ~20 vision-language models. A GPT-4o judge scores each model response head-to-head against a fixed claude-3-sonnet reference on a five-level preference scale (Better++/Better/Tie/Worse/Worse++), mapped to a win-score in [0, 1] per (model, item).
425items
20subjects
CC-BY-4.0license
preferencedomain
imagemodality
textmodality
item-level responses released
Saturation status: Unknown
Response matrix
Fit to width. Hover for subject & item; click a cell for details.

lowhighUnobserved
Scale: {0.0, 0.25, 0.5, 0.75, 1.0}
Subjects
- 1gpt-4o0.782
- 2gpt-4-vision-preview0.697
- 3Reka-Flash0.5945
- 4claude-3-opus-202402290.5675
- 5yi-vl-plus0.536
- 6liuhaotian/llava-v1.6-34b0.5125
- 7claude-3-sonnet-202402290.5005
- 8claude-3-haiku-202403070.4175
- 9gemini-pro-vision0.395
- 10deepseek-ai/deepseek-vl-7b-chat0.394
- 11liuhaotian/llava-v1.6-vicuna-13b0.393
- 12THUDM/cogvlm-chat-hf0.368
- 13liuhaotian/llava-v1.6-vicuna-7b0.343
- 14idefics2-8b-chatty0.321
- 15Qwen/Qwen-VL-Chat0.2605
- 16liuhaotian/llava-v1.5-13b0.2375
- 17BAAI/Bunny-v1_0-3B0.228
- 18openbmb/MiniCPM-V0.2125
- 19bczhou/tiny-llava-v1-hf0.169
- 20unum-cloud/uform-gen2-qwen-500m0.1575