Reasoning
MMBench V1.1
Build MMBench_V11 response matrix from VLMEval/OpenVLMRecords
3,579items
251subjects
Apache-2.0license
reasoningdomain
textmodality
imagemodality
item-level responses released
Saturation status: Yes
Response matrix
Fit to width. Hover for subject & item; click a cell for details.
_trial_1.png)
Correct (1)Incorrect (0)Unobserved
Scale: 1 = correct · 0 = incorrect
Subjects
- 1HunYuan-Standard-Vision0.9415
- 2SenseChat-Vision0.9165
- 3InternVL2_5-78B0.9154
- 4InternVL2_5-78B-MPO0.9149
- 5Qwen2.5-VL-72B0.9119
- 6Qwen2.5-VL-72B-Instruct0.9111
- 7InternVL2_5-38B-MPO0.9109
- 8InternVL2_5-38B0.91
- 9ChatGPT4o0.9089
- 10Step1o0.9086
- 11GLM4V_PLUS_202501110.9075
- 12DoubaoVL0.9071
- 13GPT4.50.9061
- 14ola0.9059
- 15Qwen-VL-Max-08090.904
- 16InternVL2-76B0.904
- 17Qwen2-VL-72B-Instruct0.9027
- 18GPT4o_202411200.9017
- 19TeleMM0.9016
- 20Ovis2-34B0.9012
- 21GLM4V_PLUS0.8998
- 22Taiyi0.8995
- 23GeminiFlash2-00.8991
- 24GeminiPro1-5-0020.8972
- 25GPT4o_202408060.8956
- 26BlueLM_V0.895
- 27InternVL2_5-26B0.8939
- 28GeminiPro2-00.8938
- 29InternVL2-40B0.8937
- 30GPT4o_HIGH0.8934
- 31llava_onevision_qwen2_72b_si0.8934
- 32Step1V0.8933
- 33InternVL2_5-26B-MPO0.8933
- 34Claude3-5V_Sonnet_202410220.892
- 35Ovis2-16B0.892
- 36llava_onevision_qwen2_72b_ov0.8913