Reasoning

MMBench V1.1

Build MMBench_V11 response matrix from VLMEval/OpenVLMRecords

3,579items
251subjects
Apache-2.0license
reasoningdomain
textmodality
imagemodality
item-level responses released
Saturation status: Yes

Response matrix

Fit to width. Hover for subject & item; click a cell for details.

MMBench V1.1 response matrix: AI models (rows) against items (columns)
Correct (1)Incorrect (0)Unobserved

Scale: 1 = correct · 0 = incorrect

Subjects

  1. 1HunYuan-Standard-Vision0.9415
  2. 2SenseChat-Vision0.9165
  3. 3InternVL2_5-78B0.9154
  4. 4InternVL2_5-78B-MPO0.9149
  5. 5Qwen2.5-VL-72B0.9119
  6. 6Qwen2.5-VL-72B-Instruct0.9111
  7. 7InternVL2_5-38B-MPO0.9109
  8. 8InternVL2_5-38B0.91
  9. 9ChatGPT4o0.9089
  10. 10Step1o0.9086
  11. 11GLM4V_PLUS_202501110.9075
  12. 12DoubaoVL0.9071
  13. 13GPT4.50.9061
  14. 14ola0.9059
  15. 15Qwen-VL-Max-08090.904
  16. 16InternVL2-76B0.904
  17. 17Qwen2-VL-72B-Instruct0.9027
  18. 18GPT4o_202411200.9017
  19. 19TeleMM0.9016
  20. 20Ovis2-34B0.9012
  21. 21GLM4V_PLUS0.8998
  22. 22Taiyi0.8995
  23. 23GeminiFlash2-00.8991
  24. 24GeminiPro1-5-0020.8972
  25. 25GPT4o_202408060.8956
  26. 26BlueLM_V0.895
  27. 27InternVL2_5-26B0.8939
  28. 28GeminiPro2-00.8938
  29. 29InternVL2-40B0.8937
  30. 30GPT4o_HIGH0.8934
  31. 31llava_onevision_qwen2_72b_si0.8934
  32. 32Step1V0.8933
  33. 33InternVL2_5-26B-MPO0.8933
  34. 34Claude3-5V_Sonnet_202410220.892
  35. 35Ovis2-16B0.892
  36. 36llava_onevision_qwen2_72b_ov0.8913