General

MMBench V1.1

Build MMBench_V11 response matrix from VLMEval/OpenVLMRecords

3,579items
251subjects
Apache-2.0license
reasoningdomain
textmodality
imagemodality
item-level responses released
Saturation status: Unknown

Response matrix

Fit to width. Hover for subject & item; click a cell for details.

MMBench V1.1 response matrix: AI models (rows) against items (columns)
Correct (1)Incorrect (0)Unobserved

Scale: 1 = correct · 0 = incorrect

Subjects

  1. 1HunYuan-Standard-Vision0.9424
  2. 2SenseChat-Vision0.9174
  3. 3InternVL2_5-78B0.9169
  4. 4InternVL2_5-78B-MPO0.9165
  5. 5Qwen2.5-VL-72B0.9151
  6. 6Qwen2.5-VL-72B-Instruct0.9139
  7. 7ChatGPT4o0.9131
  8. 8Step1o0.9126
  9. 9InternVL2_5-38B-MPO0.912
  10. 10InternVL2_5-38B0.9116
  11. 11DoubaoVL0.9113
  12. 12GLM4V_PLUS_202501110.9098
  13. 13GPT4.50.9096
  14. 14ola0.9085
  15. 15GPT4o_202411200.9062
  16. 16InternVL2-76B0.9048
  17. 17Qwen-VL-Max-08090.9041
  18. 18Ovis2-34B0.904
  19. 19Qwen2-VL-72B-Instruct0.903
  20. 20GLM4V_PLUS0.9025
  21. 21Taiyi0.9022
  22. 22TeleMM0.9022
  23. 23GeminiFlash2-00.9014
  24. 24GeminiPro1-5-0020.9012
  25. 25GPT4o_202408060.9005
  26. 26GPT4o_HIGH0.8984
  27. 27llava_onevision_qwen2_72b_si0.8964
  28. 28GeminiPro2-00.8963
  29. 29BlueLM_V0.8958
  30. 30Step1V0.8953
  31. 31InternVL2-40B0.8946
  32. 32Ovis2-16B0.8942
  33. 33abab7-preview0.8934
  34. 34InternVL2_5-26B-MPO0.8929
  35. 35InternVL2_5-26B0.8927
  36. 36llava_onevision_qwen2_72b_ov0.8925