Reasoning
MMBench V1.1
MMBench V1.1 tests whether vision language models can answer multiple choice questions about a picture, repeating each question with its answer options rotated so a model cannot lean on option position.
3,579items
251subjects
Apache-2.0license
reasoningdomain
textmodality
imagemodality
item-level responses released
Saturation status: Yes
Response matrix
Rasch analysis p = σ(θ − z + c)
1,180,617 responses, 80/20 split over cells · 251 subjects · 3,579 items · 6 conditions
AUC train
0.905
AUC test
0.903
Loading session strips…
Correct (1)Incorrect (0)Unobserved
Scale: 1 = correct · 0 = incorrect