Skip to main content

Search AIMS

Find pages, publications, events, course projects, and software. Results update as you type; press Enter to open the first result.

Reasoning

MMBench V1.1

MMBench V1.1 tests whether vision language models can answer multiple choice questions about a picture, repeating each question with its answer options rotated so a model cannot lean on option position.

3,579items
251subjects
Apache-2.0license
reasoningdomain
textmodality
imagemodality
item-level responses released
Saturation status: Yes

Response matrix

Rasch analysis p = σ(θ − z + c)

1,180,617 responses, 80/20 split over cells · 251 subjects · 3,579 items · 6 conditions

AUC train
0.905
AUC test
0.903

Loading session strips…

Correct (1)Incorrect (0)Unobserved

Scale: 1 = correct · 0 = incorrect