General
MME
MME asks vision language models yes or no questions about an image, covering both perception and cognition, and scores an answer correct when it matches the reference.
1,983items
232subjects
unknownlicense
generaldomain
textmodality
imagemodality
item-level responses released
Saturation status: Yes
Response matrix
Rasch analysis p = σ(θ − z + c)
535,810 responses, 80/20 split over cells · 232 subjects · 1,983 items · 14 conditions
1 of 232 subjects answered every item alike, so its θ is unbounded and sits pegged to the column edge.
AUC train
0.854
AUC test
0.846
Loading session strips…
Correct (1)Incorrect (0)Unobserved
Scale: 1 = correct · 0 = incorrect