Measurement Data Bank

A curated data bank of AI evaluation results from public benchmarks. Coverage spans multiple domains, such as reasoning, coding, agentic, multimodal, safety, and human preference benchmarks.

01Automatic Benchmarks

48

Benchmarks curated

123K

Evaluation items

13.4M

Model × item responses

Meta-analyses

Provider Coverage, Ranked By The Number Of Items Evaluated

Each provider's most-measured model, ranked by how many distinct benchmark items it has been evaluated on.

#Provider · most-measured modelItems asked
  1. 1OpenAI

    OpenAI

    GPT-4o

    79,267
  2. 2Anthropic

    Anthropic

    Claude 3.5 Sonnet

    76,047
  3. 3Mistral

    Mistral

    Mixtral 8x7B

    60,527
  4. 4Meta

    Meta

    Llama 3.1 8B

    56,593
  5. 5Google

    Google

    Gemini 2.0 Flash

    54,990
  6. 6DeepSeek

    DeepSeek

    V3

    46,756
  7. 7Alibaba

    Alibaba

    Qwen2.5 72B

    38,966
  8. 8A

    AllenAI

    OLMo 2 13B

    29,473
  9. 9S

    Stanford

    Alpaca 7B

    28,891
  10. 10Microsoft

    Microsoft

    Phi-4

    27,929
  11. 110

    01.AI

    Yi 1.5 34B

    26,529
  12. 12B

    ByteDance

    Doubao 1.5 Pro

    26,529
Domain Coverage, Ranked By The Number Of Items Evaluated

Benchmarks that span multiple domains count toward each, so the totals overlap rather than partition the shelf.

#DomainItems
  1. 1

    Reasoning

    12 benchmarks

    69,321
  2. 2

    Knowledge

    7 benchmarks

    61,488
  3. 3

    Science

    2 benchmarks

    29,617
  4. 4

    Law

    3 benchmarks

    17,510
  5. 5

    Multilingual

    4 benchmarks

    12,642
  6. 6

    Safety

    10 benchmarks

    11,691
  7. 7

    Medicine

    3 benchmarks

    9,238
  8. 8

    Agents & Tool Use

    10 benchmarks

    6,194
  9. 9

    General

    4 benchmarks

    5,248
  10. 10

    Software Engineering

    9 benchmarks

    2,607
  11. 11

    Mathematics

    5 benchmarks

    2,085
  12. 12

    ML Engineering

    1 benchmark

    212
  13. 13

    Cultural

    1 benchmark

    152
  14. 14

    NLP Tasks

    1 benchmark

    152
Benchmark Release Timeline

Every month-dated benchmark on this shelf, one dot per benchmark, stacked by release month and colored by primary domain.

SafetyMedicineLawMathematicsSoftware EngineeringAgents & Tool UseScienceMultilingualKnowledgeReasoningGeneral
2019 · 1
2020 · 2
2021 · 4
2022 · 1
2023 · 10
2024 · 11
2025 · 8
2026 · 10

Not shown (no month-dated release): AI2D_TEST.

Model Release Timeline

Every dated model evaluated on this shelf. One dot per model, stacked by release month and colored by provider.

OpenAIAnthropicMistralMetaGoogleDeepSeekAlibabaAllenAIStanfordMicrosoft01.AIByteDanceOther
2019 · 4
2020 · 6
2021 · 13
2022 · 20
2023 · 153
2024 · 268
2025 · 154
2026 · 76

130 models with no recorded release date are not shown.

Model × Domain Scores

Every model’s average score across each benchmark domain, over binary-scored benchmarks only. A benchmark that spans several domains counts toward each of them, so neighbouring columns can rest on the same responses. A hatched cell means the model was never measured in that domain.

Accuracy0 – 100
Not measured
Model
  1. 37
    75
    75
    55
    64
    68
    51
    62
    37
    82
    62
  2. 77
    72
    72
    77
    69
    66
    71
    23
    39
    71
    23
  3. 81
    69
    69
    51
    54
    53
    70
    46
    47
    70
    37
  4. 77
    78
    78
    54
    66
    64
    60
    51
    39
    31
    49
  5. 81
    72
    72
    49
    57
    54
    64
    29
    40
    68
    41
  6. 71
    75
    75
    72
    70
    72
    59
    11
    33
    41
  7. 80
    78
    89
    73
    59
    60
    37
    87
    52
  8. 73
    73
    60
    60
    62
    69
    60
    87
    48
  9. 79
    79
    63
    70
    63
    65
    13
    51
    77
  10. 75
    75
    52
    52
    53
    39
    45
    84
    52
  11. 67
    56
    56
    72
    44
    41
    51
    42
    78
  12. 75
    75
    47
    47
    50
    63
    42
    59
    42
  13. 71
    77
    77
    88
    85
    83
    50
    83
  14. 85
    74
    74
    85
    90
    81
    47
    72
  15. 67
    65
    65
    70
    73
    70
    33
    89
  16. 78
    71
    71
    79
    87
    70
    10
    37
  17. 73
    73
    62
    62
    63
    46
    75
    46
  18. 64
    66
    66
    66
    62
    59
    32
    84
  19. 74
    74
    9
    75
    62
    50
    77
    50
  20. 75
    75
    6
    71
    54
    42
    74
    42
  21. 62
    62
    45
    62
    55
    58
    0
    66
  22. 60
    60
    22
    32
    32
    62
    55
    56
  23. 74
    74
    38
    38
    41
    60
    12
    12
  24. 50
    50
    18
    25
    25
    56
    52
    49
  25. 77
    83
    76
    74
    71
    46
    81
  26. 76
    76
    76
    62
    66
    90
    59
  27. 73
    73
    69
    74
    63
    86
    58
  28. 74
    82
    74
    70
    67
    48
    79
  29. 77
    74
    84
    74
    51
    46
    61
  30. 69
    66
    66
    71
    78
    61
    34
  31. 73
    73
    54
    55
    60
    41
    74
  32. 72
    72
    45
    45
    47
    70
    64
  33. 80
    77
    89
    77
    22
    40
    22
  34. 72
    72
    35
    46
    47
    48
    59
  35. 73
    73
    69
    20
    27
    79
    27
  36. 62
    62
    21
    30
    30
    46
    71
  37. 68
    68
    2
    61
    15
    74
    15
  38. 76
    76
    43
    43
    45
    71
  39. 72
    72
    72
    31
    71
    31
  40. 72
    72
    62
    18
    66
    18
  41. 70
    70
    27
    27
    31
    63
  42. 41
    41
    41
    64
    32
    62
  43. 82
    91
    90
    91
    72
  44. 83
    89
    90
    89
    75
  45. 81
    88
    90
    88
    76
  46. 84
    89
    90
    89
    68
  47. 80
    86
    89
    86
    74
  48. 81
    87
    88
    87
    72
  49. 81
    87
    87
    87
    71
  50. 78
    87
    89
    87
    72
  51. 81
    89
    91
    89
    58
  52. 80
    87
    87
    87
    69
  53. 76
    86
    86
    86
    74
  54. 80
    86
    89
    86
    67
  55. 79
    85
    89
    85
    66
  56. 79
    88
    91
    88
    58
  57. 74
    86
    84
    86
    73
  58. 77
    84
    88
    84
    67
  59. 73
    86
    87
    86
    68
  60. 81
    96
    94
    96
    31
  61. 70
    88
    91
    88
    60
  62. 77
    84
    88
    84
    63
  63. 80
    87
    89
    87
    53
  64. 77
    85
    87
    85
    64
  65. 81
    89
    92
    89
    43
  66. 81
    85
    88
    85
    55
  67. 78
    83
    86
    83
    61
  68. 72
    83
    83
    83
    69
  69. 71
    85
    90
    85
    59
  70. 78
    89
    87
    89
    48
  71. 75
    81
    87
    81
    65
  72. 78
    85
    87
    85
    54
  73. 84
    89
    91
    89
    35
  74. 74
    86
    89
    86
    52
  75. 82
    85
    91
    85
    45
  76. 71
    83
    85
    83
    64
  77. 82
    79
    88
    79
    60
  78. 82
    88
    79
    88
    51
  79. 76
    82
    86
    82
    61
  80. 76
    82
    84
    82
    62
  81. 76
    81
    86
    81
    62
  82. 75
    83
    85
    83
    60
  83. 77
    82
    86
    82
    57
  84. 76
    82
    86
    82
    59
  85. 74
    85
    88
    85
    53
  86. 78
    83
    86
    83
    52
  87. 76
    84
    70
    84
    63
  88. 75
    82
    84
    82
    56
  89. 68
    83
    83
    83
    62
  90. 77
    88
    90
    88
    34
  91. 71
    79
    85
    79
    61
  92. 76
    84
    85
    84
    46
  93. 74
    81
    85
    81
    53
  94. 75
    83
    70
    83
    63
  95. 60
    88
    92
    88
    44
  96. 51
    84
    87
    84
    65
  97. 82
    83
    71
    83
    53
  98. 78
    84
    72
    84
    52
  99. 73
    79
    83
    79
    55
  100. 70
    77
    83
    77
    62
  101. 72
    81
    86
    81
    47
  102. 56
    86
    91
    86
    49
  103. 76
    78
    85
    78
    52
  104. 75
    81
    85
    81
    45
  105. 51
    82
    87
    82
    64
  106. 71
    81
    83
    81
    49
  107. 76
    78
    82
    78
    52
  108. 74
    79
    81
    79
    50
  109. 73
    76
    83
    76
    55
  110. 71
    81
    87
    81
    42
  111. 73
    77
    82
    77
    54
  112. 77
    81
    62
    81
    58
  113. 73
    80
    88
    80
    38
  114. 68
    75
    82
    75
    58
  115. 76
    79
    85
    79
    38
  116. 71
    78
    87
    78
    42
  117. 72
    77
    80
    77
    50
  118. 71
    75
    80
    75
    54
  119. 77
    74
    84
    74
    46
  120. 65
    76
    78
    76
    58
  121. 71
    78
    83
    78
    44
  122. 74
    78
    68
    78
    53
  123. 76
    81
    69
    81
    45
  124. 42
    61
    77
    75
    95
  125. 72
    75
    80
    75
    48
  126. 77
    80
    69
    80
    43
  127. 75
    76
    85
    76
    38
  128. 74
    76
    85
    76
    38
  129. 75
    78
    83
    78
    35
  130. 70
    78
    79
    78
    43
  131. 70
    74
    81
    74
    47
  132. 78
    82
    86
    82
    17
  133. 70
    77
    78
    77
    40
  134. 75
    74
    82
    74
    38
  135. 76
    79
    83
    67
    38
  136. 71
    74
    83
    74
    41
  137. 71
    75
    82
    75
    40
  138. 76
    72
    84
    72
    38
  139. 79
    78
    60
    78
    46
  140. 66
    71
    76
    71
    56
  141. 72
    71
    81
    71
    45
  142. 70
    73
    82
    73
    40
  143. 78
    76
    64
    76
    46
  144. 66
    75
    80
    75
    43
  145. 66
    75
    81
    75
    41
  146. 70
    71
    76
    71
    49
  147. 70
    73
    82
    73
    39
  148. 77
    79
    64
    79
    37
  149. 68
    75
    78
    75
    40
  150. 72
    72
    82
    72
    36
  151. 66
    78
    67
    78
    44
  152. 68
    74
    78
    74
    39
  153. 68
    70
    76
    70
    50
  154. 70
    78
    77
    78
    30
  155. 69
    74
    82
    74
    32
  156. 69
    74
    77
    74
    37
  157. 69
    69
    78
    69
    43
  158. 66
    70
    80
    70
    41
  159. 70
    73
    78
    73
    33
  160. 69
    69
    78
    69
    41
  161. 70
    70
    81
    70
    34
  162. 71
    67
    80
    67
    39
  163. 68
    73
    78
    73
    31
  164. 72
    69
    79
    69
    35
  165. 68
    70
    75
    70
    39
  166. 71
    74
    66
    74
    37
  167. 61
    72
    70
    72
    47
  168. 74
    69
    82
    69
    26
  169. 68
    70
    76
    70
    37
  170. 66
    72
    77
    72
    31
  171. 46
    74
    83
    74
    38
  172. 64
    69
    77
    69
    36
  173. 59
    71
    75
    71
    38
  174. 74
    63
    79
    63
    35
  175. 75
    66
    77
    66
    30
  176. 72
    72
    49
    72
    47
  177. 68
    65
    80
    65
    35
  178. 72
    65
    76
    65
    34
  179. 72
    64
    79
    64
    33
  180. 69
    66
    75
    66
    32
  181. 67
    69
    77
    69
    29
  182. 71
    65
    79
    65
    28
  183. 68
    67
    76
    67
    29
  184. 69
    64
    76
    64
    33
  185. 65
    65
    81
    65
    28
  186. 67
    64
    75
    64
    32
  187. 65
    66
    45
    78
    45
  188. 66
    63
    76
    63
    32
  189. 66
    64
    78
    64
    29
  190. 71
    61
    76
    61
    29
  191. 67
    62
    75
    62
    33
  192. 70
    60
    76
    60
    31
  193. 71
    60
    76
    60
    29
  194. 69
    61
    77
    61
    27
  195. 67
    60
    74
    60
    34
  196. 68
    61
    76
    61
    27
  197. 70
    65
    54
    65
    40
  198. 64
    59
    72
    59
    36
  199. 65
    68
    53
    68
    35
  200. 68
    69
    54
    69
    30
  201. 72
    58
    76
    58
    25
  202. 45
    66
    78
    66
    32
  203. 69
    57
    76
    57
    29
  204. 69
    56
    76
    56
    29
  205. 66
    58
    74
    58
    31
  206. 68
    57
    75
    57
    29
  207. 65
    57
    77
    57
    29
  208. 59
    59
    75
    59
    33
  209. 70
    57
    69
    57
    30
  210. 65
    58
    75
    58
    27
  211. 70
    56
    73
    56
    25
  212. 62
    56
    70
    56
    34
  213. 59
    59
    65
    59
    36
  214. 61
    57
    67
    57
    36
  215. 68
    56
    73
    56
    25
  216. 61
    54
    69
    54
    39
  217. 71
    71
    44
    43
    46
  218. 64
    51
    74
    51
    30
  219. 62
    53
    68
    53
    33
  220. 64
    49
    75
    49
    29
  221. 62
    62
    63
    40
    38
  222. 59
    49
    71
    49
    24
  223. 70
    53
    48
    53
    27
  224. 58
    55
    62
    55
    21
  225. 31
    58
    69
    58
    30
  226. 81
    13
    43
    64
    43
  227. 54
    46
    59
    46
    34
  228. 58
    47
    58
    47
    29
  229. 60
    42
    66
    42
    22
  230. 63
    41
    62
    41
    24
  231. 68
    38
    64
    38
    23
  232. 36
    59
    66
    59
    6
  233. 59
    42
    60
    42
    22
  234. 41
    41
    41
    61
    39
  235. 55
    39
    49
    39
    33
  236. 54
    41
    57
    41
    16
  237. 29
    29
    31
    59
    61
  238. 41
    46
    53
    46
    20
  239. 56
    38
    62
    38
    12
  240. 29
    29
    29
    64
    51
  241. 59
    33
    56
    33
    12
  242. 46
    46
    36
    36
    28
  243. 50
    30
    50
    30
    18
  244. 34
    49
    0
    0
    92
  245. 46
    27
    49
    27
    16
  246. 46
    31
    31
    31
    17
  247. 51
    26
    30
    26
    21
  248. 43
    27
    26
    27
    12
  249. 76
    77
    72
    97
  250. 74
    78
    73
    91
  251. 74
    63
    63
    97
  252. 56
    67
    73
    91
  253. 55
    71
    82
    79
  254. 78
    79
    79
    51
  255. 92
    74
    56
    56
  256. 82
    71
    53
    53
  257. 90
    86
    37
    37
  258. 54
    71
    72
    49
  259. 66
    64
    63
    40
  260. 90
    74
    26
    26
  261. 57
    57
    47
    56
  262. 44
    42
    58
    58
  263. 60
    60
    41
    40
  264. 61
    61
    42
    37
  265. 46
    85
    57
    5
  266. 50
    49
    59
    30
  267. 39
    39
    39
    70
  268. 35
    35
    35
    69
  269. 66
    20
    68
    20
  270. 42
    41
    53
    27
  271. 47
    64
    23
    23
  272. 24
    24
    24
    63
  273. 26
    26
    26
    52
  274. 19
    19
    19
    74
  275. 23
    23
    23
    57
  276. 55
    55
    4
    4
  277. 17
    17
    17
    66
  278. 19
    19
    19
    57
  279. 19
    19
    19
    31
  280. 11
    11
    11
    34
  281. 10
    10
    16
    28
  282. 81
    64
    83
  283. 83
    71
    71
  284. 81
    70
    69
  285. 72
    68
    68
  286. 61
    68
    80
  287. 69
    69
    69
  288. 47
    76
    85
  289. 82
    60
    62
  290. 41
    68
    95
  291. 67
    67
    67
  292. 71
    64
    64
  293. 66
    66
    66
  294. 65
    65
    65
  295. 64
    64
    67
  296. 64
    64
    64
  297. 71
    69
    43
  298. 60
    60
    60
  299. 60
    60
    60
  300. 60
    60
    60
  301. 70
    69
    40
  302. 46
    34
    98
  303. 64
    62
    39
  304. 68
    46
    51
  305. 49
    67
    48
  306. 73
    44
    44
  307. 53
    53
    53
  308. 77
    41
    41
  309. 47
    47
    64
  310. 50
    58
    51
  311. 61
    60
    35
  312. 70
    42
    42
  313. 56
    55
    41
  314. 75
    38
    38
  315. 50
    50
    50
  316. 58
    57
    33
  317. 55
    54
    37
  318. 49
    49
    49
  319. 57
    56
    33
  320. 56
    55
    33
  321. 57
    55
    32
  322. 53
    52
    38
  323. 54
    53
    35
  324. 23
    54
    64
  325. 68
    35
    35
  326. 55
    54
    30
  327. 51
    50
    36
  328. 51
    50
    36
  329. 75
    31
    31
  330. 54
    53
    29
  331. 45
    45
    45
  332. 45
    45
    45
  333. 49
    49
    35
  334. 50
    50
    33
  335. 49
    48
    32
  336. 49
    48
    33
  337. 49
    48
    31
  338. 49
    48
    30
  339. 65
    31
    31
  340. 47
    47
    33
  341. 46
    45
    36
  342. 47
    47
    33
  343. 46
    46
    33
  344. 48
    47
    29
  345. 37
    37
    49
  346. 70
    23
    30
  347. 58
    58
    5
  348. 46
    45
    30
  349. 71
    24
    24
  350. 42
    42
    34
  351. 43
    42
    30
  352. 42
    41
    32
  353. 42
    41
    31
  354. 42
    41
    32
  355. 38
    38
    38
  356. 74
    19
    19
  357. 77
    18
    18
  358. 61
    25
    25
  359. 40
    39
    30
  360. 29
    30
    47
  361. 72
    14
    20
  362. 28
    26
    52
  363. 40
    40
    26
  364. 49
    49
    7
  365. 40
    39
    24
  366. 38
    37
    27
  367. 38
    37
    25
  368. 33
    33
    33
  369. 49
    49
    2
  370. 39
    38
    23
  371. 48
    48
    4
  372. 6
    40
    49
  373. 36
    35
    23
  374. 5
    33
    54
  375. 40
    40
    8
  376. 43
    43
    3
  377. 31
    31
    25
  378. 71
    8
    8
  379. 28
    28
    28
  380. 28
    28
    27
  381. 40
    40
    3
  382. 27
    27
    28
  383. 28
    28
    25
  384. 27
    27
    24
  385. 27
    27
    24
  386. 26
    26
    26
  387. 29
    28
    21
  388. 26
    26
    26
  389. 25
    25
    27
  390. 27
    27
    24
  391. 26
    26
    25
  392. 27
    26
    24
  393. 26
    26
    25
  394. 26
    26
    24
  395. 25
    25
    27
  396. 25
    25
    24
  397. 26
    26
    24
  398. 26
    26
    23
  399. 26
    26
    24
  400. 25
    25
    25
  401. 26
    26
    24
  402. 26
    25
    24
  403. 25
    25
    25
  404. 69
    3
    3
  405. 25
    25
    24
  406. 26
    26
    21
  407. 25
    25
    24
  408. 23
    23
    28
  409. 26
    26
    22
  410. 26
    26
    21
  411. 25
    25
    23
  412. 24
    24
    25
  413. 25
    25
    23
  414. 27
    26
    20
  415. 25
    25
    22
  416. 26
    26
    21
  417. 26
    26
    20
  418. 25
    25
    23
  419. 25
    24
    23
  420. 24
    24
    23
  421. 23
    23
    26
  422. 26
    25
    21
  423. 26
    26
    20
  424. 25
    25
    21
  425. 25
    24
    22
  426. 23
    23
    23
  427. 24
    24
    22
  428. 23
    23
    22
  429. 22
    22
    22
  430. 27
    27
    13
  431. 0
    29
    36
  432. 48
    8
    8
  433. 62
    0
    0
  434. 21
    21
    21
  435. 25
    24
    10
  436. 19
    19
    19
  437. 18
    18
    18
  438. 17
    17
    17
  439. 97
    77
  440. 81
    81
  441. 90
    72
  442. 81
    81
  443. 80
    80
  444. 69
    89
  445. 79
    79
  446. 73
    84
  447. 86
    68
  448. 77
    77
  449. 56
    98
  450. 76
    76
  451. 55
    97
  452. 75
    75
  453. 56
    93
  454. 75
    75
  455. 58
    91
  456. 58
    91
  457. 74
    74
  458. 73
    73
  459. 73
    73
  460. 73
    73
  461. 71
    71
  462. 71
    71
  463. 70
    70
  464. 70
    70
  465. 70
    70
  466. 44
    96
  467. 70
    70
  468. 63
    76
  469. 69
    69
  470. 68
    68
  471. 68
    68
  472. 68
    68
  473. 59
    76
  474. 60
    74
  475. 67
    67
  476. 66
    66
  477. 66
    66
  478. 66
    66
  479. 65
    65
  480. 63
    63
  481. 63
    63
  482. 63
    63
  483. 61
    61
  484. 59
    59
  485. 59
    59
  486. 59
    59
  487. 58
    58
  488. 58
    58
  489. 57
    57
  490. 49
    62
  491. 55
    55
  492. 54
    54
  493. 54
    54
  494. 53
    53
  495. 53
    53
  496. 52
    52
  497. 52
    52
  498. 52
    52
  499. 39
    64
  500. 51
    51
  501. 51
    51
  502. 51
    51
  503. 50
    50
  504. 49
    49
  505. 49
    49
  506. 49
    49
  507. 47
    47
  508. 47
    47
  509. 53
    41
  510. 46
    46
  511. 46
    46
  512. 46
    46
  513. 45
    45
  514. 45
    45
  515. 45
    45
  516. 45
    45
  517. 45
    45
  518. 44
    44
  519. 43
    43
  520. 42
    42
  521. 42
    42
  522. 42
    42
  523. 41
    41
  524. 41
    41
  525. 40
    40
  526. 40
    40
  527. 38
    40
  528. 38
    38
  529. 73
    1
  530. 37
    37
  531. 37
    37
  532. 36
    36
  533. 35
    35
  534. 33
    33
  535. 33
    33
  536. 32
    32
  537. 27
    27
  538. 26
    26
  539. 26
    26
  540. 23
    23
  541. 14
    14
  542. 13
    13
  543. 1
    1
  544. 1
    1
  545. 1
    1
  546. 91
  547. 90
  548. 88
  549. 86
  550. 85
  551. 84
  552. 84
  553. 83
  554. 83
  555. 81
  556. 79
  557. 78
  558. 78
  559. 77
  560. 77
  561. 76
  562. 75
  563. 75
  564. 75
  565. 74
  566. 74
  567. 74
  568. 74
  569. 73
  570. 72
  571. 72
  572. 72
  573. 72
  574. 71
  575. 71
  576. 71
  577. 71
  578. 70
  579. 69
  580. 69
  581. 69
  582. 69
  583. 69
  584. 68
  585. 68
  586. 68
  587. 68
  588. 68
  589. 67
  590. 67
  591. 67
  592. 66
  593. 66
  594. 66
  595. 66
  596. 65
  597. 65
  598. 64
  599. 64
  600. 64
  601. 64
  602. 64
  603. 64
  604. 63
  605. 63
  606. 63
  607. 63
  608. 63
  609. 63
  610. 63
  611. 62
  612. 62
  613. 62
  614. 62
  615. 62
  616. 61
  617. 61
  618. 61
  619. 60
  620. 60
  621. 60
  622. 59
  623. 59
  624. 59
  625. 59
  626. 59
  627. 59
  628. 59
  629. 58
  630. 57
  631. 57
  632. 57
  633. 57
  634. 57
  635. 57
  636. 57
  637. 57
  638. 57
  639. 57
  640. 56
  641. 56
  642. 56
  643. 56
  644. 56
  645. 56
  646. 55
  647. 55
  648. 55
  649. 55
  650. 55
  651. 55
  652. 54
  653. 54
  654. 54
  655. 54
  656. 54
  657. 53
  658. 53
  659. 52
  660. 52
  661. 51
  662. 51
  663. 50
  664. 50
  665. 50
  666. 50
  667. 50
  668. 49
  669. 49
  670. 49
  671. 49
  672. 48
  673. 48
  674. 48
  675. 48
  676. 48
  677. 47
  678. 47
  679. 47
  680. 47
  681. 47
  682. 47
  683. 47
  684. 47
  685. 46
  686. 46
  687. 46
  688. 46
  689. 46
  690. 45
  691. 45
  692. 45
  693. 45
  694. 45
  695. 44
  696. 43
  697. 43
  698. 43
  699. 43
  700. 43
  701. 41
  702. 40
  703. 39
  704. 39
  705. 39
  706. 39
  707. 39
  708. 38
  709. 38
  710. 38
  711. 38
  712. 37
  713. 37
  714. 37
  715. 36
  716. 36
  717. 35
  718. 35
  719. 35
  720. 35
  721. 34
  722. 34
  723. 33
  724. 32
  725. 31
  726. 30
  727. 30
  728. 29
  729. 29
  730. 28
  731. 27
  732. 27
  733. 26
  734. 26
  735. 23
  736. 22
  737. 21
  738. 20
  739. 20
  740. 20
  741. 19
  742. 15
  743. 14
  744. 14
  745. 12
  746. 11
  747. 10
  748. 10
  749. 9
  750. 9
  751. 8
  752. 8
  753. 8
  754. 8
  755. 7
  756. 7
  757. 7
  758. 7
  759. 7
  760. 7
  761. 7
  762. 7
  763. 6
  764. 6
  765. 6
  766. 6
  767. 6
  768. 6
  769. 6
  770. 6
  771. 6
  772. 6
  773. 5
  774. 5
  775. 4
  776. 4
  777. 3
  778. 2
  779. 2
  780. 2
  781. 2
  782. 2
  783. 1
  784. 1
  785. 0
  786. 0
Domain
Scale
Responses
Saturation
Released
Min
Sort by
What do the card tags mean?
item-level responses released
The benchmark’s per-model, per-item response data is fully public and downloadable from Hugging Face. Every (model × item) cell in the response matrix is available, not just aggregate scores. Benchmarks without this tag only provide item content and subject lists, or aggregate-level results.
Nitems / subjects / attacks
The count of distinct evaluation items (prompts/questions), AI subjects (models/agents) evaluated, or adversarial attacks in the bank’s standardized matrix for this benchmark.
Saturation status: Yes
The top-3 models average ≥ 90% on this benchmark, suggesting the task may be approaching a performance ceiling.
Saturation status: No
No: top models still score below 90%, so meaningful headroom remains. Benchmarks without enough item-level response data to judge saturation carry no saturation tag.

2026

10
ARC-AGI-3 benchmark
A
MIT2026

ARC-AGI-3

ARC-AGI-3 drops an agent into novel grid worlds with no instructions, so it must infer the mechanics, the goal and the win condition purely by acting, and records whether it completes each level.

117items
5AI models
item-level responses released
Saturation status: No
DeepSWE benchmark
D
Apache-2.02026

DeepSWE

DeepSWE tests whether coding agents can complete long horizon software engineering tasks written from scratch against active open source repositories, graded by hand written functional verifiers the agent never sees rather than by tests shipped with an existing fix.

113items
68AI models
item-level responses released
Saturation status: No
FrontierOR benchmark
Singapore-MIT Alliance for Research and TechnologyMIT
cc-by-4.02026

FrontierOR

FrontierOR asks models to read an operations research problem stated in plain language and write an optimization program, then checks how many large instances that program solves feasibly against an expert verified reference.

179items
7AI models
item-level responses released
Saturation status: No
IKP benchmark
P
CC-BY-4.02026

IKP

IKP asks models factual questions that range from things almost anyone knows to extreme long tail trivia, counting refusals as wrong. Because accuracy tracks model size closely, it doubles as a way to estimate the parameter count of a closed model.

1,400items
201AI models
item-level responses released
Saturation status: No
MathArena Platform benchmark
ETH Zurich
MIT2026

MathArena Platform

MathArena Platform asks models to answer research mathematics questions drawn from arXiv papers, write formal proofs in Lean that a proof checker verifies, and recognize when a stated claim is deliberately false.

422items
45AI models
item-level responses released
Saturation status: No
OSWorld 2.0 benchmark
University of Hong Kong
apache-2.02026

OSWorld 2.0

OSWorld 2.0 tests whether computer use agents can finish long real world tasks on a live Ubuntu desktop and its web applications, working from screenshots, with the final state of the machine checked against every task requirement.

108items
8AI models
item-level responses released
Saturation status: No
ProgramBench benchmark
MetaStanford University
MIT2026

ProgramBench

ProgramBench asks an agent to rebuild a working program from nothing but its compiled binary and the documentation shipped with it, scoring the result by how much of a hidden behavioural test suite the rebuilt program passes.

200items
12AI models
item-level responses released
Saturation status: No
tau-Voice benchmark
Sierra
MIT2026

tau-Voice

tau-Voice tests whether voice agents can handle live customer service calls in retail, airline and telecom settings, where a call passes only if the database ends in the required state and the agent conveys the required information.

278items
6AI models
item-level responses released
Saturation status: No
Terminal-Bench benchmark
Stanford UniversityLaude Institute
Apache-2.02026

Terminal-Bench

Terminal-Bench tests whether agents working in a command line environment can carry out real engineering tasks such as building software from source, fixing a broken script or analyzing a dataset, checked by the task's own test suite.

89items
192AI models
item-level responses released
Saturation status: Yes
Terminal-Bench 2.1 benchmark
Stanford UniversityLaude Institute
Apache-2.02026

Terminal-Bench 2.1

Terminal-Bench 2.1 tests whether an agent can finish a real task inside a Linux container it drives through a shell, passing only when a test suite run against the state it leaves behind succeeds entirely.

89items
7AI models
item-level responses released
Saturation status: No

2025

8
GenAI Learning benchmark
University of Pennsylvania
unknown2025

GenAI Learning

GenAI Learning measures whether high school students answer maths practice problems correctly when working alone, with an unrestricted chatbot, or with a tutor bot that refuses to give the full solution. A language model also answers the same problems on its own.

57items
944AI models
item-level responses released
Saturation status: Yes
ICPC-2 Code Selector benchmark
University of São Paulo
MIT2025

ICPC-2 Code Selector

ICPC-2 Code Selector tests whether models can pick the primary care code that best matches a short clinical expression in Brazilian Portuguese from candidates returned by a search engine, or correctly abstain when none match.

2,175items
40AI models
item-level responses released
Saturation status: No
MathArena benchmark
ETH Zurich
MIT2025

MathArena

MathArena scores models on problems from recent mathematics competitions, checking the final answer where there is one and having judges grade written proofs criterion by criterion where the problem asks for a proof.

427items
104AI models
item-level responses released
MMDocRAG benchmark
Huawei
CC-BY-4.02025

MMDocRAG

MMDocRAG tests whether models can answer questions about documents by weaving retrieved text and images into a cited answer, which a judge model rates for fluency, citation quality, coherence, reasoning and factuality.

2,000items
68AI models
item-level responses released
Saturation status: No
REAL benchmark
The AGI CompanyMercor
unknown2025

REAL

REAL tests whether web agents can complete user tasks on deterministic clones of real websites, counting an attempt as successful only when every one of the task's automatic checks passes.

233items
44AI models
item-level responses released
Saturation status: No
ResearchCodeBench benchmark
Stanford University
MIT2025

ResearchCodeBench

ResearchCodeBench tests whether models can implement a missing piece of a recent machine learning paper's code, given the paper itself and the surrounding file, with curated tests deciding whether the filled in code works.

212items
32AI models
item-level responses released
Saturation status: No
SuperGPQA benchmark
M-A-PByteDance
CC-BY-NC-SA-4.02025

SuperGPQA

SuperGPQA asks graduate level multiple choice questions drawn from a wide range of academic disciplines, counting an answer correct only when the model picks the intended option.

26,529items
49AI models
item-level responses released
Saturation status: No
tau2-bench benchmark
Sierra
MIT2025

tau2-bench

tau2-bench puts a model in the role of a customer support agent for an airline, retail or telecom account, using the company tools and database while talking to a simulated customer it must also guide through steps only they can perform.

423items
15AI models
item-level responses released
Saturation status: No

2024

11
AfriMed-QA benchmark
IGeorgia Institute of Technology
CC-BY-NC-SA-4.02024

AfriMed-QA

AfriMed-QA asks models multiple choice medical questions set in African healthcare contexts, spanning many clinical specialties, and scores whether the option a model picks matches the correct one.

6,911items
31AI models
item-level responses released
Saturation status: No
Alignment Faking (RL) benchmark
Anthropic
CC-BY-4.02024

Alignment Faking (RL)

Alignment Faking (RL) records whether a model pretends to go along with a request it would otherwise refuse, and whether that changes with how far its training has progressed and which kind of user it is told it is talking to.

244items
155AI models
item-level responses released
Saturation status: Yes
BFCL benchmark
UC Berkeley
Apache-2.02024

BFCL

BFCL tests whether models can answer a user request by calling the right function with the right arguments, including requests that no available function can satisfy and ones spread over several conversational turns.

4,133items
93AI models
item-level responses released
Saturation status: No
CharXiv benchmark
Princeton University
CC-BY-SA-4.02024

CharXiv

CharXiv tests whether multimodal models can read real charts from scientific papers, answering both short descriptive questions and an open ended reasoning question, with a judge model grading the descriptive answers.

5,000items
37AI models
item-level responses released
Saturation status: No
AIR-Bench 2024 benchmark
UC Berkeley
Apache-2.02024

AIR-Bench 2024

AIR-Bench tests whether models give safe replies to prompts written against a taxonomy of risks drawn from government regulations and company policies, with every reply marked safe or unsafe.

5,692items
66AI models
item-level responses released
Saturation status: Yes
HarmBench benchmark
University of Illinois Urbana-ChampaignCenter for AI Safety
Apache-2.02024

HarmBench

HarmBench measures whether a model refuses harmful requests, covering behaviors such as cybercrime, misinformation and chemical or biological harm, with a classifier judging each reply as safe or unsafe.

400items
87AI models
item-level responses released
Saturation status: Yes
LexEval benchmark
Tsinghua University
MIT2024

LexEval

LexEval tests how well models answer Chinese legal multiple-choice questions covering legal knowledge, reasoning, discrimination and ethics, grading each answer by exact match of the extracted option letters against the gold answer using the official LexEval grader.

12,468items
38AI models
item-level responses released
Saturation status: No
LiveCodeBench benchmark
UC Berkeley
CC-BY-4.02024

LiveCodeBench

LiveCodeBench tests whether models can solve competitive programming problems by writing a program that produces the expected output on the problem's test cases.

1,055items
72AI models
item-level responses released
Saturation status: No
MMLU-Pro benchmark
University of Waterloo
MIT2024

MMLU-Pro

MMLU-Pro poses multiple choice exam questions across fields such as law, physics and psychology, using a longer list of answer options and harder questions than the original MMLU.

13,542items
48AI models
item-level responses released
Saturation status: No
OSWorld benchmark
University of Hong Kong
Apache-2.02024

OSWorld

OSWorld tests whether agents can complete real tasks on an Ubuntu desktop driving applications such as Chrome, LibreOffice and VS Code, with each task graded by a scorer that can award partial credit.

369items
67AI models
item-level responses released
Saturation status: No
SWE-bench Java benchmark
Chinese Academy of Sciences
Apache-2.02024

SWE-bench Java

SWE-bench Java tests whether coding agents can resolve real GitHub issues in Java repositories by submitting a code patch, with each issue recorded as resolved or unresolved.

170items
54AI models
item-level responses released
Saturation status: No

2023

10
SimpleSafetyTests benchmark
Patronus AI
Apache-2.02023

SimpleSafetyTests

SimpleSafetyTests checks whether a model responds safely to blunt requests touching severe harms such as self harm, violence and illegal activity, with a classifier marking each reply safe or unsafe.

100items
87AI models
item-level responses released
Saturation status: Yes
HELM Thai Exam benchmark
SCB 10X
Apache-2.02023

HELM Thai Exam

HELM Thai Exam asks models questions drawn from Thai standardized examinations, written in Thai, and counts an answer correct only when it exactly matches the official answer.

561items
42AI models
item-level responses released
Saturation status: No
XSTest benchmark
Bocconi UniversityUniversity of Oxford
Apache-2.02023

XSTest

XSTest checks whether models refuse harmless prompts that merely sound dangerous, pairing them with genuinely unsafe lookalikes so that both excessive caution and unsafe compliance show up.

450items
87AI models
item-level responses released
Saturation status: Yes
LawBench benchmark
Nanjing UniversityAmazon
Apache-2.02023

LawBench

LawBench measures how well models handle Chinese legal work; this curation ingests the six subtasks graded by per-item binary accuracy — judicial-exam and case-analysis multiple choice, dispute-focus and consultation classification, argumentation mining, and crime-amount extraction — regraded from the released raw predictions with the official evaluation code.

2,995items
51AI models
item-level responses released
Saturation status: No
LegalBench benchmark
Stanford University
Apache-2.02023

LegalBench

LegalBench tests whether models can answer legal reasoning questions, with a response counted correct when it matches the reference answer closely enough to pass an approximate string match.

2,047items
31AI models
item-level responses released
Saturation status: No
MathVista MINI benchmark
UCLA
CC-BY-SA-4.02023

MathVista MINI

MathVista MINI tests whether vision language models can answer mathematical questions about an accompanying image, some multiple choice and some free response, with an answer counted correct when it matches the ground truth.

1,000items
263AI models
item-level responses released
Saturation status: No
MMBench V1.1 benchmark
S
Apache-2.02023

MMBench V1.1

MMBench V1.1 tests whether vision language models can answer multiple choice questions about a picture, repeating each question with its answer options rotated so a model cannot lean on option position.

4,876items
251AI models
item-level responses released
Saturation status: Yes
MME benchmark
Nanjing UniversityTencent
unknown2023

MME

MME asks vision language models yes or no questions about an image, covering both perception and cognition, and scores an answer correct when it matches the reference.

1,983items
232AI models
item-level responses released
Saturation status: Yes
MMMU (dev+val) benchmark
IUniversity of Waterloo
Apache-2.02023

MMMU (dev+val)

MMMU asks models college level multiple choice questions that pair text with images, scoring an answer correct only when the option it picks matches the correct one.

896items
253AI models
item-level responses released
Saturation status: No
SWE-bench Verified benchmark
Princeton University
MIT2023

SWE-bench Verified

SWE-bench Verified tests whether coding agents can resolve real issues reported against open source software projects by submitting a code patch that the benchmark's automatic evaluation accepts as a fix.

500items
134AI models
item-level responses released
Saturation status: No

Earlier (pre-2023)

9
AI2D_TEST benchmark
Allen Institute for AI
CC-BY-SA-4.02016

AI2D_TEST

AI2D_TEST tests whether vision language models can answer multiple choice questions about science diagrams, scored by whether the answer letter they give matches the correct choice.

3,088items
254AI models
item-level responses released
Saturation status: Yes
ARC-AGI benchmark
Google
Apache-2.02019

ARC-AGI

ARC-AGI tests whether a model can infer the transformation rule behind a handful of example input and output grids, then produce exactly the right output grid for a new input.

520items
74AI models
item-level responses released
Saturation status: Yes
HAIID benchmark
Stanford University
MIT2021

HAIID

HAIID measures whether people revise their answers on classification tasks, such as spotting sarcasm or judging skin lesion photographs, after seeing advice presented as coming either from an AI algorithm or from a human peer. The advice itself is identical in both cases.

152items
1,127AI models
item-level responses released
Saturation status: Yes
Anthropic Red Team benchmark
Anthropic
Apache-2.02022

Anthropic Red Team

Anthropic Red Team tests whether models avoid complying with adversarial prompts taken from human red teaming conversations, with a safety classifier judging each response as safe or unsafe.

995items
87AI models
item-level responses released
Saturation status: Yes
BBQ benchmark
New York University
Apache-2.02021

BBQ

BBQ tests whether models pick the correct answer to multiple choice questions about people from different social groups, asking each question once with too little context to answer it and once with the context that settles it.

999items
87AI models
item-level responses released
Saturation status: Yes
BOLD benchmark
AmazonUC Santa Barbara
Apache-2.02021

BOLD

BOLD measures whether models produce toxic text when asked to continue open ended prompts, marking a prompt as toxic when any of the model's continuations comes back toxic.

994items
42AI models
item-level responses released
Saturation status: No
RealToxicityPrompts benchmark
University of Washington
Apache-2.02020

RealToxicityPrompts

RealToxicityPrompts measures how often a model continues a sentence fragment with text a toxicity classifier flags as toxic, using prompts drawn from both toxic and innocuous starting text.

1,000items
42AI models
item-level responses released
Saturation status: No
MMLU benchmark
UC BerkeleyColumbia University
MIT2020

MMLU

MMLU tests broad knowledge and reasoning with multiple choice exam questions spanning school and professional topics in the sciences, humanities and social sciences, scored by whether the model picks the right option.

13,937items
150AI models
item-level responses released
Saturation status: No
TruthfulQA-MC benchmark
University of OxfordOpenAI
Apache-2.02021

TruthfulQA-MC

TruthfulQA-MC tests whether models select the single true answer rather than a plausible falsehood, on questions built around misconceptions that people commonly repeat.

817items
150AI models
item-level responses released
Saturation status: No

02Human-centered AI Benchmarks

Every response is scored by a judgment: a human rating, a model rating, or a verdict from a judge that is itself the subject under evaluation. Scores therefore carry judge disagreement in addition to subject ability. Pairwise benchmarks are indexed by model pair rather than by shared item, so they are listed separately.

13

Benchmarks curated

302K

Evaluation items

2.1M

Judged responses

Meta-analyses

Provider Coverage, Ranked By The Number Of Items Evaluated

Each provider's most-measured model, ranked by how many distinct benchmark items it has been evaluated on.

#Provider · most-measured modelItems asked
  1. 1Anthropic

    Anthropic

    Claude Opus 4

    33,516
  2. 2Google

    Google

    Gemini 2.5 Flash

    30,627
  3. 3Alibaba

    Alibaba

    Qwen3 235B A22B

    27,720
  4. 4OpenAI

    OpenAI

    GPT-4o

    23,265
  5. 5Meta

    Meta

    Llama 4 Maverick

    21,822
  6. 6Mistral

    Mistral

    Medium 3

    18,160
  7. 7xAI

    xAI

    Grok 3 Mini

    17,330
  8. 8Cohere

    Cohere

    Command A

    13,576
  9. 9DeepSeek

    DeepSeek

    R1

    13,028
  10. 10Amazon

    Amazon

    Nova Pro

    12,686
  11. 11M

    MiniMax

    M1

    10,848
  12. 12A

    AllenAI

    Tulu 2 70B

    10,163
Domain Coverage, Ranked By The Number Of Items Evaluated

Benchmarks that span multiple domains count toward each, so the totals overlap rather than partition the shelf.

#DomainItems
  1. 1

    Preference

    7 benchmarks

    263,498
  2. 2

    Software Engineering

    1 benchmark

    31,930
  3. 3

    Agents & Tool Use

    1 benchmark

    31,930
  4. 4

    Reward Modeling

    3 benchmarks

    5,159
  5. 5

    General

    1 benchmark

    764
  6. 6

    Safety

    1 benchmark

    450
  7. 7

    Multilingual

    1 benchmark

    120
Benchmark Release Timeline

Every month-dated benchmark on this shelf, one dot per benchmark, stacked by release month and colored by primary domain.

SafetySoftware EngineeringMultilingualPreferenceReward ModelingGeneral
2021 · 0
2022 · 0
2023 · 2
2024 · 7
2025 · 3
2026 · 1
Model Release Timeline

Every dated model evaluated on this shelf. One dot per model, stacked by release month and colored by provider.

AnthropicGoogleAlibabaOpenAIMetaMistralxAICohereDeepSeekAmazonMiniMaxAllenAIOther
2021 · 1
2022 · 1
2023 · 62
2024 · 93
2025 · 37
2026 · 9

5 models with no recorded release date are not shown.

Model × Domain Scores

Every model’s average score across each benchmark domain, over binary-scored benchmarks only. A benchmark that spans several domains counts toward each of them, so neighbouring columns can rest on the same responses. A hatched cell means the model was never measured in that domain.

Accuracy0 – 100
Not measured
Model
  1. 76
    76
  2. 74
    74
  3. 67
    67
  4. 56
    56
  5. 51
    60
  6. 52
    52
  7. 50
    51
  8. 48
    48
  9. 45
    45
  10. 43
    43
  11. 43
    43
  12. 42
    42
  13. 40
    40
  14. 50
    29
  15. 51
    22
  16. 50
    18
  17. 50
    17
  18. 51
    17
  19. 25
    25
  20. 61
  21. 53
  22. 50
  23. 50
  24. 50
  25. 50
  26. 49
  27. 31
  28. 28
  29. 7
  30. 7
  31. 7
  32. 6
  33. 4
Domain
Scale
Responses
Saturation
Released
Min
Sort by

2026

1
SWE-chat benchmark
Stanford University
ODC-By-1.02026

SWE-chat

SWE-chat measures whether developers accept what a coding agent did during real sessions on their own repositories, with a judge model reading each following prompt to decide whether the developer pushed back.

31,930items
34AI models
item-level responses released
Saturation status: Yes

2025

3
Arena 140K benchmark
UC Berkeley
CC-BY-4.02025

Arena 140K

Arena 140K records which of two models a person preferred after reading both replies to a prompt they wrote themselves, with ties and mutual rejection both counted as outcomes.

128,287items
53AI models
item-level responses released
Saturation status: No
Prompt-to-Leaderboard benchmark
UC Berkeley
CC-BY-4.02025

Prompt-to-Leaderboard

Prompt-to-Leaderboard ranks models on individual prompts using head to head battles in which a person sees two models' answers to the same prompt and votes for one or calls it a tie.

128,287items
53AI models
item-level responses released
Saturation status: No
RewardBench 2 benchmark
Allen Institute for AI
ODC-BY2025

RewardBench 2

RewardBench 2 tests whether reward models rank the intended best response above the alternatives, on prompts covering areas such as factuality, math and safety.

1,824items
188AI models
item-level responses released
Saturation status: No

2024

7
Arena-Hard-Auto benchmark
UC Berkeley
Apache-2.02024

Arena-Hard-Auto

Arena-Hard-Auto rates how well models answer hard user requests, many of them programming and technical questions, by having a strong model judge each answer against a fixed baseline model's answer to the same request.

500items
72AI models
item-level responses released
Saturation status: No
BiGGen-Bench benchmark
Carnegie Mellon UniversityKAIST
CC-BY-SA-4.02024

BiGGen-Bench

BiGGen-Bench rates model answers to prompts spanning reasoning, instruction following, safety, planning and other capabilities, with strong models and in some cases human annotators scoring each answer against a rubric.

764items
103AI models
item-level responses released
Saturation status: No
JudgeBench benchmark
UC Berkeley
MIT2024

JudgeBench

JudgeBench tests whether a model acting as a judge picks the correct response out of a pair, on comparisons drawn from knowledge, reasoning, math and coding, with each pair shown in both orders.

350items
33AI models
item-level responses released
Saturation status: No
RewardBench benchmark
Allen Institute for AI
ODC-BY2024

RewardBench

RewardBench tests whether reward models prefer the better of two candidate responses to a prompt, over prompts covering chat, safety and reasoning.

2,985items
151AI models
item-level responses released
Saturation status: Yes
SORRY-Bench benchmark
Princeton University
other2024

SORRY-Bench

SORRY-Bench measures whether models comply with or refuse unsafe requests, with human annotators labeling each reply, including when the request is rewritten in styles such as role play, slang or a cipher.

450items
31AI models
item-level responses released
Saturation status: Yes
Tengu-Bench benchmark
L
Apache-2.02024

Tengu-Bench

Tengu-Bench rates how well models answer Japanese tasks spanning categories such as table reading and reasoning, with several language models acting as judges and scoring each answer against a written rubric.

120items
647AI models
item-level responses released
Saturation status: No
WildBench benchmark
Allen Institute for AI
CC-BY-4.02024

WildBench

WildBench rates how well models handle challenging tasks taken from real user conversations, with a strong model acting as the judge and rating each answer on a numeric scale.

1,024items
71AI models
item-level responses released
Saturation status: No

2023

2
MT-Bench benchmark
UC Berkeley
CC-BY-4.02023

MT-Bench

MT-Bench rates how well chat models handle open ended questions asked over two conversational turns, with a strong model acting as the judge and scoring each answer on its own.

160items
34AI models
item-level responses released
Saturation status: No
Preference Dissection benchmark
Shanghai Jiao Tong University
CC-BY-NC-4.02023

Preference Dissection

Preference Dissection records which of two candidate replies to the same prompt each model judge prefers, turning the judges themselves into the systems being measured.

4,890items
32AI models
item-level responses released
Saturation status: No

Citation

Cite this work

If you use the data we curated, please cite the following reference. The curation is released under CC BY-SA 4.0; individual benchmarks retain their upstream licenses.

@misc{measurementdb2026,
title = {The AI Measurement Data Bank},
author = {Truong, Nhi and Truong, Sang T. and Koyejo, Sanmi},
year = {2026},
howpublished = {\url{https://aimslab.stanford.edu/measurement-db}},
note = {AIMS Lab, Stanford University}
}