
Measurement Data Bank
A curated data bank of AI evaluation results from public benchmarks. Coverage spans multiple domains, such as reasoning, coding, agentic, multimodal, safety, and human preference benchmarks.
Maintained by Nhi Truong, Sang Truong, and Sanmi Koyejo
01Automatic Benchmarks
48
Benchmarks curated
123K
Evaluation items
13.4M
Model × item responses
Meta-analyses
Provider Coverage, Ranked By The Number Of Items Evaluated
Each provider's most-measured model, ranked by how many distinct benchmark items it has been evaluated on.
- 1
OpenAI
GPT-4o
79,267 - 2
Anthropic
Claude 3.5 Sonnet
76,047 - 3
Mistral
Mixtral 8x7B
60,527 - 4
Meta
Llama 3.1 8B
56,593 - 5
Google
Gemini 2.0 Flash
54,990 - 6
DeepSeek
V3
46,756 - 7
Alibaba
Qwen2.5 72B
38,966 - 8A
AllenAI
OLMo 2 13B
29,473 - 9S
Stanford
Alpaca 7B
28,891 - 10
Microsoft
Phi-4
27,929 - 110
01.AI
Yi 1.5 34B
26,529 - 12B
ByteDance
Doubao 1.5 Pro
26,529
Domain Coverage, Ranked By The Number Of Items Evaluated
Benchmarks that span multiple domains count toward each, so the totals overlap rather than partition the shelf.
- 1
Reasoning
12 benchmarks
69,321 - 2
Knowledge
7 benchmarks
61,488 - 3
Science
2 benchmarks
29,617 - 4
Law
3 benchmarks
17,510 - 5
Multilingual
4 benchmarks
12,642 - 6
Safety
10 benchmarks
11,691 - 7
Medicine
3 benchmarks
9,238 - 8
Agents & Tool Use
10 benchmarks
6,194 - 9
General
4 benchmarks
5,248 - 10
Software Engineering
9 benchmarks
2,607 - 11
Mathematics
5 benchmarks
2,085 - 12
ML Engineering
1 benchmark
212 - 13
Cultural
1 benchmark
152 - 14
NLP Tasks
1 benchmark
152
Benchmark Release Timeline
Every month-dated benchmark on this shelf, one dot per benchmark, stacked by release month and colored by primary domain.
Model Release Timeline
Every dated model evaluated on this shelf. One dot per model, stacked by release month and colored by provider.
130 models with no recorded release date are not shown.
Model × Domain Scores
Every model’s average score across each benchmark domain, over binary-scored benchmarks only. A benchmark that spans several domains counts toward each of them, so neighbouring columns can rest on the same responses. A hatched cell means the model was never measured in that domain.
- 3775755564685162378262
- 7772727769667123397123
- 8169695154537046477037
- 7778785466646051393149
- 8172724957546429406841
- 71757572707259113341
- 807889735960378752
- 737360606269608748
- 797963706365135177
- 757552525339458452
- 675656724441514278
- 757547475063425942
- 7177778885835083
- 8574748590814772
- 6765657073703389
- 7871717987701037
- 7373626263467546
- 6466666662593284
- 747497562507750
- 757567154427442
- 626245625558066
- 6060223232625556
- 7474383841601212
- 5050182525565249
- 77837674714681
- 76767662669059
- 73736974638658
- 74827470674879
- 77748474514661
- 69666671786134
- 73735455604174
- 72724545477064
- 80778977224022
- 72723546474859
- 73736920277927
- 62622130304671
- 6868261157415
- 767643434571
- 727272317131
- 727262186618
- 707027273163
- 414141643262
- 8291909172
- 8389908975
- 8188908876
- 8489908968
- 8086898674
- 8187888772
- 8187878771
- 7887898772
- 8189918958
- 8087878769
- 7686868674
- 8086898667
- 7985898566
- 7988918858
- 7486848673
- 7784888467
- 7386878668
- 8196949631
- 7088918860
- 7784888463
- 8087898753
- 7785878564
- 8189928943
- 8185888555
- 7883868361
- 7283838369
- 7185908559
- 7889878948
- 7581878165
- 7885878554
- 8489918935
- 7486898652
- 8285918545
- 7183858364
- 8279887960
- 8288798851
- 7682868261
- 7682848262
- 7681868162
- 7583858360
- 7782868257
- 7682868259
- 7485888553
- 7883868352
- 7684708463
- 7582848256
- 6883838362
- 7788908834
- 7179857961
- 7684858446
- 7481858153
- 7583708363
- 6088928844
- 5184878465
- 8283718353
- 7884728452
- 7379837955
- 7077837762
- 7281868147
- 5686918649
- 7678857852
- 7581858145
- 5182878264
- 7181838149
- 7678827852
- 7479817950
- 7376837655
- 7181878142
- 7377827754
- 7781628158
- 7380888038
- 6875827558
- 7679857938
- 7178877842
- 7277807750
- 7175807554
- 7774847446
- 6576787658
- 7178837844
- 7478687853
- 7681698145
- 4261777595
- 7275807548
- 7780698043
- 7576857638
- 7476857638
- 7578837835
- 7078797843
- 7074817447
- 7882868217
- 7077787740
- 7574827438
- 7679836738
- 7174837441
- 7175827540
- 7672847238
- 7978607846
- 6671767156
- 7271817145
- 7073827340
- 7876647646
- 6675807543
- 6675817541
- 7071767149
- 7073827339
- 7779647937
- 6875787540
- 7272827236
- 6678677844
- 6874787439
- 6870767050
- 7078777830
- 6974827432
- 6974777437
- 6969786943
- 6670807041
- 7073787333
- 6969786941
- 7070817034
- 7167806739
- 6873787331
- 7269796935
- 6870757039
- 7174667437
- 6172707247
- 7469826926
- 6870767037
- 6672777231
- 4674837438
- 6469776936
- 5971757138
- 7463796335
- 7566776630
- 7272497247
- 6865806535
- 7265766534
- 7264796433
- 6966756632
- 6769776929
- 7165796528
- 6867766729
- 6964766433
- 6565816528
- 6764756432
- 6566457845
- 6663766332
- 6664786429
- 7161766129
- 6762756233
- 7060766031
- 7160766029
- 6961776127
- 6760746034
- 6861766127
- 7065546540
- 6459725936
- 6568536835
- 6869546930
- 7258765825
- 4566786632
- 6957765729
- 6956765629
- 6658745831
- 6857755729
- 6557775729
- 5959755933
- 7057695730
- 6558755827
- 7056735625
- 6256705634
- 5959655936
- 6157675736
- 6856735625
- 6154695439
- 7171444346
- 6451745130
- 6253685333
- 6449754929
- 6262634038
- 5949714924
- 7053485327
- 5855625521
- 3158695830
- 8113436443
- 5446594634
- 5847584729
- 6042664222
- 6341624124
- 6838643823
- 365966596
- 5942604222
- 4141416139
- 5539493933
- 5441574116
- 2929315961
- 4146534620
- 5638623812
- 2929296451
- 5933563312
- 4646363628
- 5030503018
- 34490092
- 4627492716
- 4631313117
- 5126302621
- 4327262712
- 76777297
- 74787391
- 74636397
- 56677391
- 55718279
- 78797951
- 92745656
- 82715353
- 90863737
- 54717249
- 66646340
- 90742626
- 57574756
- 44425858
- 60604140
- 61614237
- 4685575
- 50495930
- 39393970
- 35353569
- 66206820
- 42415327
- 47642323
- 24242463
- 26262652
- 19191974
- 23232357
- 555544
- 17171766
- 19191957
- 19191931
- 11111134
- 10101628
- 816483
- 837171
- 817069
- 726868
- 616880
- 696969
- 477685
- 826062
- 416895
- 676767
- 716464
- 666666
- 656565
- 646467
- 646464
- 716943
- 606060
- 606060
- 606060
- 706940
- 463498
- 646239
- 684651
- 496748
- 734444
- 535353
- 774141
- 474764
- 505851
- 616035
- 704242
- 565541
- 753838
- 505050
- 585733
- 555437
- 494949
- 575633
- 565533
- 575532
- 535238
- 545335
- 235464
- 683535
- 555430
- 515036
- 515036
- 753131
- 545329
- 454545
- 454545
- 494935
- 505033
- 494832
- 494833
- 494831
- 494830
- 653131
- 474733
- 464536
- 474733
- 464633
- 484729
- 373749
- 702330
- 58585
- 464530
- 712424
- 424234
- 434230
- 424132
- 424131
- 424132
- 383838
- 741919
- 771818
- 612525
- 403930
- 293047
- 721420
- 282652
- 404026
- 49497
- 403924
- 383727
- 383725
- 333333
- 49492
- 393823
- 48484
- 64049
- 363523
- 53354
- 40408
- 43433
- 313125
- 7188
- 282828
- 282827
- 40403
- 272728
- 282825
- 272724
- 272724
- 262626
- 292821
- 262626
- 252527
- 272724
- 262625
- 272624
- 262625
- 262624
- 252527
- 252524
- 262624
- 262623
- 262624
- 252525
- 262624
- 262524
- 252525
- 6933
- 252524
- 262621
- 252524
- 232328
- 262622
- 262621
- 252523
- 242425
- 252523
- 272620
- 252522
- 262621
- 262620
- 252523
- 252423
- 242423
- 232326
- 262521
- 262620
- 252521
- 252422
- 232323
- 242422
- 232322
- 222222
- 272713
- 02936
- 4888
- 6200
- 212121
- 252410
- 191919
- 181818
- 171717
- 9777
- 8181
- 9072
- 8181
- 8080
- 6989
- 7979
- 7384
- 8668
- 7777
- 5698
- 7676
- 5597
- 7575
- 5693
- 7575
- 5891
- 5891
- 7474
- 7373
- 7373
- 7373
- 7171
- 7171
- 7070
- 7070
- 7070
- 4496
- 7070
- 6376
- 6969
- 6868
- 6868
- 6868
- 5976
- 6074
- 6767
- 6666
- 6666
- 6666
- 6565
- 6363
- 6363
- 6363
- 6161
- 5959
- 5959
- 5959
- 5858
- 5858
- 5757
- 4962
- 5555
- 5454
- 5454
- 5353
- 5353
- 5252
- 5252
- 5252
- 3964
- 5151
- 5151
- 5151
- 5050
- 4949
- 4949
- 4949
- 4747
- 4747
- 5341
- 4646
- 4646
- 4646
- 4545
- 4545
- 4545
- 4545
- 4545
- 4444
- 4343
- 4242
- 4242
- 4242
- 4141
- 4141
- 4040
- 4040
- 3840
- 3838
- 731
- 3737
- 3737
- 3636
- 3535
- 3333
- 3333
- 3232
- 2727
- 2626
- 2626
- 2323
- 1414
- 1313
- 11
- 11
- 11
- 91
- 90
- 88
- 86
- 85
- 84
- 84
- 83
- 83
- 81
- 79
- 78
- 78
- 77
- 77
- 76
- 75
- 75
- 75
- 74
- 74
- 74
- 74
- 73
- 72
- 72
- 72
- 72
- 71
- 71
- 71
- 71
- 70
- 69
- 69
- 69
- 69
- 69
- 68
- 68
- 68
- 68
- 68
- 67
- 67
- 67
- 66
- 66
- 66
- 66
- 65
- 65
- 64
- 64
- 64
- 64
- 64
- 64
- 63
- 63
- 63
- 63
- 63
- 63
- 63
- 62
- 62
- 62
- 62
- 62
- 61
- 61
- 61
- 60
- 60
- 60
- 59
- 59
- 59
- 59
- 59
- 59
- 59
- 58
- 57
- 57
- 57
- 57
- 57
- 57
- 57
- 57
- 57
- 57
- 56
- 56
- 56
- 56
- 56
- 56
- 55
- 55
- 55
- 55
- 55
- 55
- 54
- 54
- 54
- 54
- 54
- 53
- 53
- 52
- 52
- 51
- 51
- 50
- 50
- 50
- 50
- 50
- 49
- 49
- 49
- 49
- 48
- 48
- 48
- 48
- 48
- 47
- 47
- 47
- 47
- 47
- 47
- 47
- 47
- 46
- 46
- 46
- 46
- 46
- 45
- 45
- 45
- 45
- 45
- 44
- 43
- 43
- 43
- 43
- 43
- 41
- 40
- 39
- 39
- 39
- 39
- 39
- 38
- 38
- 38
- 38
- 37
- 37
- 37
- 36
- 36
- 35
- 35
- 35
- 35
- 34
- 34
- 33
- 32
- 31
- 30
- 30
- 29
- 29
- 28
- 27
- 27
- 26
- 26
- 23
- 22
- 21
- 20
- 20
- 20
- 19
- 15
- 14
- 14
- 12
- 11
- 10
- 10
- 9
- 9
- 8
- 8
- 8
- 8
- 7
- 7
- 7
- 7
- 7
- 7
- 7
- 7
- 6
- 6
- 6
- 6
- 6
- 6
- 6
- 6
- 6
- 6
- 5
- 5
- 4
- 4
- 3
- 2
- 2
- 2
- 2
- 2
- 1
- 1
- 0
- 0
What do the card tags mean?▶
- item-level responses released
- The benchmark’s per-model, per-item response data is fully public and downloadable from Hugging Face. Every (model × item) cell in the response matrix is available, not just aggregate scores. Benchmarks without this tag only provide item content and subject lists, or aggregate-level results.
- Nitems / subjects / attacks
- The count of distinct evaluation items (prompts/questions), AI subjects (models/agents) evaluated, or adversarial attacks in the bank’s standardized matrix for this benchmark.
- Saturation status: Yes
- The top-3 models average ≥ 90% on this benchmark, suggesting the task may be approaching a performance ceiling.
- Saturation status: No
- No: top models still score below 90%, so meaningful headroom remains. Benchmarks without enough item-level response data to judge saturation carry no saturation tag.
2026
10

DeepSWE
DeepSWE tests whether coding agents can complete long horizon software engineering tasks written from scratch against active open source repositories, graded by hand written functional verifiers the agent never sees rather than by tests shipped with an existing fix.

FrontierOR
FrontierOR asks models to read an operations research problem stated in plain language and write an optimization program, then checks how many large instances that program solves feasibly against an expert verified reference.

IKP
IKP asks models factual questions that range from things almost anyone knows to extreme long tail trivia, counting refusals as wrong. Because accuracy tracks model size closely, it doubles as a way to estimate the parameter count of a closed model.

MathArena Platform
MathArena Platform asks models to answer research mathematics questions drawn from arXiv papers, write formal proofs in Lean that a proof checker verifies, and recognize when a stated claim is deliberately false.

OSWorld 2.0
OSWorld 2.0 tests whether computer use agents can finish long real world tasks on a live Ubuntu desktop and its web applications, working from screenshots, with the final state of the machine checked against every task requirement.

ProgramBench
ProgramBench asks an agent to rebuild a working program from nothing but its compiled binary and the documentation shipped with it, scoring the result by how much of a hidden behavioural test suite the rebuilt program passes.

tau-Voice
tau-Voice tests whether voice agents can handle live customer service calls in retail, airline and telecom settings, where a call passes only if the database ends in the required state and the agent conveys the required information.

Terminal-Bench
Terminal-Bench tests whether agents working in a command line environment can carry out real engineering tasks such as building software from source, fixing a broken script or analyzing a dataset, checked by the task's own test suite.

Terminal-Bench 2.1
Terminal-Bench 2.1 tests whether an agent can finish a real task inside a Linux container it drives through a shell, passing only when a test suite run against the state it leaves behind succeeds entirely.
2025
8
GenAI Learning
GenAI Learning measures whether high school students answer maths practice problems correctly when working alone, with an unrestricted chatbot, or with a tutor bot that refuses to give the full solution. A language model also answers the same problems on its own.

ICPC-2 Code Selector
ICPC-2 Code Selector tests whether models can pick the primary care code that best matches a short clinical expression in Brazilian Portuguese from candidates returned by a search engine, or correctly abstain when none match.


MMDocRAG
MMDocRAG tests whether models can answer questions about documents by weaving retrieved text and images into a cited answer, which a judge model rates for fluency, citation quality, coherence, reasoning and factuality.


ResearchCodeBench
ResearchCodeBench tests whether models can implement a missing piece of a recent machine learning paper's code, given the paper itself and the surrounding file, with curated tests deciding whether the filled in code works.


tau2-bench
tau2-bench puts a model in the role of a customer support agent for an airline, retail or telecom account, using the company tools and database while talking to a simulated customer it must also guide through steps only they can perform.
2024
11
AfriMed-QA
AfriMed-QA asks models multiple choice medical questions set in African healthcare contexts, spanning many clinical specialties, and scores whether the option a model picks matches the correct one.

Alignment Faking (RL)
Alignment Faking (RL) records whether a model pretends to go along with a request it would otherwise refuse, and whether that changes with how far its training has progressed and which kind of user it is told it is talking to.

BFCL
BFCL tests whether models can answer a user request by calling the right function with the right arguments, including requests that no available function can satisfy and ones spread over several conversational turns.

CharXiv
CharXiv tests whether multimodal models can read real charts from scientific papers, answering both short descriptive questions and an open ended reasoning question, with a judge model grading the descriptive answers.

AIR-Bench 2024
AIR-Bench tests whether models give safe replies to prompts written against a taxonomy of risks drawn from government regulations and company policies, with every reply marked safe or unsafe.

HarmBench
HarmBench measures whether a model refuses harmful requests, covering behaviors such as cybercrime, misinformation and chemical or biological harm, with a classifier judging each reply as safe or unsafe.

LexEval
LexEval tests how well models answer Chinese legal multiple-choice questions covering legal knowledge, reasoning, discrimination and ethics, grading each answer by exact match of the extracted option letters against the gold answer using the official LexEval grader.

LiveCodeBench
LiveCodeBench tests whether models can solve competitive programming problems by writing a program that produces the expected output on the problem's test cases.


OSWorld
OSWorld tests whether agents can complete real tasks on an Ubuntu desktop driving applications such as Chrome, LibreOffice and VS Code, with each task graded by a scorer that can award partial credit.

SWE-bench Java
SWE-bench Java tests whether coding agents can resolve real GitHub issues in Java repositories by submitting a code patch, with each issue recorded as resolved or unresolved.
2023
10
SimpleSafetyTests
SimpleSafetyTests checks whether a model responds safely to blunt requests touching severe harms such as self harm, violence and illegal activity, with a classifier marking each reply safe or unsafe.

HELM Thai Exam
HELM Thai Exam asks models questions drawn from Thai standardized examinations, written in Thai, and counts an answer correct only when it exactly matches the official answer.


LawBench
LawBench measures how well models handle Chinese legal work; this curation ingests the six subtasks graded by per-item binary accuracy — judicial-exam and case-analysis multiple choice, dispute-focus and consultation classification, argumentation mining, and crime-amount extraction — regraded from the released raw predictions with the official evaluation code.

LegalBench
LegalBench tests whether models can answer legal reasoning questions, with a response counted correct when it matches the reference answer closely enough to pass an approximate string match.

MathVista MINI
MathVista MINI tests whether vision language models can answer mathematical questions about an accompanying image, some multiple choice and some free response, with an answer counted correct when it matches the ground truth.

MMBench V1.1
MMBench V1.1 tests whether vision language models can answer multiple choice questions about a picture, repeating each question with its answer options rotated so a model cannot lean on option position.


MMMU (dev+val)
MMMU asks models college level multiple choice questions that pair text with images, scoring an answer correct only when the option it picks matches the correct one.

SWE-bench Verified
SWE-bench Verified tests whether coding agents can resolve real issues reported against open source software projects by submitting a code patch that the benchmark's automatic evaluation accepts as a fix.
Earlier (pre-2023)
9


HAIID
HAIID measures whether people revise their answers on classification tasks, such as spotting sarcasm or judging skin lesion photographs, after seeing advice presented as coming either from an AI algorithm or from a human peer. The advice itself is identical in both cases.

Anthropic Red Team
Anthropic Red Team tests whether models avoid complying with adversarial prompts taken from human red teaming conversations, with a safety classifier judging each response as safe or unsafe.

BBQ
BBQ tests whether models pick the correct answer to multiple choice questions about people from different social groups, asking each question once with too little context to answer it and once with the context that settles it.


RealToxicityPrompts
RealToxicityPrompts measures how often a model continues a sentence fragment with text a toxicity classifier flags as toxic, using prompts drawn from both toxic and innocuous starting text.

MMLU
MMLU tests broad knowledge and reasoning with multiple choice exam questions spanning school and professional topics in the sciences, humanities and social sciences, scored by whether the model picks the right option.

TruthfulQA-MC
TruthfulQA-MC tests whether models select the single true answer rather than a plausible falsehood, on questions built around misconceptions that people commonly repeat.
02Human-centered AI Benchmarks
Every response is scored by a judgment: a human rating, a model rating, or a verdict from a judge that is itself the subject under evaluation. Scores therefore carry judge disagreement in addition to subject ability. Pairwise benchmarks are indexed by model pair rather than by shared item, so they are listed separately.
13
Benchmarks curated
302K
Evaluation items
2.1M
Judged responses
Meta-analyses
Provider Coverage, Ranked By The Number Of Items Evaluated
Each provider's most-measured model, ranked by how many distinct benchmark items it has been evaluated on.
- 1
Anthropic
Claude Opus 4
33,516 - 2
Google
Gemini 2.5 Flash
30,627 - 3
Alibaba
Qwen3 235B A22B
27,720 - 4
OpenAI
GPT-4o
23,265 - 5
Meta
Llama 4 Maverick
21,822 - 6
Mistral
Medium 3
18,160 - 7
xAI
Grok 3 Mini
17,330 - 8
Cohere
Command A
13,576 - 9
DeepSeek
R1
13,028 - 10
Amazon
Nova Pro
12,686 - 11M
MiniMax
M1
10,848 - 12A
AllenAI
Tulu 2 70B
10,163
Domain Coverage, Ranked By The Number Of Items Evaluated
Benchmarks that span multiple domains count toward each, so the totals overlap rather than partition the shelf.
- 1
Preference
7 benchmarks
263,498 - 2
Software Engineering
1 benchmark
31,930 - 3
Agents & Tool Use
1 benchmark
31,930 - 4
Reward Modeling
3 benchmarks
5,159 - 5
General
1 benchmark
764 - 6
Safety
1 benchmark
450 - 7
Multilingual
1 benchmark
120
Benchmark Release Timeline
Model Release Timeline
Every dated model evaluated on this shelf. One dot per model, stacked by release month and colored by provider.
5 models with no recorded release date are not shown.
Model × Domain Scores
Every model’s average score across each benchmark domain, over binary-scored benchmarks only. A benchmark that spans several domains counts toward each of them, so neighbouring columns can rest on the same responses. A hatched cell means the model was never measured in that domain.
- 7676
- 7474
- 6767
- 5656
- 5160
- 5252
- 5051
- 4848
- 4545
- 4343
- 4343
- 4242
- 4040
- 5029
- 5122
- 5018
- 5017
- 5117
- 2525
- 61
- 53
- 50
- 50
- 50
- 50
- 49
- 31
- 28
- 7
- 7
- 7
- 6
- 4
2026
1
SWE-chat
SWE-chat measures whether developers accept what a coding agent did during real sessions on their own repositories, with a judge model reading each following prompt to decide whether the developer pushed back.
2025
3
Arena 140K
Arena 140K records which of two models a person preferred after reading both replies to a prompt they wrote themselves, with ties and mutual rejection both counted as outcomes.

Prompt-to-Leaderboard
Prompt-to-Leaderboard ranks models on individual prompts using head to head battles in which a person sees two models' answers to the same prompt and votes for one or calls it a tie.

RewardBench 2
RewardBench 2 tests whether reward models rank the intended best response above the alternatives, on prompts covering areas such as factuality, math and safety.
2024
7
Arena-Hard-Auto
Arena-Hard-Auto rates how well models answer hard user requests, many of them programming and technical questions, by having a strong model judge each answer against a fixed baseline model's answer to the same request.

BiGGen-Bench
BiGGen-Bench rates model answers to prompts spanning reasoning, instruction following, safety, planning and other capabilities, with strong models and in some cases human annotators scoring each answer against a rubric.

JudgeBench
JudgeBench tests whether a model acting as a judge picks the correct response out of a pair, on comparisons drawn from knowledge, reasoning, math and coding, with each pair shown in both orders.

RewardBench
RewardBench tests whether reward models prefer the better of two candidate responses to a prompt, over prompts covering chat, safety and reasoning.

SORRY-Bench
SORRY-Bench measures whether models comply with or refuse unsafe requests, with human annotators labeling each reply, including when the request is rewritten in styles such as role play, slang or a cipher.

Tengu-Bench
Tengu-Bench rates how well models answer Japanese tasks spanning categories such as table reading and reasoning, with several language models acting as judges and scoring each answer against a written rubric.

2023
2

Preference Dissection
Preference Dissection records which of two candidate replies to the same prompt each model judge prefers, turning the judges themselves into the systems being measured.
Citation
Cite this work
If you use the data we curated, please cite the following reference. The curation is released under CC BY-SA 4.0; individual benchmarks retain their upstream licenses.
@misc{measurementdb2026, title = {The AI Measurement Data Bank}, author = {Truong, Nhi and Truong, Sang T. and Koyejo, Sanmi}, year = {2026}, howpublished = {\url{https://aimslab.stanford.edu/measurement-db}}, note = {AIMS Lab, Stanford University}}