The Science of
AI Measurement,
for science.
Current AI benchmarks often fail to generalize beyond the settings in which they are developed and can be optimized without corresponding improvements in underlying model capability. AI Measurement Science (AIMS) advances the science of AI measurement through research, teaching, and open resources.
AIMS at a glance
About AIMSRigorous measurement is the foundation of trustworthy AI.
AI claims outpace evidence
Benchmark scores are hard to interpret without explicit constructs, validated instruments, and uncertainty reporting.
Decisions depend on measurement quality
Deployment, regulation, and funding all rely on evaluation results.
The field lacks shared infrastructure
No unified community, curriculum, or software stack exists yet.
- 1
Reasoning
69 benchmarks
1,176,697 - 2
Knowledge
38 benchmarks
759,312 - 3
Safety
44 benchmarks
444,135 - 4
General
10 benchmarks
398,000 - 5
Mathematics
24 benchmarks
296,542 - 6
NLP Tasks
9 benchmarks
286,828 - 7
ML Engineering
5 benchmarks
271,701 - 8
Science
18 benchmarks
265,723 - 9
Preference
10 benchmarks
205,834 - 10
Multilingual
10 benchmarks
136,069
The AI Measurement
Data Bank
AI benchmarks are abundant, but their underlying evaluation data remain fragmented and difficult to reuse. Measurement DB curates benchmark results into a unified, item-level database spanning hundreds of evaluations and millions of model responses. Instead of reporting only leaderboard scores, it exposes the response matrices needed to study validity, reliability, model ability, item difficulty, benchmark overlap, and other fundamental measurement questions. It is designed to serve as shared infrastructure for the next generation of AI evaluation research.
Explore Measurement Data Bank01Research
We believe progress in AI depends on our ability to measure it. Our research develops the theory, methods, and infrastructure to make AI evaluation a rigorous science.
Predictive evaluation
A Measurement Science Roadmap: From Human Assessment to AI Evaluation
S Truong, N Goodman, E Brunskill, B Domingue, N Haber, S Koyejo
2026
Pretraining Scaling Laws for Generative Evaluations of Language Models
R Schaeffer, N Levi, B Miranda, S Koyejo
2025
Reliable and Efficient Amortized Model-based Evaluation
S Truong, Y Tu, P Liang, B Li, S Koyejo
2025
Validity and reliability
When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
M Akhtar, A Reuel, P Soni, S Ahuja, PS Ammanamanchi, R Rawal, et al.
2026
Quantifying the Effect of Test Set Contamination on Generative Evaluations
R Schaeffer, J Kazdan, B Abbasi, KZ Liu, B Miranda, A Ahmed, F Berez, et al.
2026
Fantastic Bugs and Where to Find Them in AI Benchmarks
S Truong, Y Tu, M Hardy, A Reuel, Z Tang, J Burapacheep, et al.
2025
Incentive, design, and governance
Strategic Evaluation: Incentivizing AI Capability Coverage with Private Benchmarks
S Truong, S Wang, N Haber, S Koyejo
2026
Public AI Benchmarks Are Broken, But Are Private Benchmarks the Answer?
S Truong, S Wang, A Wang, N Truong, A Reuel, N Haber, S Koyejo
2026
Stop Automating Peer Review Without Rigorous Evaluation
J Baumann, J Pei, S Koyejo, D Hovy
2026
Evaluation in the real world
SWE-chat: Coding Agent Interactions From Real Users in the Wild
J Baumann, V Padmakumar, X Li, J Yang, D Yang, S Koyejo
2026
In-Situ Behavioral Evaluation for LLM Fairness, Not Standardized-Test Scores
Z Tang, S Truong, D Owens, S Sharma, YJ Zhang, B Miranda, S Koyejo
2026
Holistic evaluation of large language models for medical tasks with MedHELM
S Bedi, H Cui, M Fuentes, A Unell, M Wornow, JM Banda, N Kotecha, et al.
2026
Explore our educational resources, including university course and open textbook on AI measurement science.
02Education

AI Measurement Science Course (CS321M)
This course introduces the foundations, methods, and applications of Artificial Intelligence (AI) measurement science through three connected themes: Measurement as Predictive Modeling, Measurement Validity and Reliability, and Design, Governance, and Applications.

AI Measurement Science Textbook
This textbook provides a comprehensive reference on Artificial Intelligence (AI) measurement science, covering its theoretical foundations, core methodologies, and practical applications.
03Software and Data
Open-source libraries, datasets, and interactive tools that make rigorous AI measurement practical.
News & updates
View all newsThe measurement textbook has been reorganized
The living textbook has been reorganized, and the full PDF is now available to download.
The Predictive AI Evaluation Challenge is open
AIMS's first competition invites participants to predict how models respond to unseen benchmark items from sparse observations. Submissions run through the summer.
Call for papers: AI Measurement Science Workshop at COLM 2026
AIMS is organizing a one-day workshop on rigorous AI evaluation, co-located with COLM 2026 in San Francisco. The research and competition tracks are now open.

