Skip to main content

The Science of
AI Measurement,
for science.

Current AI benchmarks often fail to generalize beyond the settings in which they are developed and can be optimized without corresponding improvements in underlying model capability. AI Measurement Science (AIMS) advances the science of AI measurement through research, teaching, and open resources.

AIMS at a glance

About AIMS
Education

Learn the field, end to end.

Software & Data

Tools and data for careful measurement.

Competition

Stress-test methods in public.

Workshop

Convene the community.

Rigorous measurement is the foundation of trustworthy AI.

01

AI claims outpace evidence

Benchmark scores are hard to interpret without explicit constructs, validated instruments, and uncertainty reporting.

02

Decisions depend on measurement quality

Deployment, regulation, and funding all rely on evaluation results.

03

The field lacks shared infrastructure

No unified community, curriculum, or software stack exists yet.

Learn more about us
#DomainItems
  1. 1

    Reasoning

    69 benchmarks

    1,176,697
  2. 2

    Knowledge

    38 benchmarks

    759,312
  3. 3

    Safety

    44 benchmarks

    444,135
  4. 4

    General

    10 benchmarks

    398,000
  5. 5

    Mathematics

    24 benchmarks

    296,542
  6. 6

    NLP Tasks

    9 benchmarks

    286,828
  7. 7

    ML Engineering

    5 benchmarks

    271,701
  8. 8

    Science

    18 benchmarks

    265,723
  9. 9

    Preference

    10 benchmarks

    205,834
  10. 10

    Multilingual

    10 benchmarks

    136,069

The AI Measurement
Data Bank

AI benchmarks are abundant, but their underlying evaluation data remain fragmented and difficult to reuse. Measurement DB curates benchmark results into a unified, item-level database spanning hundreds of evaluations and millions of model responses. Instead of reporting only leaderboard scores, it exposes the response matrices needed to study validity, reliability, model ability, item difficulty, benchmark overlap, and other fundamental measurement questions. It is designed to serve as shared infrastructure for the next generation of AI evaluation research.

Explore Measurement Data Bank

01Research

We believe progress in AI depends on our ability to measure it. Our research develops the theory, methods, and infrastructure to make AI evaluation a rigorous science.

03Software and Data

Open-source libraries, datasets, and interactive tools that make rigorous AI measurement practical.