Skip to main content

Search AIMS

Find pages, publications, events, course projects, and software. Results update as you type; press Enter to open the first result.

Search AIMS.

Find AIMS pages, publications, seminars, course projects, and software from one place.

Searches 119 AIMS pages and records, including the Measurement Data Bank entry point. The separately maintained textbook is linked at its main entry point.

Results for “measurement validity

16 results

  1. PublicationStatistical Models of Measurement

    Measurement to Meaning: A Validity-Centered Framework for AI Evaluation

    O Salaudeen, A Reuel, A Ahmed, S Bedi, Z Robertson, S Sundar, et al. · Preprint · 2025

  2. Student projectCS321M

    Inference Method as a Measurement Facet: Benchmark Validity Under Speculative Decoding

    Rami Ratl Mrad — Benchmark scores are treated as properties of models, but they are also products of a full evaluation pipeline. Each component of that pipeline—prompting, decoding, answer extraction, scoring—is a potential…

  3. PageEducation

    CS 321M: AI Measurement Science at Stanford

    Frameworks and methodologies for measuring, benchmarking, and understanding AI systems.

  4. PageCommunity

    AI Measurement Science Workshop

    A one-day workshop on the science of measuring frontier AI systems: measurement under interaction, strategic optimization, and non-stationarity.

  5. Research directionResearch

    Statistical Models of Measurement

    Foundations, models, and methods for predicting measurement outcomes, characterizing uncertainty, and assessing validity and reliability.

  6. PageEducation

    AI Measurement Science Textbook

    A comprehensive reference on the foundations, methods, and practical applications of AI measurement science.

  7. PublicationStatistical Models of Measurement

    What AI Benchmarks Actually Measure: Adapting Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks

    M Desai, S Truong, H Wallach, A Chouldechova, AF Cooper, J Garcia-Gathright, DE Ho, AZ Jacobs, S Koyejo, N Pangakis, A Wang · COLM 2026 · 2026

  8. Student projectCS321M

    Benchmark Items as Measurement Instruments: DIF Diagnostics for Response-Derived LLM Regimes on BBH

    Yiqing Liu — Large language model benchmarks are usually reported as aggregate leaderboard scores, but the underlying observations are item responses. I treat BBH as a measurement instrument: models are respondents, BBH…

  9. Student projectCS321M

    Theory of Mind or Theory of Benchmarks? Construct Coherence and Measurement Noise in LLM ToM Evaluation

    Hannah Guan, Joey Ji — Published item counts of two prominent Theory of Mind (ToM) benchmarks for LLMs overstate effective measurement by roughly 2×. We reach this conclusion through a psychometric audit of ToMBench [Chen et al.,…

  10. PageEducation

    CS321M Lecture Materials

    Browse CS321M lecture slides, discussion notebooks, and videos by class session.

  11. PageResearch

    AIMS Research

    Theory, methods, and evidence for rigorous AI measurement across modeling, validity, incentives, and real-world use.

  12. Student projectCS321M

    Beyond Stratified Accuracy: An Item Response Theory Audit of Dermatology AI Fairness

    Sonnet H Xu — We evaluate the inferential validity of Fitzpatrick skin type (FST)-stratified accuracy as a fairness metric for dermatology AI by applying amortized Item Response Theory (IRT) to the Diverse Dermatology Images…

  13. BlogBlog

    Compass and certificate

    A benchmark can guide progress without certifying readiness.

  14. Student projectCS321M

    Do Medical QA Benchmarks Measure Medically Grounded Reasoning? A Natural Language Autoencoder Study of Latent Explanations

    Blake Edward Masters — Medical multiple-choice question answering benchmarks are often reported as evidence that a language model has acquired medically grounded reasoning. Final-answer accuracy is useful, but it is an incomplete…

  15. Student projectCS321M

    Measuring the Reliability of LLM Judges for Marketplace Copy Review

    Maximilian Schaum — LLM-as-a-judge is a tempting first-pass tool for product and UI copy review. The question this paper asks is not whether LLM judges can “edit copy” but whether the judging workflow produces stable enough triage…

  16. Student projectCS321M

    Probing Deceptive Alignment: Linear Separability

    Gabriel Alexander Eidelman — Linear probes on residual stream activations detect strategic deception in large language models with high accuracy. Prior work has demonstrated that probes can generalize to a variety of contexts, so they hold…