Search AIMS.
Find AIMS pages, publications, seminars, course projects, and software from one place.
Results for “measurement validity”
16 results
Measurement to Meaning: A Validity-Centered Framework for AI Evaluation
O Salaudeen, A Reuel, A Ahmed, S Bedi, Z Robertson, S Sundar, et al. · Preprint · 2025
Inference Method as a Measurement Facet: Benchmark Validity Under Speculative Decoding
Rami Ratl Mrad — Benchmark scores are treated as properties of models, but they are also products of a full evaluation pipeline. Each component of that pipeline—prompting, decoding, answer extraction, scoring—is a potential…
CS 321M: AI Measurement Science at Stanford
Frameworks and methodologies for measuring, benchmarking, and understanding AI systems.
AI Measurement Science Workshop
A one-day workshop on the science of measuring frontier AI systems: measurement under interaction, strategic optimization, and non-stationarity.
Statistical Models of Measurement
Foundations, models, and methods for predicting measurement outcomes, characterizing uncertainty, and assessing validity and reliability.
AI Measurement Science Textbook
A comprehensive reference on the foundations, methods, and practical applications of AI measurement science.
What AI Benchmarks Actually Measure: Adapting Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks
M Desai, S Truong, H Wallach, A Chouldechova, AF Cooper, J Garcia-Gathright, DE Ho, AZ Jacobs, S Koyejo, N Pangakis, A Wang · COLM 2026 · 2026
Benchmark Items as Measurement Instruments: DIF Diagnostics for Response-Derived LLM Regimes on BBH
Yiqing Liu — Large language model benchmarks are usually reported as aggregate leaderboard scores, but the underlying observations are item responses. I treat BBH as a measurement instrument: models are respondents, BBH…
Theory of Mind or Theory of Benchmarks? Construct Coherence and Measurement Noise in LLM ToM Evaluation
Hannah Guan, Joey Ji — Published item counts of two prominent Theory of Mind (ToM) benchmarks for LLMs overstate effective measurement by roughly 2×. We reach this conclusion through a psychometric audit of ToMBench [Chen et al.,…
CS321M Lecture Materials
Browse CS321M lecture slides, discussion notebooks, and videos by class session.
AIMS Research
Theory, methods, and evidence for rigorous AI measurement across modeling, validity, incentives, and real-world use.
Beyond Stratified Accuracy: An Item Response Theory Audit of Dermatology AI Fairness
Sonnet H Xu — We evaluate the inferential validity of Fitzpatrick skin type (FST)-stratified accuracy as a fairness metric for dermatology AI by applying amortized Item Response Theory (IRT) to the Diverse Dermatology Images…
Compass and certificate
A benchmark can guide progress without certifying readiness.
Do Medical QA Benchmarks Measure Medically Grounded Reasoning? A Natural Language Autoencoder Study of Latent Explanations
Blake Edward Masters — Medical multiple-choice question answering benchmarks are often reported as evidence that a language model has acquired medically grounded reasoning. Final-answer accuracy is useful, but it is an incomplete…
Measuring the Reliability of LLM Judges for Marketplace Copy Review
Maximilian Schaum — LLM-as-a-judge is a tempting first-pass tool for product and UI copy review. The question this paper asks is not whether LLM judges can “edit copy” but whether the judging workflow produces stable enough triage…
Probing Deceptive Alignment: Linear Separability
Gabriel Alexander Eidelman — Linear probes on residual stream activations detect strategic deception in large language models with high accuracy. Prior work has demonstrated that probes can generalize to a variety of contexts, so they hold…