Search AIMS.
Find AIMS pages, publications, seminars, course projects, and software from one place.
Results for “predictive evaluation”
29 results
The Predictive AI Evaluation Competition at NeurIPS 2026
Predict whether an AI system answered a benchmark question correctly. Submissions are Python code, evaluated with Brier scores reported for each AI subject and benchmark combination.
The Predictive AI Evaluation Competition is open
The competition accepts code submissions through 20 November 2026 at 23:59 AoE.
A Measurement Science Roadmap: From Human Assessment to AI Evaluation
S Truong, N Goodman, E Brunskill, B Domingue, N Haber, S Koyejo · Preprint · 2026
Measurement to Meaning: A Validity-Centered Framework for AI Evaluation
O Salaudeen, A Reuel, A Ahmed, S Bedi, Z Robertson, S Sundar, et al. · Preprint · 2025
Reliable and Efficient Amortized Model-based Evaluation
S Truong, Y Tu, P Liang, B Li, S Koyejo · ICML 2025 · 2025
Quantifying Variance in Evaluation Benchmarks
L Madaan, AK Singh, R Schaeffer, A Poulton, S Koyejo, P Stenetorp, et al. · Preprint · 2024
Quantifying the Effect of Test Set Contamination on Generative Evaluations
R Schaeffer, J Kazdan, B Abbasi, KZ Liu, B Miranda, A Ahmed, F Berez, et al. · Preprint · 2026
Correlating and Predicting Human Evaluations of Language Models from Natural Language Processing Benchmarks
R Schaeffer, PS Koura, B Tang, R Subramanian, AK Singh, T Mihaylov, et al. · Preprint · 2025
Pretraining Scaling Laws for Generative Evaluations of Language Models
R Schaeffer, N Levi, B Miranda, S Koyejo · ICLR 2026 · 2025
Investigating Single-Turn Evaluation as a Predictor of Multi-Turn Child Safety
Tracy Y Wei — Child-AI safety benchmarks almost universally rely on single-turn evaluation, but real interactions are conversational. I re-score 225 multi-turn conversations from the ChildSafe corpus (Claude Sonnet 4,…
Signal-Detection IRT for AI Evaluation on Success-vs-Multi-Failure Tasks
Lezhi (Carrie) Tan, Luna Lyu — A recurring class of AI evaluation tasks has an asymmetric label structure: a privileged “success” class and a heterogeneous set of failure subtypes that can co-occur within a single input — multi-label…
Predictive Validity of Objective Audio Metrics for Human Preference in Music Generation
Si Qi Chen — Evaluating music generation models is difficult because the target outcome is often human preference, and there is no single objective metric that fully captures the perceptual and aesthetic quality of…
Latent Capability Estimation Under Sparse AI Evaluation
Jacob Alan Rubenstein, Shane Robinson Mion, Vania Chow — Item Response Theory (IRT) and low-rank matrix completion are both used to estimate latent model capability from (model, item) response matrices, but they are typically benchmarked on dense, curated data rather…
AIMS Research
Theory, methods, and evidence for rigorous AI measurement across modeling, validity, incentives, and real-world use.
CS 321M: AI Measurement Science at Stanford
Frameworks and methodologies for measuring, benchmarking, and understanding AI systems.
Statistical Models of Measurement
Foundations, models, and methods for predicting measurement outcomes, characterizing uncertainty, and assessing validity and reliability.
AI Cartography: Mapping the Latent Landscape of AI Benchmark Ecosystems
M Hardy, A Reuel, L Zhang, JM Casabianca, S Truong, Y Dave, H Lee, B Domingue, S Koyejo · ICML 2026 · 2026
Item Response Scaling Laws: A Measurement Theory Approach for Efficient and Generalizable Neural Scaling Estimation
S Truong, Y Tu, R Schaeffer, S Koyejo · ICML 2026 · 2026
What AI Benchmarks Actually Measure: Adapting Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks
M Desai, S Truong, H Wallach, A Chouldechova, AF Cooper, J Garcia-Gathright, DE Ho, AZ Jacobs, S Koyejo, N Pangakis, A Wang · COLM 2026 · 2026
When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
M Akhtar, A Reuel, P Soni, S Ahuja, PS Ammanamanchi, R Rawal, et al. · ICML 2026 · 2026
Why Do Safety Guardrails Degrade Across Languages?
M Zhang, A Patel, S Truong, S Koyejo · COLM 2026 · 2026
Fantastic Bugs and Where to Find Them in AI Benchmarks
S Truong, Y Tu, M Hardy, A Reuel, Z Tang, J Burapacheep, et al. · NeurIPS · 2025
How Do Large Language Monkeys Get Their Power (Laws)?
R Schaeffer, J Kazdan, J Hughes, J Juravsky, S Price, A Lynch, E Jones, et al. · ICML 2025 · 2025
Beyond the Transcript: Reliability and Bias in LLM-Based Depression Judgments
Kerui Lu, Linyin Lyu — LLM-as-judge systems are increasingly used to automate subjective evaluation, but in high-stakes settings it is unclear whether their judgments reflect the intended construct or artifacts of the measurement…
When Model Rankings Diverge: Operational Validity of Fraud Detection Metrics
Deonna Owens — Fraud detection systems are commonly evaluated using aggregate predictive metrics such as AUROC, precision, recall, and F1 score. While these metrics summarize average classification performance, they may not…
AI Measurement Science Workshop
A one-day workshop on the science of measuring frontier AI systems: measurement under interaction, strategic optimization, and non-stationarity.
AIMS Resources
The textbook, the course, the competition, the tools, and the community channels. Pick a starting point.
CS321M Lecture Materials
Browse CS321M lecture slides, discussion notebooks, and videos by class session.
Call for papers: AI Measurement Science Workshop at COLM 2026
AIMS is organizing a one-day workshop on rigorous AI measurement, co-located with COLM 2026 in San Francisco. The research and competition tracks are now open.