Skip to main content

Search AIMS

Find pages, publications, events, course projects, and software. Results update as you type; press Enter to open the first result.

Search AIMS.

Find AIMS pages, publications, seminars, course projects, and software from one place.

Searches 119 AIMS pages and records, including the Measurement Data Bank entry point. The separately maintained textbook is linked at its main entry point.

Results for “predictive evaluation

29 results

  1. PageCommunity

    The Predictive AI Evaluation Competition at NeurIPS 2026

    Predict whether an AI system answered a benchmark question correctly. Submissions are Python code, evaluated with Brier scores reported for each AI subject and benchmark combination.

  2. NewsNews

    The Predictive AI Evaluation Competition is open

    The competition accepts code submissions through 20 November 2026 at 23:59 AoE.

  3. PublicationStatistical Models of Measurement

    A Measurement Science Roadmap: From Human Assessment to AI Evaluation

    S Truong, N Goodman, E Brunskill, B Domingue, N Haber, S Koyejo · Preprint · 2026

  4. PublicationStatistical Models of Measurement

    Measurement to Meaning: A Validity-Centered Framework for AI Evaluation

    O Salaudeen, A Reuel, A Ahmed, S Bedi, Z Robertson, S Sundar, et al. · Preprint · 2025

  5. PublicationStatistical Models of Measurement

    Reliable and Efficient Amortized Model-based Evaluation

    S Truong, Y Tu, P Liang, B Li, S Koyejo · ICML 2025 · 2025

  6. PublicationStatistical Models of Measurement

    Quantifying Variance in Evaluation Benchmarks

    L Madaan, AK Singh, R Schaeffer, A Poulton, S Koyejo, P Stenetorp, et al. · Preprint · 2024

  7. PublicationStatistical Models of Measurement

    Quantifying the Effect of Test Set Contamination on Generative Evaluations

    R Schaeffer, J Kazdan, B Abbasi, KZ Liu, B Miranda, A Ahmed, F Berez, et al. · Preprint · 2026

  8. PublicationStatistical Models of Measurement

    Correlating and Predicting Human Evaluations of Language Models from Natural Language Processing Benchmarks

    R Schaeffer, PS Koura, B Tang, R Subramanian, AK Singh, T Mihaylov, et al. · Preprint · 2025

  9. PublicationStatistical Models of Measurement

    Pretraining Scaling Laws for Generative Evaluations of Language Models

    R Schaeffer, N Levi, B Miranda, S Koyejo · ICLR 2026 · 2025

  10. Student projectCS321M

    Investigating Single-Turn Evaluation as a Predictor of Multi-Turn Child Safety

    Tracy Y Wei — Child-AI safety benchmarks almost universally rely on single-turn evaluation, but real interactions are conversational. I re-score 225 multi-turn conversations from the ChildSafe corpus (Claude Sonnet 4,…

  11. Student projectCS321M

    Signal-Detection IRT for AI Evaluation on Success-vs-Multi-Failure Tasks

    Lezhi (Carrie) Tan, Luna Lyu — A recurring class of AI evaluation tasks has an asymmetric label structure: a privileged “success” class and a heterogeneous set of failure subtypes that can co-occur within a single input — multi-label…

  12. Student projectCS321M

    Predictive Validity of Objective Audio Metrics for Human Preference in Music Generation

    Si Qi Chen — Evaluating music generation models is difficult because the target outcome is often human preference, and there is no single objective metric that fully captures the perceptual and aesthetic quality of…

  13. Student projectCS321M

    Latent Capability Estimation Under Sparse AI Evaluation

    Jacob Alan Rubenstein, Shane Robinson Mion, Vania Chow — Item Response Theory (IRT) and low-rank matrix completion are both used to estimate latent model capability from (model, item) response matrices, but they are typically benchmarked on dense, curated data rather…

  14. PageResearch

    AIMS Research

    Theory, methods, and evidence for rigorous AI measurement across modeling, validity, incentives, and real-world use.

  15. PageEducation

    CS 321M: AI Measurement Science at Stanford

    Frameworks and methodologies for measuring, benchmarking, and understanding AI systems.

  16. Research directionResearch

    Statistical Models of Measurement

    Foundations, models, and methods for predicting measurement outcomes, characterizing uncertainty, and assessing validity and reliability.

  17. PublicationStatistical Models of Measurement

    AI Cartography: Mapping the Latent Landscape of AI Benchmark Ecosystems

    M Hardy, A Reuel, L Zhang, JM Casabianca, S Truong, Y Dave, H Lee, B Domingue, S Koyejo · ICML 2026 · 2026

  18. PublicationStatistical Models of Measurement

    Item Response Scaling Laws: A Measurement Theory Approach for Efficient and Generalizable Neural Scaling Estimation

    S Truong, Y Tu, R Schaeffer, S Koyejo · ICML 2026 · 2026

  19. PublicationStatistical Models of Measurement

    What AI Benchmarks Actually Measure: Adapting Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks

    M Desai, S Truong, H Wallach, A Chouldechova, AF Cooper, J Garcia-Gathright, DE Ho, AZ Jacobs, S Koyejo, N Pangakis, A Wang · COLM 2026 · 2026

  20. PublicationStatistical Models of Measurement

    When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

    M Akhtar, A Reuel, P Soni, S Ahuja, PS Ammanamanchi, R Rawal, et al. · ICML 2026 · 2026

  21. PublicationStatistical Models of Measurement

    Why Do Safety Guardrails Degrade Across Languages?

    M Zhang, A Patel, S Truong, S Koyejo · COLM 2026 · 2026

  22. PublicationStatistical Models of Measurement

    Fantastic Bugs and Where to Find Them in AI Benchmarks

    S Truong, Y Tu, M Hardy, A Reuel, Z Tang, J Burapacheep, et al. · NeurIPS · 2025

  23. PublicationStatistical Models of Measurement

    How Do Large Language Monkeys Get Their Power (Laws)?

    R Schaeffer, J Kazdan, J Hughes, J Juravsky, S Price, A Lynch, E Jones, et al. · ICML 2025 · 2025

  24. Student projectCS321M

    Beyond the Transcript: Reliability and Bias in LLM-Based Depression Judgments

    Kerui Lu, Linyin Lyu — LLM-as-judge systems are increasingly used to automate subjective evaluation, but in high-stakes settings it is unclear whether their judgments reflect the intended construct or artifacts of the measurement…

  25. Student projectCS321M

    When Model Rankings Diverge: Operational Validity of Fraud Detection Metrics

    Deonna Owens — Fraud detection systems are commonly evaluated using aggregate predictive metrics such as AUROC, precision, recall, and F1 score. While these metrics summarize average classification performance, they may not…

  26. PageCommunity

    AI Measurement Science Workshop

    A one-day workshop on the science of measuring frontier AI systems: measurement under interaction, strategic optimization, and non-stationarity.

  27. PageResources

    AIMS Resources

    The textbook, the course, the competition, the tools, and the community channels. Pick a starting point.

  28. PageEducation

    CS321M Lecture Materials

    Browse CS321M lecture slides, discussion notebooks, and videos by class session.

  29. NewsNews

    Call for papers: AI Measurement Science Workshop at COLM 2026

    AIMS is organizing a one-day workshop on rigorous AI measurement, co-located with COLM 2026 in San Francisco. The research and competition tracks are now open.