Skip to main content

Search AIMS

Find pages, publications, events, course projects, and software. Results update as you type; press Enter to open the first result.

Search AIMS.

Find AIMS pages, publications, seminars, course projects, and software from one place.

Searches 119 AIMS pages and records, including the Measurement Data Bank entry point. The separately maintained textbook is linked at its main entry point.

Results for “CS321M

30 of 60 results

  1. NewsNews

    CS321M launches as a new Stanford course

    CS321M is a new class offered at Stanford as of Spring 2026, turning measurement ideas into a living course and a way for new contributors to join.

  2. PageEducation

    CS321M Lecture Materials

    Browse CS321M lecture slides, discussion notebooks, and videos by class session.

  3. Student projectCS321M

    A Psychometric Audit of Cross-Paper Robot Foundation Model Evaluation

    Adarsh Sairam Ambati — Robot foundation models like π₀, OpenVLA, and Octo are each evaluated on different tasks by different labs, producing a sparse matrix of success rates that cannot be compared directly across papers. We…

  4. Student projectCS321M

    A Validity Audit of Hallucination-Induced Unauthorized Action in LLM Agents

    Michelle Campeau, Jake Klosowski — The case for deploying LLM agents in irreversible-action settings rests on the premise that strong scores on existing agent-safety benchmarks predict an agent will not take an explicitly forbidden action when…

  5. Student projectCS321M

    Analyzing Domain Validity in SimpleQA using Generalizability Theory

    Nevin George — SimpleQA is widely used to rank large language models on short-answer factuality, but its 4326 items are unevenly distributed across nine topical domains (135–858 items per domain). We apply generalizability…

  6. Student projectCS321M

    Benchmark Drift: Empirical Characterization of IRT Parameter Stability Across LLM Generations and Architecture Families on Open LLM Leaderboard v2

    Keao Xu — LLM benchmark scores implicitly assume measurement invariance: the same subtask should separate strong from weak models consistently regardless of when or by whom models were trained. We examine this assumption…

  7. Student projectCS321M

    Benchmark Items as Measurement Instruments: DIF Diagnostics for Response-Derived LLM Regimes on BBH

    Yiqing Liu — Large language model benchmarks are usually reported as aggregate leaderboard scores, but the underlying observations are item responses. I treat BBH as a measurement instrument: models are respondents, BBH…

  8. Student projectCS321M

    Beyond Stratified Accuracy: An Item Response Theory Audit of Dermatology AI Fairness

    Sonnet H Xu — We evaluate the inferential validity of Fitzpatrick skin type (FST)-stratified accuracy as a fairness metric for dermatology AI by applying amortized Item Response Theory (IRT) to the Diverse Dermatology Images…

  9. Student projectCS321M

    Beyond the Transcript: Reliability and Bias in LLM-Based Depression Judgments

    Kerui Lu, Linyin Lyu — LLM-as-judge systems are increasingly used to automate subjective evaluation, but in high-stakes settings it is unclear whether their judgments reflect the intended construct or artifacts of the measurement…

  10. Student projectCS321M

    Correcting for Non-Random Missingness in AI Benchmark Evaluation with Doubly Robust Estimation

    Antoine Maechler — AI model evaluation increasingly relies on benchmark leaderboards, yet the underlying response matrix (models × items) is almost never fully observed. We show that this missingness is systematically non-random:…

  11. Student projectCS321M

    Criterion Validity of General-Purpose LLM Benchmarks for Joint Military Domain Tasks

    Garrett Alarcon — General-purpose benchmark scores are the dominant way large language models (LLMs) are compared, and a natural starting point when defense organizations weigh which models to adopt. Yet, to our knowledge, no…

  12. Student projectCS321M

    De-biasing the LLM-as-a-judge: Using Latent Linguistic Complexity as a Proxy for Perplexity

    Elizabeth Sinyavin — This paper develops a methodology for mitigating the confounding factor of perplexity in LLM-as-a-judge evaluation systems, motivated by Wataoka et al.'s finding that LLM evaluators prefer responses with lower…

  13. Student projectCS321M

    Deconstructing Agentic Capabilities: A Generalizability Theory Approach to Tool-Use Evaluation

    Gilbert Gao — The current evaluation paradigm for agentic Large Language Models (LLMs) mostly rely on a monolithic end-to-end success rate. This metric is convenient, but it mixes up several distinct latent capabilities—such…

  14. Student projectCS321M

    Deployment Decision Reliability: A Generalizability-Theory Framework for Sizing Long-Horizon Agent Evaluations

    Vasundhara Srinivasan — Enterprise practitioners read agent leaderboards as if they ranked agent capability. We show, across three open agent-trace benchmarks (TheAgentCompany, τ²-bench, and AppWorld), that the agent main effect…

  15. Student projectCS321M

    Devil in the Details: Which Prompt Details Are Load-Bearing in Algorithmic Code Generation

    Austin Ho — Standard code-generation benchmarks evaluate large language models (LLMs) on behavioral correctness: whether generated code passes a given test suite. But as software is increasingly written by prompting LLMs…

  16. Student projectCS321M

    Dimensionality and Measurement Precision in HLE's Multiple-Choice Subset

    Mayank Sharma, Savira Nadela, Tyler Matteson — Humanity's Last Exam (HLE) has emerged as a prominent benchmark for evaluating advanced language models, yet widespread adoption has outpaced systematic evaluation of its measurement properties. Most benchmark…

  17. Student projectCS321M

    Do LLM Personas Have Personality? A Behavioral Benchmark of Rank-Order and Ipsative Consistency Across Situations

    Yikun Chi — LLM persona systems are increasingly used as synthetic respondents, simulated users, social agents, and role-playing characters, but most personality evaluations still ask whether a model can report a prompted…

  18. Student projectCS321M

    Do Medical QA Benchmarks Measure Medically Grounded Reasoning? A Natural Language Autoencoder Study of Latent Explanations

    Blake Edward Masters — Medical multiple-choice question answering benchmarks are often reported as evidence that a language model has acquired medically grounded reasoning. Final-answer accuracy is useful, but it is an incomplete…

  19. Student projectCS321M

    Do Strong Solvers Make Good Judges? An IRT Analysis of LLM-as-a-Judge Quality

    Sze Heng Douglas Kwok, Allison Sara John, Jiecong Tan — The LLM-as-a-judge paradigm has become a widely adopted approach for scalable model evaluation, yet it remains poorly understood whether a model's task-solving ability predicts its judgment quality. We…

  20. Student projectCS321M

    Does knowing partisan affiliation change GPT-4's plan to persuade?

    Emily Zou — This paper asks whether knowing a target's partisan affiliation changes GPT-4's plan to persuade, situating the question in research on micro-targeting — the tailoring of persuasive strategies to information…

  21. Student projectCS321M

    Does the Proxy Preserve the Decision? Auditing Multiple-Choice Reformulations of Incremental AI Benchmarks

    Ankit Aggarwal — This paper studies a common evaluation shortcut: replacing an open-ended, sequential task with a proxy that is easier to score. The testbed is pyramidal quizbowl, wherein a tossup question is revealed clue by…

  22. Student projectCS321M

    ErrorNodeBench-Interference: Why Bad-Entry Rate Misses Memory-Collapse Failures in LLM Consolidation

    Yuxiang Liu, Tongchang Zhang — Persistent textual memory is a popular alternative to weight updates for long-running LLM agents, but the rewriting step itself can corrupt useful evidence. We present ErrorNodeBench-Interference, a compact…

  23. Student projectCS321M

    Evaluating LLMs as Poker Players: An Item-Response Theory and Q-Matrix Analysis of PokerBench

    Abhinav Sattiraju — PokerBench, a recent static benchmark for evaluating large language models on no-limit Texas Hold'em poker decision-making, reports a single accuracy score per model and implicitly treats "poker skill" as a…

  24. Student projectCS321M

    GLMM G-Theory and D-Study for AI Benchmarking

    Jacob Daniel Householder — We apply G-theory and D-study analysis to three AI benchmarks (HELM Instruct, BiGGen-Bench, WildBench) under three likelihoods: the standard linear mixed model, a cumulative-logit GLMM, and a heteroscedastic…

  25. Student projectCS321M

    How Many Items Do You Really Need? IRT-Based Redundancy Analysis of LLM Benchmarks

    Dinesh Katupputhur Ramprasath — LLM benchmarks contain hundreds to thousands of evaluation items, yet it remains unclear how many are actually necessary to produce reliable model rankings. We apply Item Response Theory (IRT) to six LLM…

  26. Student projectCS321M

    Identifying Representative Subsets of Evaluation Datasets

    Anthony Maltsev — Evaluating large language models in deployment requires assessing (model, quantization, hardware) tuples, not base models alone. Full-benchmark evaluation of every tuple is expensive; we study methods for…

  27. Student projectCS321M

    Inference Method as a Measurement Facet: Benchmark Validity Under Speculative Decoding

    Rami Ratl Mrad — Benchmark scores are treated as properties of models, but they are also products of a full evaluation pipeline. Each component of that pipeline—prompting, decoding, answer extraction, scoring—is a potential…

  28. Student projectCS321M

    Investigating Single-Turn Evaluation as a Predictor of Multi-Turn Child Safety

    Tracy Y Wei — Child-AI safety benchmarks almost universally rely on single-turn evaluation, but real interactions are conversational. I re-score 225 multi-turn conversations from the ChildSafe corpus (Claude Sonnet 4,…

  29. Student projectCS321M

    Latent Capability Estimation Under Sparse AI Evaluation

    Jacob Alan Rubenstein, Shane Robinson Mion, Vania Chow — Item Response Theory (IRT) and low-rank matrix completion are both used to estimate latent model capability from (model, item) response matrices, but they are typically benchmarked on dense, curated data rather…

  30. Student projectCS321M

    Learning Prompt-Invariant Measurement Models via Sparse Twin-Tower Architectures

    Pratham Soni — The paper addresses the brittleness of LLM benchmark evaluation, where measured performance varies with superficial changes to prompt wording, formatting, or answer presentation even when the underlying…