Search AIMS.
Find AIMS pages, publications, seminars, course projects, and software from one place.
Results for “CS321M”
30 of 60 results
CS321M launches as a new Stanford course
CS321M is a new class offered at Stanford as of Spring 2026, turning measurement ideas into a living course and a way for new contributors to join.
CS321M Lecture Materials
Browse CS321M lecture slides, discussion notebooks, and videos by class session.
A Psychometric Audit of Cross-Paper Robot Foundation Model Evaluation
Adarsh Sairam Ambati — Robot foundation models like π₀, OpenVLA, and Octo are each evaluated on different tasks by different labs, producing a sparse matrix of success rates that cannot be compared directly across papers. We…
A Validity Audit of Hallucination-Induced Unauthorized Action in LLM Agents
Michelle Campeau, Jake Klosowski — The case for deploying LLM agents in irreversible-action settings rests on the premise that strong scores on existing agent-safety benchmarks predict an agent will not take an explicitly forbidden action when…
Analyzing Domain Validity in SimpleQA using Generalizability Theory
Nevin George — SimpleQA is widely used to rank large language models on short-answer factuality, but its 4326 items are unevenly distributed across nine topical domains (135–858 items per domain). We apply generalizability…
Benchmark Drift: Empirical Characterization of IRT Parameter Stability Across LLM Generations and Architecture Families on Open LLM Leaderboard v2
Keao Xu — LLM benchmark scores implicitly assume measurement invariance: the same subtask should separate strong from weak models consistently regardless of when or by whom models were trained. We examine this assumption…
Benchmark Items as Measurement Instruments: DIF Diagnostics for Response-Derived LLM Regimes on BBH
Yiqing Liu — Large language model benchmarks are usually reported as aggregate leaderboard scores, but the underlying observations are item responses. I treat BBH as a measurement instrument: models are respondents, BBH…
Beyond Stratified Accuracy: An Item Response Theory Audit of Dermatology AI Fairness
Sonnet H Xu — We evaluate the inferential validity of Fitzpatrick skin type (FST)-stratified accuracy as a fairness metric for dermatology AI by applying amortized Item Response Theory (IRT) to the Diverse Dermatology Images…
Beyond the Transcript: Reliability and Bias in LLM-Based Depression Judgments
Kerui Lu, Linyin Lyu — LLM-as-judge systems are increasingly used to automate subjective evaluation, but in high-stakes settings it is unclear whether their judgments reflect the intended construct or artifacts of the measurement…
Correcting for Non-Random Missingness in AI Benchmark Evaluation with Doubly Robust Estimation
Antoine Maechler — AI model evaluation increasingly relies on benchmark leaderboards, yet the underlying response matrix (models × items) is almost never fully observed. We show that this missingness is systematically non-random:…
Criterion Validity of General-Purpose LLM Benchmarks for Joint Military Domain Tasks
Garrett Alarcon — General-purpose benchmark scores are the dominant way large language models (LLMs) are compared, and a natural starting point when defense organizations weigh which models to adopt. Yet, to our knowledge, no…
De-biasing the LLM-as-a-judge: Using Latent Linguistic Complexity as a Proxy for Perplexity
Elizabeth Sinyavin — This paper develops a methodology for mitigating the confounding factor of perplexity in LLM-as-a-judge evaluation systems, motivated by Wataoka et al.'s finding that LLM evaluators prefer responses with lower…
Deconstructing Agentic Capabilities: A Generalizability Theory Approach to Tool-Use Evaluation
Gilbert Gao — The current evaluation paradigm for agentic Large Language Models (LLMs) mostly rely on a monolithic end-to-end success rate. This metric is convenient, but it mixes up several distinct latent capabilities—such…
Deployment Decision Reliability: A Generalizability-Theory Framework for Sizing Long-Horizon Agent Evaluations
Vasundhara Srinivasan — Enterprise practitioners read agent leaderboards as if they ranked agent capability. We show, across three open agent-trace benchmarks (TheAgentCompany, τ²-bench, and AppWorld), that the agent main effect…
Devil in the Details: Which Prompt Details Are Load-Bearing in Algorithmic Code Generation
Austin Ho — Standard code-generation benchmarks evaluate large language models (LLMs) on behavioral correctness: whether generated code passes a given test suite. But as software is increasingly written by prompting LLMs…
Dimensionality and Measurement Precision in HLE's Multiple-Choice Subset
Mayank Sharma, Savira Nadela, Tyler Matteson — Humanity's Last Exam (HLE) has emerged as a prominent benchmark for evaluating advanced language models, yet widespread adoption has outpaced systematic evaluation of its measurement properties. Most benchmark…
Do LLM Personas Have Personality? A Behavioral Benchmark of Rank-Order and Ipsative Consistency Across Situations
Yikun Chi — LLM persona systems are increasingly used as synthetic respondents, simulated users, social agents, and role-playing characters, but most personality evaluations still ask whether a model can report a prompted…
Do Medical QA Benchmarks Measure Medically Grounded Reasoning? A Natural Language Autoencoder Study of Latent Explanations
Blake Edward Masters — Medical multiple-choice question answering benchmarks are often reported as evidence that a language model has acquired medically grounded reasoning. Final-answer accuracy is useful, but it is an incomplete…
Do Strong Solvers Make Good Judges? An IRT Analysis of LLM-as-a-Judge Quality
Sze Heng Douglas Kwok, Allison Sara John, Jiecong Tan — The LLM-as-a-judge paradigm has become a widely adopted approach for scalable model evaluation, yet it remains poorly understood whether a model's task-solving ability predicts its judgment quality. We…
Does knowing partisan affiliation change GPT-4's plan to persuade?
Emily Zou — This paper asks whether knowing a target's partisan affiliation changes GPT-4's plan to persuade, situating the question in research on micro-targeting — the tailoring of persuasive strategies to information…
Does the Proxy Preserve the Decision? Auditing Multiple-Choice Reformulations of Incremental AI Benchmarks
Ankit Aggarwal — This paper studies a common evaluation shortcut: replacing an open-ended, sequential task with a proxy that is easier to score. The testbed is pyramidal quizbowl, wherein a tossup question is revealed clue by…
ErrorNodeBench-Interference: Why Bad-Entry Rate Misses Memory-Collapse Failures in LLM Consolidation
Yuxiang Liu, Tongchang Zhang — Persistent textual memory is a popular alternative to weight updates for long-running LLM agents, but the rewriting step itself can corrupt useful evidence. We present ErrorNodeBench-Interference, a compact…
Evaluating LLMs as Poker Players: An Item-Response Theory and Q-Matrix Analysis of PokerBench
Abhinav Sattiraju — PokerBench, a recent static benchmark for evaluating large language models on no-limit Texas Hold'em poker decision-making, reports a single accuracy score per model and implicitly treats "poker skill" as a…
GLMM G-Theory and D-Study for AI Benchmarking
Jacob Daniel Householder — We apply G-theory and D-study analysis to three AI benchmarks (HELM Instruct, BiGGen-Bench, WildBench) under three likelihoods: the standard linear mixed model, a cumulative-logit GLMM, and a heteroscedastic…
How Many Items Do You Really Need? IRT-Based Redundancy Analysis of LLM Benchmarks
Dinesh Katupputhur Ramprasath — LLM benchmarks contain hundreds to thousands of evaluation items, yet it remains unclear how many are actually necessary to produce reliable model rankings. We apply Item Response Theory (IRT) to six LLM…
Identifying Representative Subsets of Evaluation Datasets
Anthony Maltsev — Evaluating large language models in deployment requires assessing (model, quantization, hardware) tuples, not base models alone. Full-benchmark evaluation of every tuple is expensive; we study methods for…
Inference Method as a Measurement Facet: Benchmark Validity Under Speculative Decoding
Rami Ratl Mrad — Benchmark scores are treated as properties of models, but they are also products of a full evaluation pipeline. Each component of that pipeline—prompting, decoding, answer extraction, scoring—is a potential…
Investigating Single-Turn Evaluation as a Predictor of Multi-Turn Child Safety
Tracy Y Wei — Child-AI safety benchmarks almost universally rely on single-turn evaluation, but real interactions are conversational. I re-score 225 multi-turn conversations from the ChildSafe corpus (Claude Sonnet 4,…
Latent Capability Estimation Under Sparse AI Evaluation
Jacob Alan Rubenstein, Shane Robinson Mion, Vania Chow — Item Response Theory (IRT) and low-rank matrix completion are both used to estimate latent model capability from (model, item) response matrices, but they are typically benchmarked on dense, curated data rather…
Learning Prompt-Invariant Measurement Models via Sparse Twin-Tower Architectures
Pratham Soni — The paper addresses the brittleness of LLM benchmark evaluation, where measured performance varies with superficial changes to prompt wording, formatting, or answer presentation even when the underlying…