Interactive Measurement
Evaluation of policies, long-horizon behavior, tool-using agents, and statistical validity when observations are non-i.i.d. and path-dependent.
Treating evaluation as a first-class scientific problem for frontier AI.
We are excited to announce the inaugural AI Measurement Science Workshop, co-located with COLM 2026 at the Hilton San Francisco Union Square. The workshop is a one-day, in-person event on October 9, 2026, bringing together researchers from machine learning, statistics, psychometrics, economics, and policy to develop a unified perspective on measurement for frontier AI systems.
As AI systems become increasingly embedded in real-world workflows, existing evaluation paradigms face fundamental limitations. Contemporary benchmarks typically assume i.i.d. samples, stationary concepts, and passive models. These assumptions are increasingly violated in modern systems, which interact with environments, adapt to feedback, and are optimized against known evaluation signals.
This workshop frames evaluation as a problem of AI measurement science. We focus on three tightly coupled challenges (measurement of interactive systems, measurement under strategic optimization, and measurement under non-stationarity) and aim to define common abstractions and open research directions for this emerging area.
Submissions from educational testing, psychometrics, and related fields are strongly encouraged.
We invite work across theory, methodology, and systems, including but not limited to the three tightly coupled challenges below.
Evaluation of policies, long-horizon behavior, tool-using agents, and statistical validity when observations are non-i.i.d. and path-dependent.
Evaluation as a game between learner and evaluator; robustness to gaming and contamination; incentive-compatible, private, randomized, or adaptive leaderboards.
Covariate, label, and construct drift; adaptive test design, recalibration, and longitudinal monitoring of deployed systems.
We also invite write-ups from participants of the Predictive AI Evaluation Challenge. Papers should describe each team's approach to predicting model responses from sparse observations, with methods, results, and lessons learned. Submissions follow the same 4–8 page COLM format as the research track and are reviewed on a later timeline so that final competition results can be included. The Best Competition Paper award will be presented alongside the workshop's research awards.
Competition track deadline: August 15, 2026
Short, focused contributions in COLM format. Review is double-blind.
All deadlines are in Anywhere on Earth (AoE) time.
Submissions open now on OpenReview. Deadline June 23, 2026.
Invited speakers from academia, industry, and policy.

UIUC
Assistant professor working on agentic benchmarks, AI security evaluation, and ML deployment systems.

Harvard / UBC
Postdoctoral researcher on long-term AI impacts, counterfactual metrics for social welfare, and robust distillation.

UC Berkeley · Transluce
Professor and CEO of Transluce; LLM safety and reward alignment.

Brookings
Led NIST's AI Risk Management Framework and served as chief technologist of the U.S. AI Safety Institute; named one of TIME's 100 most influential in AI.
Tentative program for the one-day workshop. October 9, 2026
Opening remarks
Sanmi Koyejo
Invited Talk: Daniel Kang
The Predictive AI Evaluation Challenge
Sang Truong
Invited Talk: Serena Wang
Poster session 1 (40 posters)
Invited Talk: Jacob Steinhardt
Networking lunch
Invited Talk: Elham Tabassi
Contributed talks (selected from submissions)
Poster session 2 (40 posters)
Debate, panel discussion, and group discussion
Mentorship session
Berivan Isik
Closing remarks
Sanmi Koyejo

Stanford University

Stanford University

Schmidt Sciences / MIT

Stanford University

Google DeepMind

Carnegie Mellon University

Microsoft Research Asia
Get involved
Reach out to the organizing committee. We would love to hear from you.