From metric selection to metric discovery in AI evaluation
Abstract
Evaluation metrics are a cornerstone of the AI ecosystem, driving the decisions of developers, consumers, and investors alike. This talk will address two challenges in evaluation that are exacerbated in the development of AI systems. First is a computational challenge: the explosion of metrics to measure complex and hard-to-define capabilities in AI systems has led to rapidly increasing computational costs to evaluation. To mitigate this, we'll discuss metric selection, with algorithms for efficient and provably representative selection of metrics based in social choice theory. Second is an informational challenge: evaluators continue to face a fundamental information problem of not knowing whether they could be missing some important metrics entirely. Thus, moving beyond selection, we'll then discuss metric discovery through an economic model of incentives for agents to reveal unknown unknown metrics under information asymmetry.
About the speaker
Serena is an Assistant Professor in Computer Science at the University of British Columbia and a Canada CIFAR AI Chair. She was a Postdoctoral Fellow at Harvard University hosted by Ariel Procaccia, and completed her PhD at UC Berkeley in 2024 advised by Michael Jordan. Serena’s research focuses on understanding and improving the long term societal impacts of AI by rethinking algorithms and their surrounding incentives and practices. Her recent work concerns evaluation processes for AI systems and beyond, including robustness, incentives, and representation in a multi-stakeholder ecosystem. Her interdisciplinary research agenda combines ideas from machine learning, statistics, economics, and the social sciences.















