AI Cartography: Mapping the Latent Landscape of AI Benchmark Ecosystems
Abstract
While aggregate leaderboard scores drive AI development, they contain substantial measurement noise whose sources and magnitudes remain unquantified, making it unclear when rankings reflect genuine capability differences versus evaluation artifacts. We introduce a framework for measuring the latent landscape in AI benchmark ecosystems. Applying Confirmatory Factor Analysis and Generalizability Theory to 4,000+ models from the Open LLM Leaderboard, we decompose sources of ranking variance, uncovering evidence of local dependence among leaderboard items, showing that contributor metadata explains more rank-relevant variance than architecture or deployment categories, and finding that the latent general-factor slope is far more stable than manifest-score scaling laws. We provide actionable diagnostics to determine how benchmark rankings can be trusted and how benchmark design can be improved.












