Notation

This appendix collects the notation used throughout the book. Symbols are grouped thematically; the Introduced column indicates the chapter where each symbol first appears.

General

Symbol Meaning Domain Introduced
\(N\) Number of persons / models \(\mathbb{N}\) 3  Models
\(M\) Number of items / questions \(\mathbb{N}\) 3  Models
\(Y_{ij}\) Binary response of model \(i\) to item \(j\) (0 = incorrect, 1 = correct) \(\{0,1\}\) 3  Models
\(S_i = \sum_j Y_{ij}\) Sum score (total correct) for model \(i\) \(\{0, 1, \ldots, M\}\) 3  Models
\(\sigma(x) = \frac{1}{1+e^{-x}}\) Logistic sigmoid function \((0,1)\) 3  Models
\(\Phi(x)\) Standard normal CDF; also a generic monotone link in Additive Conjoint Measurement \((0,1)\) 3  Models

Item Response Theory

Symbol Meaning Domain Introduced
\(\theta_i\) or \(U_i\) Latent ability of model \(i\) \(\mathbb{R}\) 3  Models
\(\beta_j\) or \(V_j\) Difficulty of item \(j\) \(\mathbb{R}\) 3  Models
\(\boldsymbol{\beta}_j\) Item-parameter vector \((\beta_{j1}, \beta_{j2}, \beta_{j3})\) = (difficulty, discrimination, guessing); \(\beta_{j1} = \beta_j\) \(\mathbb{R}^n\) 3  Models
\(p_{ij}\), \(\hat{p}_{ij}\) Response probability \(P(Y_{ij}=1)\) and its empirical estimate \([0,1]\) 3  Models
\(f, g\) Additive conjoint component functions (ability, difficulty factors) \(\mathbb{R}\) 3  Models
\(\beta_{j2}\) Discrimination parameter of item \(j\) (component of \(\boldsymbol{\beta}_j\); conventionally \(a_j\)) \(\mathbb{R}^+\) 3  Models
\(\beta_{j3}\) Guessing (pseudo-chance) parameter of item \(j\) (component of \(\boldsymbol{\beta}_j\); conventionally \(c_j\)) \([0,1]\) 3  Models
\(d_j\) Discrimination parameter (alternative notation, as in 1PL/2PL) \(\mathbb{R}^+\) 3  Models
\(z_j\) Difficulty parameter (alternative notation) \(\mathbb{R}\) 3  Models
\(I_j(\theta)\) Fisher information for item \(j\) at ability \(\theta\) \(\mathbb{R}^+\) 6  Efficiency
\(\mathcal{I}(\theta)\) Fisher information matrix \(\mathbb{R}^{K \times K}\) 6  Efficiency

Learning and Estimation

Symbol Meaning Domain Introduced
\(\ell(\theta, \beta)\) Log-likelihood function \(\mathbb{R}\) 4  Learning
\(\nabla_\theta \ell\) Gradient of log-likelihood w.r.t. ability parameters \(\mathbb{R}^N\) 4  Learning
\(\pi(\theta)\) Prior distribution over abilities 4  Learning
\(\pi(\beta)\) Prior distribution over difficulties 4  Learning
\(\hat{\theta}_{\text{MLE}}\) Maximum likelihood estimate of ability \(\mathbb{R}^N\) 4  Learning
\(\hat{\theta}_{\text{MAP}}\) Maximum a posteriori estimate of ability \(\mathbb{R}^N\) 4  Learning
\(\eta\) Learning rate \(\mathbb{R}^+\) 4  Learning

Reliability

Symbol Meaning Domain Introduced
\(X_{ij}\) Observed score for model \(i\) on occasion / item \(j\) \(\mathbb{R}\) 5  Reliability
\(T_i\) True score for model \(i\) (CTT) \(\mathbb{R}\) 5  Reliability
\(E_{ij}\) Error component (CTT) \(\mathbb{R}\) 5  Reliability
\(\rho_{XX'}\) Reliability coefficient \([0,1]\) 5  Reliability
\(\alpha\) Cronbach’s alpha \((-\infty, 1]\) 5  Reliability
\(\sigma^2_p, \sigma^2_i, \sigma^2_r\) Variance components: person (model), item, rater \(\mathbb{R}^+\) 5  Reliability
\(G\) Generalizability coefficient \([0,1]\) 5  Reliability
\(n_r, n_i\) Number of raters, items in a D-study design \(\mathbb{N}\) 5  Reliability
\(\kappa\) Cohen’s kappa (inter-rater agreement) \([-1, 1]\) 5  Reliability

Validity

Symbol Meaning Domain Introduced
\(g \in \{0,1\}\) Group membership indicator (DIF analysis) \(\{0,1\}\) 1  Validity
\(\alpha_{MH}\) Mantel-Haenszel odds ratio \(\mathbb{R}^+\) 1  Validity
\(\lambda_k\) \(k\)-th eigenvalue of the correlation matrix \(\mathbb{R}\) 1  Validity
\(\text{MNSQ}_i\) Mean-square fit statistic for item \(i\) \(\mathbb{R}^+\) 1  Validity
\(r_{ij}\) Correlation between trait \(i\) measured by method \(j\) (MTMM) \([-1,1]\) 1  Validity

Causality and Distribution Shift

Symbol Meaning Domain Introduced
\(\text{do}(X = x)\) Intervention setting variable \(X\) to value \(x\) 9  Intervention
\(P^{(s)}, P^{(t)}\) Source (benchmark) / target (deployment) distribution 9  Intervention
\(\pi_0(a \mid x)\) Logging (benchmark) policy \([0,1]\) 9  Intervention
\(\pi(a \mid x)\) Target (deployment) policy \([0,1]\) 9  Intervention
\(w(x)\) Importance weight \(P^{(t)}(x)/P^{(s)}(x)\) \(\mathbb{R}^+\) 9  Intervention
\(\hat{r}(x, a)\) Reward model (e.g., IRT prediction) \([0,1]\) 9  Intervention
\(\hat{V}_{\text{DR}}\) Doubly robust value estimator \(\mathbb{R}\) 9  Intervention
\(C_\alpha(x)\) Conformal prediction set at level \(\alpha\) subset of \(\mathcal{Y}\) 9  Intervention

Information and Mechanism Design

Symbol Meaning Domain Introduced
\(\mathcal{C}\) Construct: the latent capability being measured What Are We Trying to Hide?
\(D\) Operationalizing distribution (formally \(\pi_E\)) distribution on \(F\) What Are We Trying to Hide?
\(D', D''\) Public proxy distributions for \(\mathcal{C}\) distributions on \(F\) Publishing the Construct: The Construct-Validity Protocol
\(S\) Realized evaluation set, \(S \sim D\) \(S \subseteq F\) What Are We Trying to Hide?
\(\mu\) Social-relevance distribution (construct fully operationalized; \(\text{Uniform}(F)\) as a special case) distribution on \(F\) What Are We Trying to Hide?
\(F\) Universe of tasks finite set, \(\|F\| = N\) The Evaluation Game
\(F_E, F_M\) Evaluator’s / builder’s task sets \(F_E, F_M \subseteq F\) The Evaluation Game
\(\pi_E, \pi_M\) Evaluator’s / builder’s sampling distributions distributions on \(F\) The Evaluation Game
\(f(\theta)\) Task performance function \([0,1]\) The Evaluation Game
\(u_E(\theta)\) Evaluator’s utility: \(\mathbb{E}_{f \sim \mu}[f(\theta)]\) \(\mathbb{R}^+\) The Evaluation Game
\(k, k^*\) Number of (optimal) sampled evaluation tasks per round \(\mathbb{N}\) The Evaluation Game
\(\rho, \sigma_c\) Distribution correction rate and bandwidth \((0,1] \times \mathbb{R}^+\) The Evaluation Game
\(\Delta_t\) Residual misalignment: \(\max_\theta u_E(\theta) - u_E(\theta_t)\) \(\mathbb{R}^+\) The Evaluation Game
\(\tilde{\pi}_E\) Builder’s belief about the evaluator’s sampling distribution on \(F\) The Evaluation Game
\(C\) Agent cost (principal-agent model) \(\mathbb{R}^+\) The Evaluation Game
\(b\) Principal’s value from agent effort \(\mathbb{R}^+\) The Evaluation Game

Red-Teaming and Adversarial Evaluation

Symbol Meaning Domain Introduced
\(\theta_j^{(\text{adv})}\) Adversarial robustness of model \(j\) \(\mathbb{R}\) 12  Application
\(\theta_j^{(\text{std})}\) Standard accuracy ability of model \(j\) \(\mathbb{R}\) 12  Application
\(\beta_i^{(\text{atk})}\) Attack strength (difficulty) of adversarial item \(i\) \(\mathbb{R}\) 12  Application
\(\delta\) Perturbation applied to an input \(\mathcal{B}_\epsilon\) 12  Application
\(\alpha_{s,\mathcal{D}}\) Attack success probability under criterion \(s\) and goal distribution \(\mathcal{D}\) \([0,1]\) 12  Application
\(J\) Operational judge of attack success \(\{0,1\}\) 12  Application
\(s\) Oracle (true) success criterion \(\{0,1\}\) 12  Application
\(K\) Number of repeated samples in Top-1 aggregation \(\mathbb{N}\) 12  Application
\(\hat{\mu}_{\text{syn}}\) Mean performance estimated from synthetic items \([0,1]\) 12  Application
\(\hat{\mu}_{\text{human}}\) Mean performance estimated from human-authored items \([0,1]\) 12  Application
\(\hat{\mu}_{\text{PPI}}\) Prediction-powered inference estimator \([0,1]\) 12  Application
\(n\) Number of human-evaluated items \(\mathbb{N}\) 12  Application