Reasoning

OOD-Prediction (Internal Causal Mechanisms)

OOD-Prediction correctness dataset (Huang et al., ICML 2025): released prompts for 5 tasks (IOI, PriceTag, RAVEL, MMLU, Unlearn Harry Potter) under in-distribution and OOD settings, each prompt labelled correct/wrong by the target model. subject = target model; item = the prompt; response = 1 correct / 0 wrong.

142,805items
2subjects
MITlicense
knowledgedomain
reasoningdomain
textmodality
item-level responses released
Saturation status: Unknown

Response matrix

Fit to width. Hover for subject & item; click a cell for details.

OOD-Prediction (Internal Causal Mechanisms) response matrix: AI models (rows) against items (columns)
Correct (1)Incorrect (0)Unobserved

Scale: 1 = correct · 0 = incorrect

Subjects

  1. 1Meta-Llama-3-8B-Instruct0.5105
  2. 2gpt20.5