Reasoning
Few-Shot TTT (BIG-Bench Hard)
Released five-example previews for BIG-Bench Hard evaluations of Llama-3.1-8B-Instruct under eleven zero-shot, few-shot ICL and test-time training interventions. These previews cover 7,425 observations, not the complete evaluation sets used to calculate the accompanying task accuracies.
1,310items
11subjects
MITlicense
reasoningdomain
textmodality
item-level responses released
Saturation status: No
Response matrix
Loading response matrix…
Correct (1)Incorrect (0)Unobserved
Scale: 1 = correct · 0 = incorrect