Skip to main content

Search AIMS

Find pages, publications, events, course projects, and software. Results update as you type; press Enter to open the first result.

Reasoning

Few-Shot TTT (BIG-Bench Hard)

Released five-example previews for BIG-Bench Hard evaluations of Llama-3.1-8B-Instruct under eleven zero-shot, few-shot ICL and test-time training interventions. These previews cover 7,425 observations, not the complete evaluation sets used to calculate the accompanying task accuracies.

1,310items
11subjects
MITlicense
reasoningdomain
textmodality
item-level responses released
Saturation status: No

Response matrix

Loading response matrix…

Correct (1)Incorrect (0)Unobserved

Scale: 1 = correct · 0 = incorrect