Reasoning
Few-Shot TTT (BIG-Bench Hard)
Per-item BIG-Bench Hard predictions for Llama-3.1-8B-Instruct under zero-shot, few-shot ICL, and test-time training (TTT) interventions.
1,310items
11subjects
MITlicense
reasoningdomain
textmodality
item-level responses released
Saturation status: Unknown
Response matrix
Fit to width. Hover for subject & item; click a cell for details.

Correct (1)Incorrect (0)Unobserved
Scale: 1 = correct · 0 = incorrect
Subjects
- 1Llama-3.1-8B-Instruct + TTT (10-shot, main)0.5822
- 2Llama-3.1-8B-Instruct + TTT (10-shot, majority vote)0.5822
- 3Llama-3.1-8B-Instruct + Shared TTT (10-shot)0.5763
- 4Llama-3.1-8B-Instruct + TTT (10-shot, no shuffle)0.5644
- 5Llama-3.1-8B-Instruct + TTT (10-shot, no shuffle, majority vote)0.557
- 6Llama-3.1-8B-Instruct + TTT (10-shot, loss on all tokens)0.5467
- 7Llama-3.1-8B-Instruct (10-shot ICL)0.5274
- 8Llama-3.1-8B-Instruct + Shared E2E direct I/O (10-shot)0.5244
- 9Llama-3.1-8B-Instruct + TTT (10-shot, loss on last output)0.5215
- 10Llama-3.1-8B-Instruct (10-shot ICL, majority vote)0.5081
- 11Llama-3.1-8B-Instruct (zero-shot)0.4267