Reasoning

Few-Shot TTT (BIG-Bench Hard)

Per-item BIG-Bench Hard predictions for Llama-3.1-8B-Instruct under zero-shot, few-shot ICL, and test-time training (TTT) interventions.

1,310items
11subjects
MITlicense
reasoningdomain
textmodality
item-level responses released
Saturation status: Unknown

Response matrix

Fit to width. Hover for subject & item; click a cell for details.

Few-Shot TTT (BIG-Bench Hard) response matrix: AI models (rows) against items (columns)
Correct (1)Incorrect (0)Unobserved

Scale: 1 = correct · 0 = incorrect

Subjects

  1. 1Llama-3.1-8B-Instruct + TTT (10-shot, main)0.5822
  2. 2Llama-3.1-8B-Instruct + TTT (10-shot, majority vote)0.5822
  3. 3Llama-3.1-8B-Instruct + Shared TTT (10-shot)0.5763
  4. 4Llama-3.1-8B-Instruct + TTT (10-shot, no shuffle)0.5644
  5. 5Llama-3.1-8B-Instruct + TTT (10-shot, no shuffle, majority vote)0.557
  6. 6Llama-3.1-8B-Instruct + TTT (10-shot, loss on all tokens)0.5467
  7. 7Llama-3.1-8B-Instruct (10-shot ICL)0.5274
  8. 8Llama-3.1-8B-Instruct + Shared E2E direct I/O (10-shot)0.5244
  9. 9Llama-3.1-8B-Instruct + TTT (10-shot, loss on last output)0.5215
  10. 10Llama-3.1-8B-Instruct (10-shot ICL, majority vote)0.5081
  11. 11Llama-3.1-8B-Instruct (zero-shot)0.4267