Reasoning
ARC-AGI
ARC-AGI v1 + v2 per-model per-task solve data from the ARC Prize Foundation evaluation repos, plus leaderboard aggregate scores.
520items
73subjects
Apache-2.0license
reasoningdomain
gridmodality
item-level responses released
Saturation status: Unknown
Response matrix
Fit to width. Hover for subject & item; click a cell for details.

Correct (1)Incorrect (0)Unobserved
Scale: 1 = correct · 0 = incorrect
Subjects
- 1gpt-5-4-pro-xhigh0.9825
- 2gpt-5-2-pro-2025-12-11-xhigh0.9775
- 3grok-4.20-multi-agent-beta-0309-xhigh0.9674
- 4claude-opus-4-8-max0.9626
- 5gemini-3-1-pro-preview0.9606
- 6gpt-5-4-high0.9575
- 7grok-4.20-beta-0309b-reasoning0.955
- 8claude-opus-4-6-thinking-120K-max0.9342
- 9claude-opus-4-6-thinking-120K-high0.9327
- 10claude-opus-4-8-high0.9269
- 11gpt-5-4-medium0.92
- 12claude-opus-4-8-medium0.9173
- 13claude-opus-4-6-thinking-120K-medium0.9077
- 14public_eval0.8972
- 15gpt-5-2-2025-12-11-thinking-xhigh0.8856
- 16claude-opus-4-8-low0.8615
- 17gpt-5-2-pro-2025-12-11-high0.8574
- 18claude-opus-4-6-thinking-120K-low0.8382
- 19gemini-3-deep-think-preview0.8372
- 20gpt-5-2-2025-12-11-thinking-high0.8015
- 21gpt-5-4-low0.8
- 22gpt-5-2-pro-2025-12-11-medium0.7923
- 23gpt-5-4-mini-xhigh0.7738
- 24gpt-5-pro-2025-10-060.77
- 25gemini-3-flash-preview-thinking-high0.7533
- 26claude-opus-4-5-20251101-thinking-32k0.7437
- 27gemini-3-pro-preview0.7437
- 28claude-sonnet-4-5-20250929-thinking-32k0.74
- 29gpt-5-4-nano-xhigh0.7385
- 30kimi-k2.50.7325
- 31gpt-5-2-2025-12-11-thinking-medium0.6981
- 32claude-opus-4-5-20251101-thinking-64k0.695
- 33gpt-5-4-mini-high0.665
- 34gpt-5-1-2025-11-13-thinking-high0.6462
- 35claude-sonnet-4-5-20250929-thinking-16k0.6375
- 36claude-haiku-4-5-20251001-thinking-32k0.6325