Reasoning

ARC-AGI

ARC-AGI v1 + v2 per-model per-task solve data from the ARC Prize Foundation evaluation repos, plus leaderboard aggregate scores.

520items
73subjects
Apache-2.0license
reasoningdomain
gridmodality
item-level responses released
Saturation status: Unknown

Response matrix

Fit to width. Hover for subject & item; click a cell for details.

ARC-AGI response matrix: AI models (rows) against items (columns)
Correct (1)Incorrect (0)Unobserved

Scale: 1 = correct · 0 = incorrect

Subjects

  1. 1gpt-5-4-pro-xhigh0.9825
  2. 2gpt-5-2-pro-2025-12-11-xhigh0.9775
  3. 3grok-4.20-multi-agent-beta-0309-xhigh0.9674
  4. 4claude-opus-4-8-max0.9626
  5. 5gemini-3-1-pro-preview0.9606
  6. 6gpt-5-4-high0.9575
  7. 7grok-4.20-beta-0309b-reasoning0.955
  8. 8claude-opus-4-6-thinking-120K-max0.9342
  9. 9claude-opus-4-6-thinking-120K-high0.9327
  10. 10claude-opus-4-8-high0.9269
  11. 11gpt-5-4-medium0.92
  12. 12claude-opus-4-8-medium0.9173
  13. 13claude-opus-4-6-thinking-120K-medium0.9077
  14. 14public_eval0.8972
  15. 15gpt-5-2-2025-12-11-thinking-xhigh0.8856
  16. 16claude-opus-4-8-low0.8615
  17. 17gpt-5-2-pro-2025-12-11-high0.8574
  18. 18claude-opus-4-6-thinking-120K-low0.8382
  19. 19gemini-3-deep-think-preview0.8372
  20. 20gpt-5-2-2025-12-11-thinking-high0.8015
  21. 21gpt-5-4-low0.8
  22. 22gpt-5-2-pro-2025-12-11-medium0.7923
  23. 23gpt-5-4-mini-xhigh0.7738
  24. 24gpt-5-pro-2025-10-060.77
  25. 25gemini-3-flash-preview-thinking-high0.7533
  26. 26claude-opus-4-5-20251101-thinking-32k0.7437
  27. 27gemini-3-pro-preview0.7437
  28. 28claude-sonnet-4-5-20250929-thinking-32k0.74
  29. 29gpt-5-4-nano-xhigh0.7385
  30. 30kimi-k2.50.7325
  31. 31gpt-5-2-2025-12-11-thinking-medium0.6981
  32. 32claude-opus-4-5-20251101-thinking-64k0.695
  33. 33gpt-5-4-mini-high0.665
  34. 34gpt-5-1-2025-11-13-thinking-high0.6462
  35. 35claude-sonnet-4-5-20250929-thinking-16k0.6375
  36. 36claude-haiku-4-5-20251001-thinking-32k0.6325