Science

SuperGPQA

SuperGPQA: graduate-level multiple-choice knowledge & reasoning across 285 disciplines (26,529 questions, avg 9.67 options/question). Per-(model, question) verdicts for 49 LLMs from the upstream zero-shot records dump; response is the binary correct/incorrect outcome.

26,529items
49subjects
CC-BY-NC-SA-4.0license
knowledgedomain
reasoningdomain
sciencedomain
textmodality
item-level responses released
Saturation status: Unknown

Response matrix

Fit to width. Hover for subject & item; click a cell for details.

SuperGPQA response matrix: AI models (rows) against items (columns)
Correct (1)Incorrect (0)Unobserved

Scale: 1 = correct · 0 = incorrect

Subjects

  1. 1DeepSeek-R10.6182
  2. 2o1-2024-12-170.6024
  3. 3DeepSeek-R1-Zero0.6024
  4. 4o3-mini-2025-01-31-high0.5522
  5. 5Doubao-1.5-pro-32k-2501150.5509
  6. 6o3-mini-2025-01-31-medium0.5269
  7. 7Doubao-1.5-pro-32k-2412250.5093
  8. 8qwen-max-2025-01-250.5008
  9. 9claude-3-5-sonnet-202410220.4816
  10. 10o3-mini-2025-01-31-low0.4803
  11. 11gemini-2.0-flash0.4773
  12. 12DeepSeek-V30.474
  13. 13o1-mini-2024-09-120.4522
  14. 14MiniMax-Text-010.4511
  15. 15gpt-4o-2024-11-200.444
  16. 16QwQ0.4359
  17. 17Llama-3.1-405B-Instruct0.4314
  18. 18gpt-4o-2024-08-060.4164
  19. 19Qwen2.5-72B-Instruct0.4075
  20. 20Mistral-Large-Instruct-24110.4065
  21. 21qwen-max-2024-09-190.3996
  22. 22gpt-4o-2024-05-130.3976
  23. 23Qwen2.5-32B-Instruct0.3876
  24. 24Llama-3.3-70B-Instruct0.3769
  25. 25phi-40.3765
  26. 26Qwen2.5-14B-Instruct0.3515
  27. 27Llama-3.1-70B-Instruct0.3486
  28. 28Yi-Lighting0.3342
  29. 29Mixtral-8x22B-Instruct-v0.10.2923
  30. 30Qwen2.5-7B-Instruct0.2878
  31. 31gemma-2-27b-it0.2743
  32. 32Yi-1.5-34B-Chat0.2603
  33. 33Mistral-Small-Instruct-24090.2589
  34. 34gemma-2-9b-it0.2404
  35. 35Qwen2.5-3B-Instruct0.2331
  36. 36Yi-1.5-9B-Chat0.2317