Reasoning

PlanBench

PlanBench: evaluating LLMs on automated-planning and reasoning-about-change tasks over classic IPC domains (Blocksworld, Logistics, Sokoban) plus mystery / obfuscated / randomized / unsolvable variants. We ingest the REAL per-(model, item) graded outputs from the maintained official repo: each response is one LLM's binary correctness on one planning instance (plan generation), graded by the VAL plan verifier or the run's recorded correctness flag. Items are the real natural-language planning prompts; correct_answer is the ground-truth plan.

10,703items
17subjects
MITlicense
reasoningdomain
textmodality
item-level responses released
Saturation status: Unknown

Response matrix

Fit to width. Hover for subject & item; click a cell for details.

PlanBench response matrix: AI models (rows) against items (columns)
Correct (1)Incorrect (0)Unobserved

Scale: 1 = correct · 0 = incorrect

Subjects

  1. 1gemini-2.0-flash-thinking-exp0.678
  2. 2deepseek-r1-api0.5688
  3. 3o1-preview_chat0.4464
  4. 4claude-3.5-sonnet_aws0.289
  5. 5claude-3-opus0.272
  6. 6llama-3.1-405b_aws0.2696
  7. 7o1-mini_chat0.2643
  8. 8gemini-1.5-pro0.2005
  9. 9gpt-4o_chat0.1616
  10. 10gpt-4-turbo_chat0.1603
  11. 11gpt-4_chat0.1494
  12. 12llama3-70b-8192_groq0.1239
  13. 13qwen-qwq0.118
  14. 14gemini-1.5-flash0.1003
  15. 15gpt-3.5-turbo-instruct0.0345
  16. 16gemini-pro0.0332
  17. 17gpt-4o-mini-2024-07-18_chat0