Reasoning
PlanBench
PlanBench: evaluating LLMs on automated-planning and reasoning-about-change tasks over classic IPC domains (Blocksworld, Logistics, Sokoban) plus mystery / obfuscated / randomized / unsolvable variants. We ingest the REAL per-(model, item) graded outputs from the maintained official repo: each response is one LLM's binary correctness on one planning instance (plan generation), graded by the VAL plan verifier or the run's recorded correctness flag. Items are the real natural-language planning prompts; correct_answer is the ground-truth plan.
10,703items
17subjects
MITlicense
reasoningdomain
textmodality
item-level responses released
Saturation status: Unknown
Response matrix
Fit to width. Hover for subject & item; click a cell for details.

Correct (1)Incorrect (0)Unobserved
Scale: 1 = correct · 0 = incorrect
Subjects
- 1gemini-2.0-flash-thinking-exp0.678
- 2deepseek-r1-api0.5688
- 3o1-preview_chat0.4464
- 4claude-3.5-sonnet_aws0.289
- 5claude-3-opus0.272
- 6llama-3.1-405b_aws0.2696
- 7o1-mini_chat0.2643
- 8gemini-1.5-pro0.2005
- 9gpt-4o_chat0.1616
- 10gpt-4-turbo_chat0.1603
- 11gpt-4_chat0.1494
- 12llama3-70b-8192_groq0.1239
- 13qwen-qwq0.118
- 14gemini-1.5-flash0.1003
- 15gpt-3.5-turbo-instruct0.0345
- 16gemini-pro0.0332
- 17gpt-4o-mini-2024-07-18_chat0