General

AppWorld

AppWorld: interactive coding agent benchmark. Per-(agent, task) pass/fail decrypted from the public leaderboard's experiment bundles; real per-task natural-language instructions from the released task data.

521items
8subjects
Apache-2.0license
agents_and_tool_usedomain
textmodality
gui_screenshotmodality
item-level responses released
Saturation status: Unknown

Response matrix

Fit to width. Hover for subject & item; click a cell for details.

AppWorld response matrix: AI models (rows) against items (columns)
Correct (1)Incorrect (0)Unobserved

Scale: 1 = correct · 0 = incorrect

Subjects

  1. 1Qwen3-14B0.7316
  2. 2gpt-4.1-2025-04-140.6205
  3. 3Qwen-2.5-32B-Instruct0.5453
  4. 4gpt-4o-2024-08-060.4735
  5. 5gpt-4o-2024-05-130.2697
  6. 6gpt-4-turbo-2024-04-090.1821
  7. 7meta-llama/Llama-3-70b-chat-hf0.0821
  8. 8deepseek-ai/deepseek-coder-33b-instruct0.0433