General
AppWorld
AppWorld: interactive coding agent benchmark. Per-(agent, task) pass/fail decrypted from the public leaderboard's experiment bundles; real per-task natural-language instructions from the released task data.
521items
8subjects
Apache-2.0license
agents_and_tool_usedomain
textmodality
gui_screenshotmodality
item-level responses released
Saturation status: Unknown
Response matrix
Fit to width. Hover for subject & item; click a cell for details.

Correct (1)Incorrect (0)Unobserved
Scale: 1 = correct · 0 = incorrect
Subjects
- 1Qwen3-14B0.7316
- 2gpt-4.1-2025-04-140.6205
- 3Qwen-2.5-32B-Instruct0.5453
- 4gpt-4o-2024-08-060.4735
- 5gpt-4o-2024-05-130.2697
- 6gpt-4-turbo-2024-04-090.1821
- 7meta-llama/Llama-3-70b-chat-hf0.0821
- 8deepseek-ai/deepseek-coder-33b-instruct0.0433