General
AgentBoard
AgentBoard: analytical evaluation of multi-turn LLM agents across 9 interactive environments (alfworld, scienceworld, babyai, jericho, pddl, webarena, webshop, tool-query, tool-operation). Each item is a task instance; response is per-example success_rate (1 if the agent completed the full task).
867items
13subjects
GPL-3.0license
agents_and_tool_usedomain
textmodality
item-level responses released
Saturation status: Unknown
Response matrix
Fit to width. Hover for subject & item; click a cell for details.

Correct (1)Incorrect (0)Unobserved
Scale: 1 = correct · 0 = incorrect
Subjects
- 1gpt-40.4071
- 2claude20.2619
- 3gpt-35-turbo0.2144
- 4deepseek-67b0.2099
- 5text-davinci-0030.1571
- 6gpt-35-turbo-16k0.1542
- 7codellama-13b0.1147
- 8codellama-34b0.0998
- 9lemur-70b0.0751
- 10vicuna-13b-16k0.0741
- 11llama2-70b0.07
- 12mistral-7b0.0619
- 13llama2-13b0.0158