General

AgentBoard

AgentBoard: analytical evaluation of multi-turn LLM agents across 9 interactive environments (alfworld, scienceworld, babyai, jericho, pddl, webarena, webshop, tool-query, tool-operation). Each item is a task instance; response is per-example success_rate (1 if the agent completed the full task).

867items
13subjects
GPL-3.0license
agents_and_tool_usedomain
textmodality
item-level responses released
Saturation status: Unknown

Response matrix

Fit to width. Hover for subject & item; click a cell for details.

AgentBoard response matrix: AI models (rows) against items (columns)
Correct (1)Incorrect (0)Unobserved

Scale: 1 = correct · 0 = incorrect

Subjects

  1. 1gpt-40.4071
  2. 2claude20.2619
  3. 3gpt-35-turbo0.2144
  4. 4deepseek-67b0.2099
  5. 5text-davinci-0030.1571
  6. 6gpt-35-turbo-16k0.1542
  7. 7codellama-13b0.1147
  8. 8codellama-34b0.0998
  9. 9lemur-70b0.0751
  10. 10vicuna-13b-16k0.0741
  11. 11llama2-70b0.07
  12. 12mistral-7b0.0619
  13. 13llama2-13b0.0158