Safety

AgentDojo

AgentDojo: prompt-injection evaluation across many tool-using agents and environments.

1,081items
29subjects
MITlicense
agents_and_tool_usedomain
safetydomain
textmodality
item-level responses released
Saturation status: No

Response matrix

Fit to width. Hover for subject & item; click a cell for details.

AgentDojo response matrix: AI models (rows) against items (columns)
Correct (1)Incorrect (0)Unobserved

Scale: 1 = correct · 0 = incorrect

Subjects

  1. 1gpt-4o-2024-05-13 (spotlighting_with_delimiting)0.5073
  2. 2gpt-4o-2024-05-13 (repeat_user_prompt)0.5066
  3. 3gpt-4-0125-preview0.5037
  4. 4claude-3-5-sonnet-202406200.4531
  5. 5gpt-4o-2024-05-130.4418
  6. 6gpt-4-turbo-2024-04-090.4334
  7. 7gpt-4o-mini-2024-07-180.4165
  8. 8claude-3-7-sonnet-202502190.4002
  9. 9claude-3-sonnet-20240229 (repeat_user_prompt)0.4
  10. 10claude-3-5-sonnet-202410220.3982
  11. 11Meta-SecAlign-70B (repeat_user_prompt)0.3543
  12. 12gpt-4o-2024-05-13 (tool_filter)0.3499
  13. 13claude-3-opus-202402290.3492
  14. 14gemini-1.5-pro-0020.3455
  15. 15Meta-SecAlign-70B0.3414
  16. 16claude-3-sonnet-202402290.3192
  17. 17gemini-1.5-pro-0010.3038
  18. 18gemini-2.0-flash-exp0.3009
  19. 19meta-llama_Llama-3.3-70B-Instruct0.2809
  20. 20gemini-2.0-flash-0010.2587
  21. 21gemini-1.5-flash-0010.2504
  22. 22gpt-3.5-turbo-01250.2452
  23. 23meta-llama_Llama-3-70b-chat-hf0.2335
  24. 24meta-llama_Llama-3.3-70B-Instruct (repeat_user_prompt)0.233
  25. 25claude-3-haiku-202403070.2291
  26. 26gemini-1.5-flash-0020.1991
  27. 27command-r0.1867
  28. 28gpt-4o-2024-05-13 (transformers_pi_detector)0.1684
  29. 29command-r-plus0.1574