Safety

HarmBench

HarmBench standardized red-teaming results: ~29 target models x 140 behaviors x 16 attack methods, with the HarmBench Llama-2-13B classifier verdict and the AdvBench refusal-string heuristic as two separate judges. Covers the standard and contextual behavior categories from the HarmBench website playground.

140items
29subjects
MITlicense
safetydomain
textmodality
item-level responses released
Saturation status: Unknown

Response matrix

Fit to width. Hover for subject & item; click a cell for details.

HarmBench response matrix: AI models (rows) against items (columns)
Correct (1)Incorrect (0)Unobserved

Scale: 1 = correct · 0 = incorrect

Subjects

  1. 1zephyr_7b0.7226
  2. 2starling_7b0.677
  3. 3mixtral_8x7b0.6416
  4. 4openchat_3_5_12100.6371
  5. 5solar_10_7b_instruct0.6313
  6. 6mistral_7b_v20.6202
  7. 7orca_2_7b0.6052
  8. 8koala_13b0.5929
  9. 9koala_7b0.5897
  10. 10orca_2_13b0.5804
  11. 11baichuan2_13b0.5584
  12. 12vicuna_7b_v1_50.5326
  13. 13baichuan2_7b0.5253
  14. 14gpt-3.5-turbo-06130.4897
  15. 15vicuna_13b_v1_50.4863
  16. 16qwen_7b_chat0.4782
  17. 17qwen_14b_chat0.4656
  18. 18qwen_72b_chat0.4248
  19. 19gpt-4-06130.4222
  20. 20gpt-3.5-turbo-11060.3441
  21. 21gemini0.3332
  22. 22gpt-4-1106-preview0.3021
  23. 23zephyr_7b_robust6_data20.2698
  24. 24llama2_70b0.1929
  25. 25llama2_13b0.1849
  26. 26llama2_7b0.1822
  27. 27claude-instant-10.1296
  28. 28claude-2.10.0931
  29. 29claude-20.089