Safety
HarmBench
HarmBench standardized red-teaming results: ~29 target models x 140 behaviors x 16 attack methods, with the HarmBench Llama-2-13B classifier verdict and the AdvBench refusal-string heuristic as two separate judges. Covers the standard and contextual behavior categories from the HarmBench website playground.
140items
29subjects
MITlicense
safetydomain
textmodality
item-level responses released
Saturation status: Unknown
Response matrix
Fit to width. Hover for subject & item; click a cell for details.

Correct (1)Incorrect (0)Unobserved
Scale: 1 = correct · 0 = incorrect
Subjects
- 1zephyr_7b0.7226
- 2starling_7b0.677
- 3mixtral_8x7b0.6416
- 4openchat_3_5_12100.6371
- 5solar_10_7b_instruct0.6313
- 6mistral_7b_v20.6202
- 7orca_2_7b0.6052
- 8koala_13b0.5929
- 9koala_7b0.5897
- 10orca_2_13b0.5804
- 11baichuan2_13b0.5584
- 12vicuna_7b_v1_50.5326
- 13baichuan2_7b0.5253
- 14gpt-3.5-turbo-06130.4897
- 15vicuna_13b_v1_50.4863
- 16qwen_7b_chat0.4782
- 17qwen_14b_chat0.4656
- 18qwen_72b_chat0.4248
- 19gpt-4-06130.4222
- 20gpt-3.5-turbo-11060.3441
- 21gemini0.3332
- 22gpt-4-1106-preview0.3021
- 23zephyr_7b_robust6_data20.2698
- 24llama2_70b0.1929
- 25llama2_13b0.1849
- 26llama2_7b0.1822
- 27claude-instant-10.1296
- 28claude-2.10.0931
- 29claude-20.089