Safety
HELM HarmBench
HELM HarmBench: per-(model, behavior) safety_score in {0,1} (1=safe) from HELM's safety classifier. 87 models x HarmBench harmful behaviors. HELM-sourced; complements the standalone harmbench build.
400items
81subjects
Apache-2.0license
safetydomain
textmodality
item-level responses released
Saturation status: Unknown
Response matrix
Fit to width. Hover for subject & item; click a cell for details.

Correct (1)Incorrect (0)Unobserved
Scale: 1 = correct · 0 = incorrect
Subjects
- 1openai/gpt-oss-120b1
- 2openai/gpt-oss-20b0.9825
- 3openai/o3-2025-04-160.9825
- 4anthropic/claude-sonnet-4-20250514-thinking-10k0.9725
- 5anthropic/claude-3-5-sonnet-202406200.9675
- 6openai/o4-mini-2025-04-160.9675
- 7openai/gpt-5-nano-2025-08-070.965
- 8moonshotai/kimi-k2-instruct0.965
- 9openai/gpt-5-2025-08-070.955
- 10openai/gpt-5.1-2025-11-130.955
- 11openai/o1-2024-12-170.955
- 12openai/gpt-5-mini-2025-08-070.9525
- 13openai/gpt-4.5-preview-2025-02-270.9525
- 14openai/o3-mini-2025-01-310.95
- 15anthropic/claude-3-opus-202402290.9475
- 16anthropic/claude-haiku-4-5-202510010.9375
- 17anthropic/claude-3-sonnet-202402290.9275
- 18openai/gpt-4.1-2025-04-140.9075
- 19anthropic/claude-sonnet-4-202505140.905
- 20anthropic/claude-sonnet-4-5-202509290.8725
- 21openai/gpt-4-turbo-2024-04-090.87
- 22anthropic/claude-3-haiku-202403070.8525
- 23openai/gpt-4.1-nano-2025-04-140.8525
- 24anthropic/claude-opus-4-20250514-thinking-10k0.8425
- 25openai/gpt-4.1-mini-2025-04-140.84
- 26openai/gpt-4o-mini-2024-07-180.8225
- 27anthropic/claude-opus-4-202505140.8125
- 28allenai/olmo-2-0325-32b-instruct0.805
- 29ibm/granite-4.0-h-small-with-guardian0.7975
- 30openai/gpt-4o-2024-05-130.7925
- 31ibm/granite-4.0-micro-with-guardian0.79
- 32openai/o1-mini-2024-09-120.7875
- 33qwen/qwen3-next-80b-a3b-thinking0.7825
- 34writer/palmyra-fin0.7775
- 35anthropic/claude-3-7-sonnet-202502190.745
- 36qwen/qwen3-235b-a22b-instruct-2507-fp80.74