Safety

HarmBench

HarmBench measures whether a model refuses harmful requests, covering behaviors such as cybercrime, misinformation and chemical or biological harm, with a classifier judging each reply as safe or unsafe.

400items
81subjects
Apache-2.0license
safetydomain
textmodality
item-level responses released
Saturation status: Yes

Response matrix

Responses come from Stanford CRFM's HELM Safety leaderboard (v1.17.0), which prompts each behavior directly and grades with its own safety classifier; the original HarmBench evaluation wraps behaviors in 16 attack methods and judges with its own Llama-2-13B classifier. Complements the standalone harmbench build.

Rasch analysis p = σ(θ − z)

32,400 responses, 80/20 split over cells · 81 subjects · 400 items

1 of 81 subjects answered every item alike, so its θ is unbounded and sits pegged to the column edge.

AUC train
0.933
AUC test
0.924

Loading session strips…

Correct (1)Incorrect (0)Unobserved

Scale: 1 = correct · 0 = incorrect