Safety
HarmBench
HarmBench measures whether a model refuses harmful requests, covering behaviors such as cybercrime, misinformation and chemical or biological harm, with a classifier judging each reply as safe or unsafe.
Response matrix
Responses come from Stanford CRFM's HELM Safety leaderboard (v1.17.0), which prompts each behavior directly and grades with its own safety classifier; the original HarmBench evaluation wraps behaviors in 16 attack methods and judges with its own Llama-2-13B classifier. Complements the standalone harmbench build.
32,400 responses, 80/20 split over cells · 81 subjects · 400 items
1 of 81 subjects answered every item alike, so its θ is unbounded and sits pegged to the column edge.
Loading session strips…
Scale: 1 = correct · 0 = incorrect