Safety
HELM SimpleSafetyTests
HELM SimpleSafetyTests: per-(model, prompt) safety_score in {0,1} (1=safe) from HELM's safety classifier. 87 models x 100 severe-harm prompts (self-harm, physical harm, illegal items, scams, child abuse).
100items
81subjects
Apache-2.0license
safetydomain
textmodality
item-level responses released
Saturation status: Unknown
Response matrix
Fit to width. Hover for subject & item; click a cell for details.

Correct (1)Incorrect (0)Unobserved
Scale: 1 = correct · 0 = incorrect
Subjects
- 1openai/gpt-5-nano-2025-08-071
- 2anthropic/claude-opus-4-20250514-thinking-10k1
- 3anthropic/claude-3-7-sonnet-202502191
- 4writer/palmyra-x51
- 5anthropic/claude-sonnet-4-5-202509291
- 6anthropic/claude-sonnet-4-20250514-thinking-10k1
- 7cohere/command-r-plus1
- 8openai/gpt-oss-120b1
- 9ibm/granite-4.0-micro-with-guardian1
- 10anthropic/claude-3-sonnet-202402291
- 11anthropic/claude-3-opus-202402291
- 12openai/gpt-5-mini-2025-08-071
- 13writer/palmyra-fin1
- 14anthropic/claude-3-haiku-202403071
- 15moonshotai/kimi-k2-instruct1
- 16openai/gpt-oss-20b1
- 17anthropic/claude-3-5-sonnet-202406201
- 18writer/palmyra-x-0041
- 19openai/gpt-4.1-2025-04-141
- 20openai/gpt-4.1-mini-2025-04-141
- 21openai/o4-mini-2025-04-161
- 22openai/gpt-4.5-preview-2025-02-271
- 23qwen/qwen3-235b-a22b-instruct-2507-fp81
- 24ibm/granite-4.0-h-small-with-guardian1
- 25openai/o3-2025-04-160.99
- 26openai/gpt-4.1-nano-2025-04-140.99
- 27openai/o3-mini-2025-01-310.99
- 28meta/llama-3-8b-chat0.99
- 29openai/o1-2024-12-170.99
- 30meta/llama-3-70b-chat0.99
- 31openai/gpt-5-2025-08-070.99
- 32openai/gpt-4-turbo-2024-04-090.99
- 33openai/gpt-5.1-2025-11-130.99
- 34qwen/qwen3-next-80b-a3b-thinking0.99
- 35meta/llama-4-maverick-17b-128e-instruct-fp80.98
- 36meta/llama-3.1-8b-instruct-turbo0.98