Safety

HELM SimpleSafetyTests

HELM SimpleSafetyTests: per-(model, prompt) safety_score in {0,1} (1=safe) from HELM's safety classifier. 87 models x 100 severe-harm prompts (self-harm, physical harm, illegal items, scams, child abuse).

100items
81subjects
Apache-2.0license
safetydomain
textmodality
item-level responses released
Saturation status: Unknown

Response matrix

Fit to width. Hover for subject & item; click a cell for details.

HELM SimpleSafetyTests response matrix: AI models (rows) against items (columns)
Correct (1)Incorrect (0)Unobserved

Scale: 1 = correct · 0 = incorrect

Subjects

  1. 1openai/gpt-5-nano-2025-08-071
  2. 2anthropic/claude-opus-4-20250514-thinking-10k1
  3. 3anthropic/claude-3-7-sonnet-202502191
  4. 4writer/palmyra-x51
  5. 5anthropic/claude-sonnet-4-5-202509291
  6. 6anthropic/claude-sonnet-4-20250514-thinking-10k1
  7. 7cohere/command-r-plus1
  8. 8openai/gpt-oss-120b1
  9. 9ibm/granite-4.0-micro-with-guardian1
  10. 10anthropic/claude-3-sonnet-202402291
  11. 11anthropic/claude-3-opus-202402291
  12. 12openai/gpt-5-mini-2025-08-071
  13. 13writer/palmyra-fin1
  14. 14anthropic/claude-3-haiku-202403071
  15. 15moonshotai/kimi-k2-instruct1
  16. 16openai/gpt-oss-20b1
  17. 17anthropic/claude-3-5-sonnet-202406201
  18. 18writer/palmyra-x-0041
  19. 19openai/gpt-4.1-2025-04-141
  20. 20openai/gpt-4.1-mini-2025-04-141
  21. 21openai/o4-mini-2025-04-161
  22. 22openai/gpt-4.5-preview-2025-02-271
  23. 23qwen/qwen3-235b-a22b-instruct-2507-fp81
  24. 24ibm/granite-4.0-h-small-with-guardian1
  25. 25openai/o3-2025-04-160.99
  26. 26openai/gpt-4.1-nano-2025-04-140.99
  27. 27openai/o3-mini-2025-01-310.99
  28. 28meta/llama-3-8b-chat0.99
  29. 29openai/o1-2024-12-170.99
  30. 30meta/llama-3-70b-chat0.99
  31. 31openai/gpt-5-2025-08-070.99
  32. 32openai/gpt-4-turbo-2024-04-090.99
  33. 33openai/gpt-5.1-2025-11-130.99
  34. 34qwen/qwen3-next-80b-a3b-thinking0.99
  35. 35meta/llama-4-maverick-17b-128e-instruct-fp80.98
  36. 36meta/llama-3.1-8b-instruct-turbo0.98