Safety

HELM BBQ

HELM BBQ: per-(model, question) exact_match in {0,1} (1=correct) on 1000 BBQ social-bias multiple-choice questions. 87 models.

999items
81subjects
Apache-2.0license
safetydomain
textmodality
item-level responses released
Saturation status: Unknown

Response matrix

Fit to width. Hover for subject & item; click a cell for details.

HELM BBQ response matrix: AI models (rows) against items (columns)
Correct (1)Incorrect (0)Unobserved

Scale: 1 = correct · 0 = incorrect

Subjects

  1. 1anthropic/claude-opus-4-202505140.993
  2. 2anthropic/claude-sonnet-4-5-202509290.989
  3. 3openai/gpt-oss-120b0.985
  4. 4google/gemini-3-pro-preview0.984
  5. 5anthropic/claude-sonnet-4-202505140.979
  6. 6openai/o3-2025-04-160.979
  7. 7zai-org/glm-4.5-air-fp80.978
  8. 8openai/gpt-5-nano-2025-08-070.976
  9. 9google/gemini-2.5-flash-preview-04-170.975
  10. 10openai/o1-2024-12-170.973
  11. 11anthropic/claude-opus-4-20250514-thinking-10k0.971
  12. 12openai/o1-mini-2024-09-120.969
  13. 13openai/gpt-5-2025-08-070.968
  14. 14openai/gpt-oss-20b0.967
  15. 15xai/grok-3-mini-beta0.967
  16. 16deepseek-ai/deepseek-v30.967
  17. 17deepseek-ai/deepseek-r10.966
  18. 18deepseek-ai/deepseek-r1-hide-reasoning0.966
  19. 19qwen/qwen3-235b-a22b-fp8-tput0.966
  20. 20anthropic/claude-sonnet-4-20250514-thinking-10k0.966
  21. 21google/gemini-2.5-pro-preview-03-250.964
  22. 22deepseek-ai/deepseek-r1-05280.963
  23. 23openai/gpt-5-mini-2025-08-070.963
  24. 24qwen/qwen3-235b-a22b-instruct-2507-fp80.962
  25. 25writer/palmyra-x-0040.955
  26. 26meta/llama-3.1-70b-instruct-turbo0.954
  27. 27qwen/qwen2-72b-instruct0.951
  28. 28openai/gpt-4o-2024-05-130.951
  29. 29google/gemini-2.5-flash-lite0.949
  30. 30moonshotai/kimi-k2-instruct0.949
  31. 31anthropic/claude-3-5-sonnet-202406200.949
  32. 32writer/palmyra-x50.948
  33. 33google/gemini-1.5-flash-0010.947
  34. 34meta/llama-3.1-405b-instruct-turbo0.945
  35. 35google/gemini-1.5-pro-0010.945
  36. 36writer/palmyra-fin0.942