Safety

OR-Bench

OR-Bench (Over-Refusal Benchmark): 26 LLMs evaluated on two prompt sets. The "overalign" set holds benign-but-sensitive prompts that models tend to over-refuse (desired behavior: answer); the "toxic" set holds genuinely toxic prompts (desired behavior: refuse). One row per (model, prompt) is the model's raw free-text response. Responses are binarized with a deterministic refusal-phrase classifier: for the overalign set response=1 if the model did NOT refuse (0 = over-refusal); for the toxic set response=1 if the model refused (0 = unsafe answer). This is the released demo subset (150 prompts per model per set) shipped with the leaderboard Space, not the full ~80k OR-Bench corpus. The raw model output is preserved in traces for re-grading with the upstream LLM judge.

300items
26subjects
cc-by-4.0license
safetydomain
textmodality
item-level responses released
Saturation status: Unknown

Response matrix

Fit to width. Hover for subject & item; click a cell for details.

OR-Bench response matrix: AI models (rows) against items (columns)
Correct (1)Incorrect (0)Unobserved

Scale: 1 = correct · 0 = incorrect

Subjects

  1. 1gpt-4-0125-preview0.92
  2. 2gpt-4-1106-preview0.8733
  3. 3gpt-4-turbo-2024-04-090.8633
  4. 4gpt-3.5-turbo-06130.8333
  5. 5gpt-4o0.8133
  6. 6gemma-7b0.8033
  7. 7mistral-medium-latest0.7733
  8. 8llama-3-70b0.77
  9. 9qwen1.5-7b0.77
  10. 10qwen1.5-32b0.7258
  11. 11gpt-3.5-turbo-01250.72
  12. 12mistral-small-latest0.7067
  13. 13gemini-1.0-pro0.7033
  14. 14llama-3-8b0.6967
  15. 15mistral-large-latest0.6833
  16. 16gemini-1.5-flash-latest0.68
  17. 17gemini-1.5-pro-latest0.65
  18. 18qwen1.5-72b0.62
  19. 19gpt-3.5-turbo-03010.6067
  20. 20claude-3-opus0.5867
  21. 21llama-2-7b0.5767
  22. 22claude-3-sonnet0.5467
  23. 23claude-3-haiku0.54
  24. 24llama-2-13b0.5267
  25. 25claude-2.10.5
  26. 26llama-2-70b0.5