Safety
ReEval
Full binary response matrix (183 LMs x 22 HELM scenarios) released with the amortized IRT evaluation paper.
84,327items
183subjects
apache-2.0license
safetydomain
mathematicsdomain
lawdomain
medicinedomain
reasoningdomain
knowledgedomain
nlp_taskdomain
multilingualdomain
textmodality
item-level responses released
Saturation status: Unknown
Response matrix
Fit to width. Hover for subject & item; click a cell for details.

Correct (1)Incorrect (0)Unobserved
Scale: 1 = correct · 0 = incorrect
Subjects
- 1google/gemini-1.5-pro-001-safety-block-none0.8674
- 2anthropic/claude-3-5-sonnet-202410220.8663
- 3anthropic/claude-3-5-sonnet-202406200.8525
- 4anthropic/claude-3-opus-202402290.83
- 5openai/o1-2024-12-170.8275
- 6qwen/qwen2.5-72b-instruct-turbo0.8204
- 7google/gemini-1.5-flash-001-safety-block-none0.8194
- 8google/gemini-1.5-pro-0020.8091
- 9google/gemini-1.5-pro-preview-04090.8083
- 10amazon/nova-pro-v1:00.8059
- 11google/gemini-1.5-pro-0010.8045
- 12meta/llama-3.2-90b-vision-instruct-turbo0.7973
- 13openai/gpt-4o-2024-08-060.7953
- 14openai/gpt-4-06130.7941
- 15google/gemini-2.0-flash-exp0.7926
- 16openai/gpt-4-turbo-2024-04-090.7896
- 17meta/llama-3.3-70b-instruct-turbo0.7891
- 18openai/gpt-4-1106-preview0.7856
- 19meta/llama-3-70b0.7784
- 20meta/llama-3.1-405b-instruct-turbo0.7771
- 21openai/o3-mini-2025-01-310.7764
- 22google/gemini-1.5-flash-preview-05140.7746
- 23qwen/qwen2-72b-instruct0.7738
- 24writer/palmyra-x-v30.7677
- 25openai/gpt-4o-2024-05-130.7674
- 26upstage/solar-pro-2411260.7654
- 27ai21/jamba-1.5-large0.7639
- 28google/text-unicorn@0010.7599
- 29qwen/qwen1.5-72b0.7572
- 30amazon/nova-lite-v1:00.7571
- 31mistralai/mixtral-8x22b0.7561
- 32microsoft/phi-3-medium-4k-instruct0.7533
- 33deepseek-ai/deepseek-v30.7519
- 3401-ai/yi-large-preview0.7506
- 35google/gemini-1.5-flash-0010.7503
- 36google/gemma-2-27b0.7474