Reasoning
Correlated Errors in LLMs (HELM MMLU)
Correlated Errors in LLMs — HELM MMLU per-(model x question) responses: 71 LLMs answering ~14K MMLU multiple-choice questions, binary correctness.
13,868items
71subjects
MITlicense
knowledgedomain
reasoningdomain
textmodality
item-level responses released
Saturation status: Unknown
Response matrix
Fit to width. Hover for subject & item; click a cell for details.

Correct (1)Incorrect (0)Unobserved
Scale: 1 = correct · 0 = incorrect
Subjects
- 1anthropic/claude-3-5-sonnet-202410220.8764
- 2anthropic/claude-3-5-sonnet-202406200.872
- 3google/gemini-1.5-pro-0020.8641
- 4anthropic/claude-3-opus-202402290.8504
- 5openai/gpt-4o-2024-08-060.8495
- 6meta/llama-3.1-405b-instruct-turbo0.849
- 7openai/gpt-4o-2024-05-130.8482
- 8openai/gpt-4-06130.8401
- 9qwen/qwen2.5-72b-instruct-turbo0.8331
- 10google/gemini-1.5-pro-0010.8274
- 11qwen/qwen2-72b-instruct0.8274
- 12openai/gpt-4-turbo-2024-04-090.8194
- 13writer/palmyra-x-0040.8176
- 14meta/llama-3.2-90b-vision-instruct-turbo0.8109
- 15openai/gpt-4-1106-preview0.8101
- 16meta/llama-3.1-70b-instruct-turbo0.8091
- 17google/gemini-1.5-pro-preview-04090.8086
- 18mistralai/mistral-large-24070.8066
- 1901-ai/yi-large-preview0.8
- 20meta/llama-3-70b0.7867
- 21ai21/jamba-1.5-large0.7824
- 22google/text-unicorn@0010.7799
- 23writer/palmyra-x-v30.7799
- 24google/gemini-1.5-flash-0010.7755
- 25google/gemini-1.5-flash-preview-05140.7746
- 26qwen/qwen1.5-110b-chat0.7744
- 27qwen/qwen1.5-72b0.774
- 28microsoft/phi-3-medium-4k-instruct0.7738
- 29mistralai/mixtral-8x22b0.7734
- 30anthropic/claude-3-sonnet-202402290.762
- 3101-ai/yi-34b0.7616
- 32microsoft/phi-3-small-8k-instruct0.7613
- 33openai/gpt-4o-mini-2024-07-180.7577
- 34google/gemma-2-27b0.7481
- 35google/gemini-1.5-flash-0020.7402
- 36qwen/qwen1.5-32b0.7394