Safety
MMLU Moral Disputes
MMLU moral_disputes per-(model, question) accuracy (acc in {0,1}) over an applied-ethics MCQ subject, from the Open LLM Leaderboard v1 details datasets. Model panel capped to 150.
346items
150subjects
MITlicense
knowledgedomain
reasoningdomain
textmodality
item-level responses released
Saturation status: Unknown
Response matrix
Fit to width. Hover for subject & item; click a cell for details.

Correct (1)Incorrect (0)Unobserved
Scale: 1 = correct · 0 = incorrect
Subjects
- 1augtoma__qCammel-70-x0.8035
- 2jordiclive__Llama-2-70b-oasst-1-2000.789
- 3quantumaikr__llama-2-70b-fb16-guanaco-1k0.7832
- 4deepnight-research__llama-2-70B-inst0.7803
- 5upstage__Llama-2-70b-instruct-v20.7803
- 6liuxiang886__llama2-70B-qlora-gpt40.7803
- 7upstage__Llama-2-70b-instruct0.7775
- 8jarradh__llama2_70b_chat_uncensored0.7717
- 9garage-bAInd__Platypus2-70B-instruct0.7572
- 10MayaPH__GodziLLa2-70B0.7543
- 11upstage__llama-65b-instruct0.7514
- 12garage-bAInd__Dolphin-Platypus2-70B0.7428
- 13HiTZ__alpaca-lora-65b-en-pt-es-ca0.737
- 14OpenBuddy__openbuddy-llama-65b-v8-bf160.7225
- 15lilloukas__Platypus-30B0.7225
- 16quantumaikr__QuantumLM-70B-hf0.7139
- 17lilloukas__GPlatty-30B0.7023
- 18upstage__llama-30b-instruct-20480.6965
- 19OptimalScale__robin-65b-v2-delta0.6908
- 20upstage__llama-30b-instruct0.6647
- 21garage-bAInd__Camel-Platypus2-13B0.6503
- 22shareAI__bimoGPT-llama2-13b0.6474
- 23OpenBuddyEA__openbuddy-llama-30b-v7.1-bf160.6474
- 24camel-ai__CAMEL-33B-Combined-Data0.6445
- 25layoric__llama-2-13b-code-alpaca0.6416
- 26Lajonbot__Llama-2-13b-hf-instruct-pl-lora_unload0.6387
- 27augtoma__qCammel-130.6387
- 28Aeala__GPT4-x-AlpacaDente2-30b0.6387
- 29Aeala__GPT4-x-AlpacaDente-30b0.6358
- 30CalderaAI__13B-Legerdemain-L20.6329
- 31shareAI__llama2-13b-Chinese-chat0.6329
- 32NousResearch__Redmond-Puffin-13B0.6301
- 33OpenBuddy__openbuddy-llama2-13b-v8.1-fp160.6069
- 34CalderaAI__30B-Lazarus0.6012
- 35WizardLM__WizardLM-13B-V1.20.5983
- 36NousResearch__Nous-Hermes-Llama2-13b0.5954