Safety

MMLU Moral Disputes

MMLU moral_disputes per-(model, question) accuracy (acc in {0,1}) over an applied-ethics MCQ subject, from the Open LLM Leaderboard v1 details datasets. Model panel capped to 150.

346items
150subjects
MITlicense
knowledgedomain
reasoningdomain
textmodality
item-level responses released
Saturation status: Unknown

Response matrix

Fit to width. Hover for subject & item; click a cell for details.

MMLU Moral Disputes response matrix: AI models (rows) against items (columns)
Correct (1)Incorrect (0)Unobserved

Scale: 1 = correct · 0 = incorrect

Subjects

  1. 1augtoma__qCammel-70-x0.8035
  2. 2jordiclive__Llama-2-70b-oasst-1-2000.789
  3. 3quantumaikr__llama-2-70b-fb16-guanaco-1k0.7832
  4. 4deepnight-research__llama-2-70B-inst0.7803
  5. 5upstage__Llama-2-70b-instruct-v20.7803
  6. 6liuxiang886__llama2-70B-qlora-gpt40.7803
  7. 7upstage__Llama-2-70b-instruct0.7775
  8. 8jarradh__llama2_70b_chat_uncensored0.7717
  9. 9garage-bAInd__Platypus2-70B-instruct0.7572
  10. 10MayaPH__GodziLLa2-70B0.7543
  11. 11upstage__llama-65b-instruct0.7514
  12. 12garage-bAInd__Dolphin-Platypus2-70B0.7428
  13. 13HiTZ__alpaca-lora-65b-en-pt-es-ca0.737
  14. 14OpenBuddy__openbuddy-llama-65b-v8-bf160.7225
  15. 15lilloukas__Platypus-30B0.7225
  16. 16quantumaikr__QuantumLM-70B-hf0.7139
  17. 17lilloukas__GPlatty-30B0.7023
  18. 18upstage__llama-30b-instruct-20480.6965
  19. 19OptimalScale__robin-65b-v2-delta0.6908
  20. 20upstage__llama-30b-instruct0.6647
  21. 21garage-bAInd__Camel-Platypus2-13B0.6503
  22. 22shareAI__bimoGPT-llama2-13b0.6474
  23. 23OpenBuddyEA__openbuddy-llama-30b-v7.1-bf160.6474
  24. 24camel-ai__CAMEL-33B-Combined-Data0.6445
  25. 25layoric__llama-2-13b-code-alpaca0.6416
  26. 26Lajonbot__Llama-2-13b-hf-instruct-pl-lora_unload0.6387
  27. 27augtoma__qCammel-130.6387
  28. 28Aeala__GPT4-x-AlpacaDente2-30b0.6387
  29. 29Aeala__GPT4-x-AlpacaDente-30b0.6358
  30. 30CalderaAI__13B-Legerdemain-L20.6329
  31. 31shareAI__llama2-13b-Chinese-chat0.6329
  32. 32NousResearch__Redmond-Puffin-13B0.6301
  33. 33OpenBuddy__openbuddy-llama2-13b-v8.1-fp160.6069
  34. 34CalderaAI__30B-Lazarus0.6012
  35. 35WizardLM__WizardLM-13B-V1.20.5983
  36. 36NousResearch__Nous-Hermes-Llama2-13b0.5954