Safety

TruthfulQA-MC

TruthfulQA-MC1 per-(model, question) correctness (mc1 bool, 1=truthful) over 817 questions, from the Open LLM Leaderboard v1 details datasets. Model panel capped to 150.

817items
150subjects
Apache-2.0license
safetydomain
knowledgedomain
textmodality
item-level responses released
Saturation status: Unknown

Response matrix

Fit to width. Hover for subject & item; click a cell for details.

TruthfulQA-MC response matrix: AI models (rows) against items (columns)
Correct (1)Incorrect (0)Unobserved

Scale: 1 = correct · 0 = incorrect

Subjects

  1. 1TigerResearch__tigerbot-7b-sft0.47
  2. 2deepnight-research__llama-2-70B-inst0.4443
  3. 3upstage__Llama-2-70b-instruct-v20.4443
  4. 4upstage__llama-65b-instruct0.4296
  5. 5MayaPH__GodziLLa2-70B0.4259
  6. 6upstage__Llama-2-70b-instruct0.4235
  7. 7CalderaAI__30B-Lazarus0.4137
  8. 8quantumaikr__llama-2-70b-fb16-guanaco-1k0.4064
  9. 9augtoma__qCammel-70-x0.4015
  10. 10upstage__llama-30b-instruct-20480.3978
  11. 11OpenBuddy__openbuddy-llama-65b-v8-bf160.388
  12. 12WizardLM__WizardLM-70B-V1.00.3868
  13. 13WizardLM__WizardLM-13B-V1.10.3843
  14. 14MayaPH__GodziLLa-30B0.3782
  15. 15jarradh__llama2_70b_chat_uncensored0.3709
  16. 16Aeala__GPT4-x-AlpacaDente-30b0.366
  17. 17Aeala__GPT4-x-Alpasta-13b0.3647
  18. 18OpenBuddyEA__openbuddy-llama-30b-v7.1-bf160.3635
  19. 19liuxiang886__llama2-70B-qlora-gpt40.3623
  20. 20kevinpro__Vicuna-13B-CoT0.3623
  21. 21jordiclive__Llama-2-70b-oasst-1-2000.3599
  22. 22quantumaikr__QuantumLM-70B-hf0.3586
  23. 23LLMs__WizardLM-13B-V1.00.3562
  24. 24edor__Stable-Platypus2-mini-7B0.3562
  25. 25lilloukas__GPlatty-30B0.355
  26. 26NousResearch__Nous-Hermes-13b0.3537
  27. 27upstage__llama-30b-instruct0.3525
  28. 28NousResearch__Nous-Hermes-Llama2-13b0.3501
  29. 29OpenBuddy__openbuddy-llama2-13b-v8.1-fp160.3501
  30. 30Lajonbot__vicuna-13b-v1.3-PL-lora_unload0.3476
  31. 31Lajonbot__tableBeluga-7B-instruct-pl-lora_unload0.3464
  32. 32mosaicml__mpt-30b-chat0.339
  33. 33OptimalScale__robin-13b-v2-delta0.3378
  34. 34HiTZ__alpaca-lora-65b-en-pt-es-ca0.3366
  35. 35NousResearch__Nous-Hermes-llama-2-7b0.3341
  36. 36camel-ai__CAMEL-13B-Combined-Data0.3341