Safety
TruthfulQA-MC
TruthfulQA-MC1 per-(model, question) correctness (mc1 bool, 1=truthful) over 817 questions, from the Open LLM Leaderboard v1 details datasets. Model panel capped to 150.
817items
150subjects
Apache-2.0license
safetydomain
knowledgedomain
textmodality
item-level responses released
Saturation status: Unknown
Response matrix
Fit to width. Hover for subject & item; click a cell for details.

Correct (1)Incorrect (0)Unobserved
Scale: 1 = correct · 0 = incorrect
Subjects
- 1TigerResearch__tigerbot-7b-sft0.47
- 2deepnight-research__llama-2-70B-inst0.4443
- 3upstage__Llama-2-70b-instruct-v20.4443
- 4upstage__llama-65b-instruct0.4296
- 5MayaPH__GodziLLa2-70B0.4259
- 6upstage__Llama-2-70b-instruct0.4235
- 7CalderaAI__30B-Lazarus0.4137
- 8quantumaikr__llama-2-70b-fb16-guanaco-1k0.4064
- 9augtoma__qCammel-70-x0.4015
- 10upstage__llama-30b-instruct-20480.3978
- 11OpenBuddy__openbuddy-llama-65b-v8-bf160.388
- 12WizardLM__WizardLM-70B-V1.00.3868
- 13WizardLM__WizardLM-13B-V1.10.3843
- 14MayaPH__GodziLLa-30B0.3782
- 15jarradh__llama2_70b_chat_uncensored0.3709
- 16Aeala__GPT4-x-AlpacaDente-30b0.366
- 17Aeala__GPT4-x-Alpasta-13b0.3647
- 18OpenBuddyEA__openbuddy-llama-30b-v7.1-bf160.3635
- 19liuxiang886__llama2-70B-qlora-gpt40.3623
- 20kevinpro__Vicuna-13B-CoT0.3623
- 21jordiclive__Llama-2-70b-oasst-1-2000.3599
- 22quantumaikr__QuantumLM-70B-hf0.3586
- 23LLMs__WizardLM-13B-V1.00.3562
- 24edor__Stable-Platypus2-mini-7B0.3562
- 25lilloukas__GPlatty-30B0.355
- 26NousResearch__Nous-Hermes-13b0.3537
- 27upstage__llama-30b-instruct0.3525
- 28NousResearch__Nous-Hermes-Llama2-13b0.3501
- 29OpenBuddy__openbuddy-llama2-13b-v8.1-fp160.3501
- 30Lajonbot__vicuna-13b-v1.3-PL-lora_unload0.3476
- 31Lajonbot__tableBeluga-7B-instruct-pl-lora_unload0.3464
- 32mosaicml__mpt-30b-chat0.339
- 33OptimalScale__robin-13b-v2-delta0.3378
- 34HiTZ__alpaca-lora-65b-en-pt-es-ca0.3366
- 35NousResearch__Nous-Hermes-llama-2-7b0.3341
- 36camel-ai__CAMEL-13B-Combined-Data0.3341