General
LLM-Uncertainty-Bench
LLM-Uncertainty-Bench: 25 LLMs on five 6-way (A-F) multiple-choice datasets (MMLU, CosmosQA, HellaSwag, Halu-Dialogue, Halu-Summarization). Per-item option logits are released; we grade argmax vs the gold letter.
49,931items
25subjects
MITlicense
knowledgedomain
reasoningdomain
safetydomain
textmodality
item-level responses released
Saturation status: Unknown
Response matrix
Fit to width. Hover for subject & item; click a cell for details.

Correct (1)Incorrect (0)Unobserved
Scale: 1 = correct · 0 = incorrect
Subjects
- 1Yi-34B0.8149
- 2Qwen-72B0.7832
- 3deepseek-llm-67b-chat0.7775
- 4Qwen-14B0.741
- 5Yi-34B-Chat0.73
- 6Llama-2-70b-hf0.724
- 7deepseek-llm-67b-base0.7171
- 8Yi-6B0.6889
- 9Mistral-7B-v0.10.6433
- 10Llama-2-13b-hf0.6027
- 11Qwen-7B0.6
- 12Yi-6B-Chat0.5976
- 13deepseek-llm-7b-chat0.575
- 14Llama-2-70b-chat-hf0.5136
- 15internlm-7b0.4917
- 16Llama-2-7b-hf0.4671
- 17deepseek-llm-7b-base0.4587
- 18Llama-2-13b-chat-hf0.4536
- 19Qwen-1_8B0.4228
- 20Llama-2-7b-chat-hf0.3811
- 21falcon-40b0.3478
- 22falcon-40b-instruct0.3381
- 23mpt-7b0.2734
- 24falcon-7b0.2487
- 25falcon-7b-instruct0.2445