General

LLM-Uncertainty-Bench

LLM-Uncertainty-Bench: 25 LLMs on five 6-way (A-F) multiple-choice datasets (MMLU, CosmosQA, HellaSwag, Halu-Dialogue, Halu-Summarization). Per-item option logits are released; we grade argmax vs the gold letter.

49,931items
25subjects
MITlicense
knowledgedomain
reasoningdomain
safetydomain
textmodality
item-level responses released
Saturation status: Unknown

Response matrix

Fit to width. Hover for subject & item; click a cell for details.

LLM-Uncertainty-Bench response matrix: AI models (rows) against items (columns)
Correct (1)Incorrect (0)Unobserved

Scale: 1 = correct · 0 = incorrect

Subjects

  1. 1Yi-34B0.8149
  2. 2Qwen-72B0.7832
  3. 3deepseek-llm-67b-chat0.7775
  4. 4Qwen-14B0.741
  5. 5Yi-34B-Chat0.73
  6. 6Llama-2-70b-hf0.724
  7. 7deepseek-llm-67b-base0.7171
  8. 8Yi-6B0.6889
  9. 9Mistral-7B-v0.10.6433
  10. 10Llama-2-13b-hf0.6027
  11. 11Qwen-7B0.6
  12. 12Yi-6B-Chat0.5976
  13. 13deepseek-llm-7b-chat0.575
  14. 14Llama-2-70b-chat-hf0.5136
  15. 15internlm-7b0.4917
  16. 16Llama-2-7b-hf0.4671
  17. 17deepseek-llm-7b-base0.4587
  18. 18Llama-2-13b-chat-hf0.4536
  19. 19Qwen-1_8B0.4228
  20. 20Llama-2-7b-chat-hf0.3811
  21. 21falcon-40b0.3478
  22. 22falcon-40b-instruct0.3381
  23. 23mpt-7b0.2734
  24. 24falcon-7b0.2487
  25. 25falcon-7b-instruct0.2445