Mathematics
MoE-CAP
MoE-CAP: a benchmark and framework for sparse Mixture-of-Experts LLM serving systems, characterizing the Cost/Accuracy/Performance trade-off with sparsity-aware utilization metrics (S-MBU, S-MFU). This ingestion captures the released lm-eval-harness per-item correctness: each MoE model under a serving framework is graded per GSM8K / MMLU question.
7,736items
14subjects
unknownlicense
mathematicsdomain
knowledgedomain
textmodality
item-level responses released
Saturation status: Unknown
Response matrix
Fit to width. Hover for subject & item; click a cell for details.

Correct (1)Incorrect (0)Unobserved
Scale: 1 = correct · 0 = incorrect
Subjects
- 1Qwen/Qwen3-30B-A3B (sglang)0.93
- 2mistralai/Mixtral-8x22B-Instruct-v0.1 (vllm_moe)0.7707
- 3mistralai/Mixtral-8x22B-Instruct-v0.1 (vllm_moe_fixbs)0.7705
- 4mistralai/Mixtral-8x22B-Instruct-v0.1 (tensorrt_llm)0.7621
- 5databricks/dbrx-instruct (vllm_moe_fixbs)0.7252
- 6databricks/dbrx-instruct (vllm_moe)0.7245
- 7mistralai/Mixtral-8x7B-Instruct-v0.1 (vllm_moe_fixbs)0.706
- 8mistralai/Mixtral-8x7B-Instruct-v0.1 (vllm_moe)0.7053
- 9Qwen/Qwen1.5-MoE-A2.7B-Chat (vllm_moe)0.5959
- 10Qwen/Qwen1.5-MoE-A2.7B-Chat (hf-chat)0.5928
- 11databricks/dbrx-instruct (tensorrt_llm)0.3027
- 12mistralai/Mixtral-8x22B-Instruct-v0.1 (hf-chat)0.2965
- 13mistralai/Mixtral-8x7B-Instruct-v0.1 (hf-chat)0.2875
- 14databricks/dbrx-instruct (hf-chat)0.2799