Software Engineering
KernelBench
LLMs generate optimized CUDA kernels for 250 PyTorch ML workloads; response is per-task functional correctness.
250items
7subjects
MITlicense
ml_engineeringdomain
software_engineeringdomain
textmodality
item-level responses released
Saturation status: Unknown
Response matrix
Fit to width. Hover for subject & item; click a cell for details.
This condition combination wasn’t evaluated — try a different attack, category, or judge.
Correct (1)Incorrect (0)Unobserved
Scale: 1 = correct · 0 = incorrect
Subjects
- 1OpenAI o10.5816
- 2DeepSeek-R10.2845
- 3Claude 3.5 Sonnet0.2805
- 4DeepSeek-V30.2439
- 5GPT-4o0.2227
- 6Llama 3.1 405B Instruct0.168
- 7Llama 3.1 70B Instruct0.076