Software Engineering

KernelBench

LLMs generate optimized CUDA kernels for 250 PyTorch ML workloads; response is per-task functional correctness.

250items
7subjects
MITlicense
ml_engineeringdomain
software_engineeringdomain
textmodality
item-level responses released
Saturation status: Unknown

Response matrix

Fit to width. Hover for subject & item; click a cell for details.

This condition combination wasn’t evaluated — try a different attack, category, or judge.
Correct (1)Incorrect (0)Unobserved

Scale: 1 = correct · 0 = incorrect

Subjects

  1. 1OpenAI o10.5816
  2. 2DeepSeek-R10.2845
  3. 3Claude 3.5 Sonnet0.2805
  4. 4DeepSeek-V30.2439
  5. 5GPT-4o0.2227
  6. 6Llama 3.1 405B Instruct0.168
  7. 7Llama 3.1 70B Instruct0.076