Software Engineering

ALE-Bench

ALE-Bench: long-horizon, score-based algorithm engineering on AtCoder Heuristic Contest optimization problems. 40 problems; 102 frontier LLMs evaluated across 5 self-refinement budgets. Per-(model, problem) judge verdict reduced to binary ACCEPTED vs. not.

40items
102subjects
CC-BY-4.0license
software_engineeringdomain
reasoningdomain
textmodality
item-level responses released
Saturation status: Unknown

Response matrix

Fit to width. Hover for subject & item; click a cell for details.

ALE-Bench response matrix: AI models (rows) against items (columns)
Correct (1)Incorrect (0)Unobserved

Scale: 1 = correct · 0 = incorrect

Subjects

  1. 1gpt-5.1-codex-high1
  2. 2gpt-5.1-thinking1
  3. 3gpt-5.1-codex-max-xhigh0.995
  4. 4gpt-5.2-high0.995
  5. 5gpt-5.2-medium0.99
  6. 6gpt-5-thinking0.99
  7. 7gpt-5.1-codex-max-high0.985
  8. 8gpt-5.4-mini-high0.98
  9. 9gpt-5.2-codex-xhigh0.98
  10. 10gpt-5.3-codex-xhigh0.98
  11. 11o4-mini-high0.98
  12. 12gpt-5.4-nano-high0.975
  13. 13gpt-5.5-xhigh0.975
  14. 14claude-4.8-opus-high0.97
  15. 15claude-4.5-opus0.97
  16. 16grok-4.20-beta0.965
  17. 17gpt-5.4-high0.96
  18. 18mercury-20.96
  19. 19claude-fable-5-high0.96
  20. 20claude-4.7-opus-no-thinking0.96
  21. 21gpt-5.5-medium0.955
  22. 22claude-4.8-opus-no-thinking0.955
  23. 23gpt-5.5-none0.95
  24. 24gpt-5.4-none0.95
  25. 25gpt-oss-120b0.945
  26. 26o3-high0.945
  27. 27gemini-3.1-pro-preview-low0.94
  28. 28glm-5.10.935
  29. 29gemini-3-flash-preview-high0.935
  30. 30claude-4.6-opus-no-thinking0.93
  31. 31claude-4.6-sonnet-medium0.93
  32. 32gpt-oss-20b0.92
  33. 33mistral-medium-3.50.92
  34. 34glm-5.2-max0.915
  35. 35deepseek-v3.1-terminus0.915
  36. 36mimo-v2.5-pro0.915