Software Engineering
ALE-Bench
ALE-Bench: long-horizon, score-based algorithm engineering on AtCoder Heuristic Contest optimization problems. 40 problems; 102 frontier LLMs evaluated across 5 self-refinement budgets. Per-(model, problem) judge verdict reduced to binary ACCEPTED vs. not.
40items
102subjects
CC-BY-4.0license
software_engineeringdomain
reasoningdomain
textmodality
item-level responses released
Saturation status: Unknown
Response matrix
Fit to width. Hover for subject & item; click a cell for details.

Correct (1)Incorrect (0)Unobserved
Scale: 1 = correct · 0 = incorrect
Subjects
- 1gpt-5.1-codex-high1
- 2gpt-5.1-thinking1
- 3gpt-5.1-codex-max-xhigh0.995
- 4gpt-5.2-high0.995
- 5gpt-5.2-medium0.99
- 6gpt-5-thinking0.99
- 7gpt-5.1-codex-max-high0.985
- 8gpt-5.4-mini-high0.98
- 9gpt-5.2-codex-xhigh0.98
- 10gpt-5.3-codex-xhigh0.98
- 11o4-mini-high0.98
- 12gpt-5.4-nano-high0.975
- 13gpt-5.5-xhigh0.975
- 14claude-4.8-opus-high0.97
- 15claude-4.5-opus0.97
- 16grok-4.20-beta0.965
- 17gpt-5.4-high0.96
- 18mercury-20.96
- 19claude-fable-5-high0.96
- 20claude-4.7-opus-no-thinking0.96
- 21gpt-5.5-medium0.955
- 22claude-4.8-opus-no-thinking0.955
- 23gpt-5.5-none0.95
- 24gpt-5.4-none0.95
- 25gpt-oss-120b0.945
- 26o3-high0.945
- 27gemini-3.1-pro-preview-low0.94
- 28glm-5.10.935
- 29gemini-3-flash-preview-high0.935
- 30claude-4.6-opus-no-thinking0.93
- 31claude-4.6-sonnet-medium0.93
- 32gpt-oss-20b0.92
- 33mistral-medium-3.50.92
- 34glm-5.2-max0.915
- 35deepseek-v3.1-terminus0.915
- 36mimo-v2.5-pro0.915