Mathematics

AlgoTune

AlgoTune: can LLM agents speed up general-purpose numerical programs? 154 coding tasks, each with a reference solver from a popular library, an input generator and a verifier. The agent must produce a correct but faster implementation. Binary response: pass iff the validated speedup over the reference is >= 1.0 (matched-or-beat the reference and was correct).

154items
18subjects
MITlicense
software_engineeringdomain
reasoningdomain
textmodality
item-level responses released
Saturation status: Unknown

Response matrix

Fit to width. Hover for subject & item; click a cell for details.

AlgoTune response matrix: AI models (rows) against items (columns)
Correct (1)Incorrect (0)Unobserved

Scale: 1 = correct · 0 = incorrect

Subjects

  1. 1gemini-3.1-pro-preview0.8701
  2. 2gpt-5.40.8312
  3. 3gpt-5.20.8312
  4. 4gpt-50.8117
  5. 5claude-opus-4.50.8117
  6. 6claude-sonnet-4-5-202509290.7792
  7. 7deepseek-reasoner0.7792
  8. 8o4-mini0.7662
  9. 9glm-4.50.7662
  10. 10gemini-3-pro-preview0.7403
  11. 11qwen3-coder0.6948
  12. 12claude-opus-4-1-202508050.6558
  13. 13gpt-5-mini0.6104
  14. 14gpt-oss-120b0.6039
  15. 15gemini-2.5-pro0.5714
  16. 16claude-opus-4-202505140.5584
  17. 17claude-opus-4.60.5455
  18. 18gpt-5-pro (medium)0.5065