Software Engineering

GSO

GSO: challenging software-optimization tasks for evaluating SWE-agents. Each task asks an agent to reproduce a real performance optimization in a repository, graded by the Opt@K metric (the agent patch must meet the optimization target under a performance script plus correctness tests). We store the real per-task descriptions and per-(model, instance) Opt@1 pass/fail outcomes from the official gso-experiments reports.

102items
21subjects
MITlicense
software_engineeringdomain
agents_and_tool_usedomain
textmodality
item-level responses released
Saturation status: Unknown

Response matrix

Fit to width. Hover for subject & item; click a cell for details.

GSO response matrix: AI models (rows) against items (columns)
Correct (1)Incorrect (0)Unobserved

Scale: 1 = correct · 0 = incorrect

Subjects

  1. 1claude-opus-4.70.4412
  2. 2claude-opus-4.6-high0.4242
  3. 3gpt-5.5-xhigh0.402
  4. 4gpt-5.4-xhigh0.3478
  5. 5claude-opus-4.60.3333
  6. 6gpt-5.20.3218
  7. 7claude-opus-4.50.2784
  8. 8gpt-5.4-high0.2574
  9. 9gemini-3.1-pro0.2447
  10. 10gemini-3-pro0.1863
  11. 11claude-sonnet-4.50.1515
  12. 12gpt-5.10.1505
  13. 13o30.125
  14. 14gemini-3-flash0.1031
  15. 15claude-opus-40.0714
  16. 16gpt-50.07
  17. 17qwen3-coder0.05
  18. 18kimi-k20.0495
  19. 19claude-sonnet-40.0417
  20. 20gemini-2.5-pro0.0404
  21. 21glm-4.50.0323