Software Engineering
GSO
GSO: challenging software-optimization tasks for evaluating SWE-agents. Each task asks an agent to reproduce a real performance optimization in a repository, graded by the Opt@K metric (the agent patch must meet the optimization target under a performance script plus correctness tests). We store the real per-task descriptions and per-(model, instance) Opt@1 pass/fail outcomes from the official gso-experiments reports.
102items
21subjects
MITlicense
software_engineeringdomain
agents_and_tool_usedomain
textmodality
item-level responses released
Saturation status: Unknown
Response matrix
Fit to width. Hover for subject & item; click a cell for details.

Correct (1)Incorrect (0)Unobserved
Scale: 1 = correct · 0 = incorrect
Subjects
- 1claude-opus-4.70.4412
- 2claude-opus-4.6-high0.4242
- 3gpt-5.5-xhigh0.402
- 4gpt-5.4-xhigh0.3478
- 5claude-opus-4.60.3333
- 6gpt-5.20.3218
- 7claude-opus-4.50.2784
- 8gpt-5.4-high0.2574
- 9gemini-3.1-pro0.2447
- 10gemini-3-pro0.1863
- 11claude-sonnet-4.50.1515
- 12gpt-5.10.1505
- 13o30.125
- 14gemini-3-flash0.1031
- 15claude-opus-40.0714
- 16gpt-50.07
- 17qwen3-coder0.05
- 18kimi-k20.0495
- 19claude-sonnet-40.0417
- 20gemini-2.5-pro0.0404
- 21glm-4.50.0323