Software Engineering

SWE-rebench

SWE-rebench: an automated, decontaminated benchmark of real-world software-engineering agent tasks. Each task is a GitHub issue + repo snapshot; an LLM agent must produce a patch verified by the repo test suite (FAIL_TO_PASS / PASS_TO_PASS). This build ingests the publicly released per-instance trajectories of the OpenHands v0.54.0 + Qwen3-Coder-480B-A35B-Instruct agent, emitting a binary resolved=1/0 response per (agent, instance, run).

6,271items
1subjects
CC-BY-4.0license
software_engineeringdomain
agents_and_tool_usedomain
textmodality
item-level responses released
Saturation status: Unknown

Response matrix

Fit to width. Hover for subject & item; click a cell for details.

SWE-rebench response matrix: AI models (rows) against items (columns)
Correct (1)Incorrect (0)Unobserved

Scale: 1 = correct · 0 = incorrect

Subjects

  1. 1OpenHands v0.54.0 + Qwen3-Coder-480B-A35B-Instruct0.4795