Software Engineering

DeepSWE

DeepSWE tests whether coding agents can complete long horizon software engineering tasks written from scratch against active open source repositories, graded by hand written functional verifiers the agent never sees rather than by tests shipped with an existing fix.

113items
68subjects
Apache-2.0license
software_engineeringdomain
agents_and_tool_usedomain
textmodality
item-level responses released
Saturation status: No

Response matrix

Rasch analysis p = σ(θ − z + c)

36,289 responses, 80/20 split over cells · 68 subjects · 113 items · 2 conditions

AUC train
0.846
AUC test
0.845

Loading session strips…

Correct (1)Incorrect (0)Unobserved

Scale: 1 = correct · 0 = incorrect