Software Engineering
DeepSWE
DeepSWE tests whether coding agents can complete long horizon software engineering tasks written from scratch against active open source repositories, graded by hand written functional verifiers the agent never sees rather than by tests shipped with an existing fix.
113items
68subjects
Apache-2.0license
software_engineeringdomain
agents_and_tool_usedomain
textmodality
item-level responses released
Saturation status: No
Response matrix
Rasch analysis p = σ(θ − z + c)
36,289 responses, 80/20 split over cells · 68 subjects · 113 items · 2 conditions
AUC train
0.846
AUC test
0.845
Loading session strips…
Correct (1)Incorrect (0)Unobserved
Scale: 1 = correct · 0 = incorrect