Software Engineering
ProgramBench
ProgramBench asks an agent to rebuild a working program from nothing but its compiled binary and the documentation shipped with it, scoring the result by how much of a hidden behavioural test suite the rebuilt program passes.
200items
12subjects
MITlicense
software_engineeringdomain
agents_and_tool_usedomain
textmodality
item-level responses released
Saturation status: No
Response matrix
Per-(model, task) outcomes for 12 models across 15 agent runs, on upstream's own Resolved / Almost-Resolved thresholds.
Loading session strips…
lowhighUnobserved
Scale: {0 = <95% of behavioural tests pass, 1 = >=95% (upstream "Almost Resolved"), 2 = 100% (upstream "Resolved")}