Software Engineering

ProgramBench

ProgramBench asks an agent to rebuild a working program from nothing but its compiled binary and the documentation shipped with it, scoring the result by how much of a hidden behavioural test suite the rebuilt program passes.

200items
12subjects
MITlicense
software_engineeringdomain
agents_and_tool_usedomain
textmodality
item-level responses released
Saturation status: No

Response matrix

Per-(model, task) outcomes for 12 models across 15 agent runs, on upstream's own Resolved / Almost-Resolved thresholds.

Loading session strips…

lowhighUnobserved

Scale: {0 = <95% of behavioural tests pass, 1 = >=95% (upstream "Almost Resolved"), 2 = 100% (upstream "Resolved")}