Agents & Tool Use

ARC-AGI-3

ARC-AGI-3 drops an agent into novel grid worlds with no instructions, so it must infer the mechanics, the goal and the win condition purely by acting, and records whether it completes each level.

117items
5subjects
MITlicense
reasoningdomain
agents_and_tool_usedomain
gridmodality
item-level responses released
Saturation status: No

Response matrix

ARC Prize publishes per-environment results for the 25 public demo environments only; the 55 semi-private and 55 fully private environments are reported as aggregate scores alone, and ARC Prize states that public-set scores are not a valid measure of progress because systems can be tuned against them. Within the public set an item is one (environment, level) pair. A level's opening frame becomes public only once a released run reaches it: 117 of the 183 levels have one and are curated here, and no evaluated model reached the other 66, so ARC Prize released no content for them. Columns are grouped by level, deepest first; the number beside each level heading is how many environments have a curated level that deep.

Rasch analysis p = σ(θ − z + c)

2,220 responses, 80/20 split over cells · 5 subjects · 117 items · 40 conditions

AUC train
0.992
AUC test
0.965

Loading session strips…

Correct (1)Incorrect (0)Unobserved

Scale: 1 = correct · 0 = incorrect