General
REAL
REAL: autonomous web agents on deterministic simulations of real websites. Tasks across high-fidelity website clones (Staynb, DashDish, ...), each with a deterministic success checker made of one or more evals. Response is the per-attempt binary task success (1 iff all of the task's evals pass).
233items
44subjects
unknownlicense
agents_and_tool_usedomain
textmodality
gui_screenshotmodality
item-level responses released
Saturation status: Unknown
Response matrix
Fit to width. Hover for subject & item; click a cell for details.

Correct (1)Incorrect (0)Unobserved
Scale: 1 = correct · 0 = incorrect
Subjects
- 1anthropic/claude-sonnet-4.50.7473
- 2anthropic-computer_use0.7273
- 3anthropic/claude-3.7-sonnet:thinking0.7179
- 4google/gemini-2.5-pro0.6782
- 5anthropic/claude-sonnet-40.6667
- 6google/gemini-2.5-flash0.6604
- 7x-ai/grok-4-fast0.6429
- 8amazon/nova-act-v1.00.6352
- 9openai/gpt-50.618
- 10GBOX0.5982
- 11KISS-10.5714
- 12brassbunny0.4902
- 13Anthropic computer use0.4783
- 14deepseek/deepseek-v3.2-exp0.4754
- 15AGI agent 00.4595
- 16meta-llama/llama-4-maverick0.4444
- 17Web_Agent_GPT-4o0.4348
- 18Claude 3.7 Sonnet:thinking0.4107
- 19Claude-Opus-4:Thinking0.4107
- 20Sonnet-4:thinking0.3929
- 21Gemini 2.5 pro0.3839
- 22gpt-4o-prsm0.3768
- 23Magellanes0.375
- 24Browser Use Claude 3.7 Sonnet:thinking0.3514
- 25o30.3482
- 26Claude 3.7 Sonnet0.3393
- 27openai/gpt-5-nano0.3382
- 28Browser Use GPT 4o0.3119
- 29GPT 4.10.2812
- 30Update Eval0.2667
- 31o3 mini0.25
- 32Stagehand Open Operator GPT 4o0.2018
- 33Deepseek chat v3 03240.1964
- 34o10.1607
- 35o1 mini0.1518
- 36GPT 4o0.1351