General

REAL

REAL: autonomous web agents on deterministic simulations of real websites. Tasks across high-fidelity website clones (Staynb, DashDish, ...), each with a deterministic success checker made of one or more evals. Response is the per-attempt binary task success (1 iff all of the task's evals pass).

233items
44subjects
unknownlicense
agents_and_tool_usedomain
textmodality
gui_screenshotmodality
item-level responses released
Saturation status: Unknown

Response matrix

Fit to width. Hover for subject & item; click a cell for details.

REAL response matrix: AI models (rows) against items (columns)
Correct (1)Incorrect (0)Unobserved

Scale: 1 = correct · 0 = incorrect

Subjects

  1. 1anthropic/claude-sonnet-4.50.7473
  2. 2anthropic-computer_use0.7273
  3. 3anthropic/claude-3.7-sonnet:thinking0.7179
  4. 4google/gemini-2.5-pro0.6782
  5. 5anthropic/claude-sonnet-40.6667
  6. 6google/gemini-2.5-flash0.6604
  7. 7x-ai/grok-4-fast0.6429
  8. 8amazon/nova-act-v1.00.6352
  9. 9openai/gpt-50.618
  10. 10GBOX0.5982
  11. 11KISS-10.5714
  12. 12brassbunny0.4902
  13. 13Anthropic computer use0.4783
  14. 14deepseek/deepseek-v3.2-exp0.4754
  15. 15AGI agent 00.4595
  16. 16meta-llama/llama-4-maverick0.4444
  17. 17Web_Agent_GPT-4o0.4348
  18. 18Claude 3.7 Sonnet:thinking0.4107
  19. 19Claude-Opus-4:Thinking0.4107
  20. 20Sonnet-4:thinking0.3929
  21. 21Gemini 2.5 pro0.3839
  22. 22gpt-4o-prsm0.3768
  23. 23Magellanes0.375
  24. 24Browser Use Claude 3.7 Sonnet:thinking0.3514
  25. 25o30.3482
  26. 26Claude 3.7 Sonnet0.3393
  27. 27openai/gpt-5-nano0.3382
  28. 28Browser Use GPT 4o0.3119
  29. 29GPT 4.10.2812
  30. 30Update Eval0.2667
  31. 31o3 mini0.25
  32. 32Stagehand Open Operator GPT 4o0.2018
  33. 33Deepseek chat v3 03240.1964
  34. 34o10.1607
  35. 35o1 mini0.1518
  36. 36GPT 4o0.1351