General

RealTime QA

RealTime QA: a dynamic multiple-choice news QA benchmark in which new questions are released weekly from CNN / THE WEEK news quizzes. We ingest the published GRANULAR per-(model, question) baseline predictions from the official realtimeqa_public repository: each response is one model's per-item correctness {0,1} on one weekly multiple-choice question (does the chosen 0-based choice index match the gold answer). Covers all weeks with released baselines (2022, 2023, 2026), both the standard (qa) and none-of-the-above (nota) MC settings, with retrieval method (closed-book / DPR / GCS) recorded as an evaluation condition. Free-text generation (_gen) outputs are excluded because they lack a gold choice index.

4,743items
16subjects
MITlicense
knowledgedomain
textmodality
item-level responses released
Saturation status: Unknown

Response matrix

Fit to width. Hover for subject & item; click a cell for details.

RealTime QA response matrix: AI models (rows) against items (columns)
Correct (1)Incorrect (0)Unobserved

Scale: 1 = correct · 0 = incorrect

Subjects

  1. 1openai/gpt-5.20.7395
  2. 2google/gemini-3-pro-preview0.7026
  3. 3openai/gpt-4.10.6342
  4. 4anthropic/claude-sonnet-4.50.6237
  5. 5anthropic/claude-opus-4.60.6071
  6. 6anthropic/claude-sonnet-4.60.5893
  7. 7openai/gpt-5.3-chat0.585
  8. 8meta-llama/llama-4-maverick0.5746
  9. 9openai/gpt-5.40.5675
  10. 10anthropic/claude-haiku-4.50.55
  11. 11meta-llama/llama-4-scout0.5484
  12. 12google/gemini-2.5-pro0.5295
  13. 13google/gemini-3.1-pro-preview0.5155
  14. 14GPT-3 (text-davinci)0.4647
  15. 15T5 (closed-book QA)0.3604
  16. 16RAG (DPR reader)0.3318