General
RealTime QA
RealTime QA: a dynamic multiple-choice news QA benchmark in which new questions are released weekly from CNN / THE WEEK news quizzes. We ingest the published GRANULAR per-(model, question) baseline predictions from the official realtimeqa_public repository: each response is one model's per-item correctness {0,1} on one weekly multiple-choice question (does the chosen 0-based choice index match the gold answer). Covers all weeks with released baselines (2022, 2023, 2026), both the standard (qa) and none-of-the-above (nota) MC settings, with retrieval method (closed-book / DPR / GCS) recorded as an evaluation condition. Free-text generation (_gen) outputs are excluded because they lack a gold choice index.
Response matrix
Fit to width. Hover for subject & item; click a cell for details.

Scale: 1 = correct · 0 = incorrect
Subjects
- 1openai/gpt-5.20.7395
- 2google/gemini-3-pro-preview0.7026
- 3openai/gpt-4.10.6342
- 4anthropic/claude-sonnet-4.50.6237
- 5anthropic/claude-opus-4.60.6071
- 6anthropic/claude-sonnet-4.60.5893
- 7openai/gpt-5.3-chat0.585
- 8meta-llama/llama-4-maverick0.5746
- 9openai/gpt-5.40.5675
- 10anthropic/claude-haiku-4.50.55
- 11meta-llama/llama-4-scout0.5484
- 12google/gemini-2.5-pro0.5295
- 13google/gemini-3.1-pro-preview0.5155
- 14GPT-3 (text-davinci)0.4647
- 15T5 (closed-book QA)0.3604
- 16RAG (DPR reader)0.3318