Science

CROP

CROP benchmark: 5,045 bilingual (Chinese/English) crop-science multiple-choice questions across three difficulty levels, derived from 2K+ crop-science academic papers. Each item asks the model to select the correct option; the response is whether the chosen option was correct. We ingest the authors' released per-(model, item) grading matrix for four commercial LLMs.

5,045items
4subjects
CC-BY-NC-4.0license
sciencedomain
knowledgedomain
multilingualdomain
textmodality
item-level responses released
Saturation status: Unknown

Response matrix

Fit to width. Hover for subject & item; click a cell for details.

CROP response matrix: AI models (rows) against items (columns)
Correct (1)Incorrect (0)Unobserved

Scale: 1 = correct · 0 = incorrect

Subjects

  1. 1claude-3-opus-202402290.9001
  2. 2qwen-max0.8664
  3. 3gpt-4-turbo-2024-04-090.8569
  4. 4gpt-3.5-turbo-01250.3286