Medicine

ClinBench

ClinBench released ablation study: per-item lung-cancer staging from TCGA pathology reports. Binary (gold == prediction) responses for 3 models across four extracted fields (pT, pN, tumor_stage, histologic_diagnosis) over 774 TCGA lung cases. The main 11-model benchmark is aggregate-F1 only with access-gated notes; this is the one released per-(model, item) artifact.

774items
3subjects
Apache-2.0license
medicinedomain
nlp_taskdomain
textmodality
item-level responses released
Saturation status: Unknown

Response matrix

Fit to width. Hover for subject & item; click a cell for details.

ClinBench response matrix: AI models (rows) against items (columns)
Correct (1)Incorrect (0)Unobserved

Scale: 1 = correct · 0 = incorrect

Subjects

  1. 1gpt-4o-2024-05-130.8751
  2. 2gpt-4o-mini0.8092
  3. 3gpt-3.5-turbo-11060.7794