Law
LexEval
LexEval tests how well models answer Chinese legal multiple-choice questions covering legal knowledge, reasoning, discrimination and ethics, grading each answer by exact match of the extracted option letters against the gold answer using the official LexEval grader.
Response matrix
LexEval contains 23 tasks and 14,150 questions in total. This page includes the 19 objective multiple-choice tasks, comprising 12,468 questions. The four L5 Generation tasks, which evaluate legal drafting, are excluded because the upstream benchmark scores them using ROUGE-L, a continuous text-similarity metric for which no categorical correctness threshold is specified. Their exclusion also explains why the ability levels shown here progress directly from L4 to L6. The upstream release provides each model’s raw answers rather than per-item correctness labels. We therefore grade responses using the authors’ published evaluation procedure: exact matching between the option letters extracted from each response and the corresponding gold answer. Unobserved cells reflect the benchmark’s original evaluation coverage.
Loading session strips…
Scale: 1 = correct · 0 = incorrect