Law
LawBench
LawBench measures how well models handle Chinese legal work; this curation ingests the six subtasks graded by per-item binary accuracy — judicial-exam and case-analysis multiple choice, dispute-focus and consultation classification, argumentation mining, and crime-amount extraction — regraded from the released raw predictions with the official evaluation code.
Response matrix
LawBench's full bank is 20 subtasks; this page carries the six whose official metric is per-item binary accuracy — judicial-exam and case-analysis multiple choice, dispute-focus and consultation classification, argumentation-mining multiple choice, and crime-amount extraction (2,995 questions). The other 14 subtasks are excluded: upstream scores them with continuous corpus metrics — ROUGE-L, F1 variants, normalized log-distance and F0.5 — that it defines no categorical threshold for. Upstream releases each model's raw predictions rather than per-item grades, so responses here are graded with the authors' own published evaluation code; a prediction naming no option counts as wrong, exactly as in the official accuracy. The zero-shot and one-shot bands are upstream's two released prompt settings. The zero-shot dispute-focus predictions were generated from an older ordering of that task file, so 12 of its 496 questions have no zero-shot observation.
Loading session strips…
Scale: 1 = correct · 0 = incorrect