General

ComplexBench

ComplexBench: complex instruction following with multiple composed constraints (NeurIPS 2024). Each instruction decomposes into binary scoring questions; this build ingests the deterministic, rule-verified subset of those questions, grading each released model generation with the benchmark's own rule evaluator.

252items
15subjects
MITlicense
generaldomain
textmodality
item-level responses released
Saturation status: Unknown

Response matrix

Fit to width. Hover for subject & item; click a cell for details.

ComplexBench response matrix: AI models (rows) against items (columns)
Correct (1)Incorrect (0)Unobserved

Scale: 1 = correct · 0 = incorrect

Subjects

  1. 1gpt3.5_turbo_11060.7063
  2. 2erniebot_40.696
  3. 3gpt4_11060.6865
  4. 4glm40.6825
  5. 5claude3_opus0.6786
  6. 6qwen1.5_72b0.619
  7. 7qwen1.5_14b0.5952
  8. 8llama3_70b0.5833
  9. 9qwen1.5_7b0.5833
  10. 10internlm2_20b0.5833
  11. 11baichuan2_13b0.5635
  12. 12chatglm3_6b0.5635
  13. 13internlm2_7b0.5595
  14. 14llama3_8b0.5119
  15. 15mistral_7b0.4603