General

BIRD (text-to-SQL)

BIRD: a large-scale, cross-domain text-to-SQL benchmark (NL question + gold SQL over 95 real databases, 37 domains), scored by execution accuracy. The official repo ships questions/gold SQL/eval harness but no baseline predictions; this build ingests petavue/NL2SQL-Benchmark's public 'Inference Level Dataset' of 43,200 per-inference rows over a 360-question subset of BIRD dev. Each response is one model's binary execution-match correctness on one question under one prompting config; items are the real NL questions, correct_answer is the gold SQL. Models served under multiple provider aliases are merged to one canonical subject, with the serving platform recorded in test_condition.

360items
16subjects
MITlicense
software_engineeringdomain
textmodality
item-level responses released
Saturation status: Unknown

Response matrix

Fit to width. Hover for subject & item; click a cell for details.

BIRD (text-to-SQL) response matrix: AI models (rows) against items (columns)
Correct (1)Incorrect (0)Unobserved

Scale: 1 = correct · 0 = incorrect

Subjects

  1. 1WizardLM/WizardCoder-33B-V1.10.31
  2. 2defog/sqlcoder-70b-alpha0.3
  3. 3databricks/dbrx-instruct0.2439
  4. 4defog/sqlcoder-7b-20.2272
  5. 5codellama/CodeLlama-34b-Instruct-hf0.1683
  6. 6mistralai/Mixtral-8x7B-Instruct-v0.10.1296
  7. 7codellama/CodeLlama-70b-Instruct-hf0.1136
  8. 8mistralai/Mistral-7B-Instruct-v0.20.0853
  9. 9mistralai/Mistral-7B-Instruct-v0.10.0353
  10. 10meta-llama/Llama-2-70b-chat-hf0.0106
  11. 11claude-3-haiku-202403070
  12. 12claude-3-sonnet-202402290
  13. 13gemini-1.0-pro0
  14. 14gpt-4-turbo-preview0
  15. 15claude-3-opus-202402290
  16. 16gpt-3.5-turbo-16k0