General
BIRD (text-to-SQL)
BIRD: a large-scale, cross-domain text-to-SQL benchmark (NL question + gold SQL over 95 real databases, 37 domains), scored by execution accuracy. The official repo ships questions/gold SQL/eval harness but no baseline predictions; this build ingests petavue/NL2SQL-Benchmark's public 'Inference Level Dataset' of 43,200 per-inference rows over a 360-question subset of BIRD dev. Each response is one model's binary execution-match correctness on one question under one prompting config; items are the real NL questions, correct_answer is the gold SQL. Models served under multiple provider aliases are merged to one canonical subject, with the serving platform recorded in test_condition.
Response matrix
Fit to width. Hover for subject & item; click a cell for details.

Scale: 1 = correct · 0 = incorrect
Subjects
- 1WizardLM/WizardCoder-33B-V1.10.31
- 2defog/sqlcoder-70b-alpha0.3
- 3databricks/dbrx-instruct0.2439
- 4defog/sqlcoder-7b-20.2272
- 5codellama/CodeLlama-34b-Instruct-hf0.1683
- 6mistralai/Mixtral-8x7B-Instruct-v0.10.1296
- 7codellama/CodeLlama-70b-Instruct-hf0.1136
- 8mistralai/Mistral-7B-Instruct-v0.20.0853
- 9mistralai/Mistral-7B-Instruct-v0.10.0353
- 10meta-llama/Llama-2-70b-chat-hf0.0106
- 11claude-3-haiku-202403070
- 12claude-3-sonnet-202402290
- 13gemini-1.0-pro0
- 14gpt-4-turbo-preview0
- 15claude-3-opus-202402290
- 16gpt-3.5-turbo-16k0