Benchmark
Bird-SQL (dev)
reasoning
text
BIRD (BIg Bench for LaRge-scale Database Grounded Text-to-SQLs) is a comprehensive text-to-SQL benchmark containing 12,751 question-SQL pairs across 95 databases (33.4 GB total) spanning 37+ professional domains. It evaluates large language models' ability to convert natural language to executable SQL queries in real-world scenarios with complex database schemas and dirty data.
语言EN
满分1
参评模型6
模型排名
| 名次 | 模型 | 机构 | 分数 | 来源 |
|---|---|---|---|---|
| 1 | Gemini 2.0 Flash-Lite | 57.4 | 来源 ↗ | |
| 2 | Gemini 2.0 Flash | 56.9 | 来源 ↗ | |
| 3 | Gemma 3 27B | 54.4 | 来源 ↗ | |
| 4 | Gemma 3 12B | 47.9 | 来源 ↗ | |
| 5 | Gemma 3 4B | 36.3 | 来源 ↗ | |
| 6 | Gemma 3 1B | 6.4 | 来源 ↗ |