Benchmark

Bird-SQL (dev)

reasoning text

BIRD (BIg Bench for LaRge-scale Database Grounded Text-to-SQLs) is a comprehensive text-to-SQL benchmark containing 12,751 question-SQL pairs across 95 databases (33.4 GB total) spanning 37+ professional domains. It evaluates large language models' ability to convert natural language to executable SQL queries in real-world scenarios with complex database schemas and dirty data.

语言EN
满分1
参评模型6

模型排名

名次 模型 机构 分数 来源
1 Gemini 2.0 Flash-Lite Google 57.4 来源 ↗
2 Gemini 2.0 Flash Google 56.9 来源 ↗
3 Gemma 3 27B Google 54.4 来源 ↗
4 Gemma 3 12B Google 47.9 来源 ↗
5 Gemma 3 4B Google 36.3 来源 ↗
6 Gemma 3 1B Google 6.4 来源 ↗