Benchmark

FACTS Grounding

reasoning text

A benchmark evaluating language models' ability to generate factually accurate and well-grounded responses based on long-form input context, comprising 1,719 examples with documents up to 32k tokens requiring detailed responses that are fully grounded in provided documents

语言EN
满分1
参评模型9

模型排名

名次 模型 机构 分数 来源
1 Gemini 2.5 Pro Preview 06-05 Google 87.8 来源 ↗
2 Gemini 2.5 Flash Google 85.3 来源 ↗
3 Gemini 2.5 Flash-Lite Google 84.1 来源 ↗
4 Gemini 2.0 Flash Google 83.6 来源 ↗
5 Gemini 2.0 Flash-Lite Google 83.6 来源 ↗
6 Gemma 3 12B Google 75.8 来源 ↗
7 Gemma 3 27B Google 74.9 来源 ↗
8 Gemma 3 4B Google 70.1 来源 ↗
9 Gemma 3 1B Google 36.4 来源 ↗