Benchmark

FActScore

reasoning text

A fine-grained atomic evaluation metric for factual precision in long-form text generation that breaks generated text into atomic facts and computes the percentage supported by reliable knowledge sources, with automated assessment using retrieval and language models

语言EN
满分1
参评模型1

模型排名

名次 模型 机构 分数 来源
1 GPT-5 OpenAI 1.0 来源 ↗