Benchmark

CRAG

reasoning search text

CRAG (Comprehensive RAG Benchmark) is a factual question answering benchmark consisting of 4,409 question-answer pairs across 5 domains (finance, sports, music, movie, open domain) and 8 question categories. The benchmark includes mock APIs to simulate web and Knowledge Graph search, designed to represent the diverse and dynamic nature of real-world QA tasks with temporal dynamism ranging from years to seconds. It evaluates retrieval-augmented generation systems for trustworthy question answering.

语言EN
满分1
参评模型3

模型排名

名次 模型 机构 分数 来源
1 Nova Pro Amazon 50.3 来源 ↗
2 Nova Lite Amazon 43.8 来源 ↗
3 Nova Micro Amazon 43.1 来源 ↗