Benchmark
TriviaQA
general
reasoning
text
A large-scale reading comprehension dataset containing over 650K question-answer-evidence triples. TriviaQA includes 95K question-answer pairs authored by trivia enthusiasts and independently gathered evidence documents (six per question on average) that provide high quality distant supervision for answering the questions. The dataset features relatively complex, compositional questions with considerable syntactic and lexical variability, requiring cross-sentence reasoning to find answers.
语言EN
满分1
参评模型13
模型排名
| 名次 | 模型 | 机构 | 分数 | 来源 |
|---|---|---|---|---|
| 1 | Kimi K2 Base | Moonshot AI | 85.1 | 来源 ↗ |
| 2 | Gemma 2 27B | 83.7 | 来源 ↗ | |
| 3 | Mistral Small 3.1 24B Base | Mistral AI | 80.5 | 来源 ↗ |
| 4 | Mistral Small 3.1 24B Instruct | Mistral AI | 80.5 | 来源 ↗ |
| 5 | Mistral Small 3 24B Base | Mistral AI | 80.3 | 来源 ↗ |
| 6 | Granite 3.3 8B Base | IBM | 78.2 | 来源 ↗ |
| 7 | Gemma 2 9B | 76.6 | 来源 ↗ | |
| 8 | Mistral NeMo Instruct | Mistral AI | 73.8 | 来源 ↗ |
| 9 | Gemma 3n E4B | 70.2 | 来源 ↗ | |
| 10 | Gemma 3n E4B Instructed LiteRT Preview | 70.2 | 来源 ↗ | |
| 11 | Ministral 8B Instruct | Mistral AI | 65.5 | 来源 ↗ |
| 12 | Gemma 3n E2B | 60.8 | 来源 ↗ | |
| 13 | Gemma 3n E2B Instructed LiteRT (Preview) | 60.8 | 来源 ↗ |