Benchmark
TruthfulQA
general
reasoning
legal
healthcare
finance
text
TruthfulQA is a benchmark to measure whether language models are truthful in generating answers to questions. It comprises 817 questions that span 38 categories, including health, law, finance and politics. The questions are crafted such that some humans would answer falsely due to a false belief or misconception, testing models' ability to avoid generating false answers learned from human texts.
语言EN
满分1
参评模型16
模型排名
| 名次 | 模型 | 机构 | 分数 | 来源 |
|---|---|---|---|---|
| 1 | Phi-3.5-MoE-instruct | Microsoft | 77.5 | 来源 ↗ |
| 2 | Granite 3.3 8B Instruct | IBM | 66.9 | 来源 ↗ |
| 3 | Phi 4 Mini | Microsoft | 66.4 | 来源 ↗ |
| 4 | Phi-3.5-mini-instruct | Microsoft | 64.0 | 来源 ↗ |
| 5 | Llama 3.1 Nemotron 70B Instruct | NVIDIA | 58.6 | 来源 ↗ |
| 6 | Qwen2.5 14B Instruct | Alibaba Cloud / Qwen Team | 58.4 | 来源 ↗ |
| 7 | Jamba 1.5 Large | AI21 Labs | 58.3 | 来源 ↗ |
| 8 | IBM Granite 4.0 Tiny Preview | IBM | 58.1 | 来源 ↗ |
| 9 | Qwen2.5 32B Instruct | Alibaba Cloud / Qwen Team | 57.8 | 来源 ↗ |
| 10 | Command R+ | Cohere | 56.3 | 来源 ↗ |
| 11 | Qwen2 72B Instruct | Alibaba Cloud / Qwen Team | 54.8 | 来源 ↗ |
| 12 | Qwen2.5-Coder 32B Instruct | Alibaba Cloud / Qwen Team | 54.2 | 来源 ↗ |
| 13 | Jamba 1.5 Mini | AI21 Labs | 54.1 | 来源 ↗ |
| 14 | Granite 3.3 8B Base | IBM | 52.2 | 来源 ↗ |
| 15 | Qwen2.5-Coder 7B Instruct | Alibaba Cloud / Qwen Team | 50.6 | 来源 ↗ |
| 16 | Mistral NeMo Instruct | Mistral AI | 50.3 | 来源 ↗ |