Benchmark

TruthfulQA

general reasoning legal healthcare finance text

TruthfulQA is a benchmark to measure whether language models are truthful in generating answers to questions. It comprises 817 questions that span 38 categories, including health, law, finance and politics. The questions are crafted such that some humans would answer falsely due to a false belief or misconception, testing models' ability to avoid generating false answers learned from human texts.

语言EN
满分1
参评模型16

模型排名

名次 模型 机构 分数 来源
1 Phi-3.5-MoE-instruct Microsoft 77.5 来源 ↗
2 Granite 3.3 8B Instruct IBM 66.9 来源 ↗
3 Phi 4 Mini Microsoft 66.4 来源 ↗
4 Phi-3.5-mini-instruct Microsoft 64.0 来源 ↗
5 Llama 3.1 Nemotron 70B Instruct NVIDIA 58.6 来源 ↗
6 Qwen2.5 14B Instruct Alibaba Cloud / Qwen Team 58.4 来源 ↗
7 Jamba 1.5 Large AI21 Labs 58.3 来源 ↗
8 IBM Granite 4.0 Tiny Preview IBM 58.1 来源 ↗
9 Qwen2.5 32B Instruct Alibaba Cloud / Qwen Team 57.8 来源 ↗
10 Command R+ Cohere 56.3 来源 ↗
11 Qwen2 72B Instruct Alibaba Cloud / Qwen Team 54.8 来源 ↗
12 Qwen2.5-Coder 32B Instruct Alibaba Cloud / Qwen Team 54.2 来源 ↗
13 Jamba 1.5 Mini AI21 Labs 54.1 来源 ↗
14 Granite 3.3 8B Base IBM 52.2 来源 ↗
15 Qwen2.5-Coder 7B Instruct Alibaba Cloud / Qwen Team 50.6 来源 ↗
16 Mistral NeMo Instruct Mistral AI 50.3 来源 ↗