Benchmark
HealthBench Hard
healthcare
text
A challenging variation of HealthBench that evaluates large language models' performance and safety in healthcare through 5,000 multi-turn conversations with particularly rigorous evaluation criteria validated by 262 physicians from 60 countries
语言EN
满分1
参评模型3
模型排名
| 名次 | 模型 | 机构 | 分数 | 来源 |
|---|---|---|---|---|
| 1 | GPT OSS 120B | OpenAI | 30.0 | 来源 ↗ |
| 2 | GPT OSS 20B | OpenAI | 10.8 | 来源 ↗ |
| 3 | GPT-5 | OpenAI | 1.6 | 来源 ↗ |