Benchmark

HealthBench

healthcare text

An open-source benchmark for measuring performance and safety of large language models in healthcare, consisting of 5,000 multi-turn conversations evaluated by 262 physicians using 48,562 unique rubric criteria across health contexts and behavioral dimensions

语言EN
满分1
参评模型2

模型排名

名次 模型 机构 分数 来源
1 GPT OSS 120B OpenAI 57.6 来源 ↗
2 GPT OSS 20B OpenAI 42.5 来源 ↗