Benchmark

MEGA MLQA

language reasoning text 多语言

MLQA as part of the MEGA (Multilingual Evaluation of Generative AI) benchmark suite. A multi-way aligned extractive QA evaluation benchmark for cross-lingual question answering across 7 languages (English, Arabic, German, Spanish, Hindi, Vietnamese, and Simplified Chinese) with over 12K QA instances in English and 5K in each other language.

语言EN
满分1
参评模型2

模型排名

名次 模型 机构 分数 来源
1 Phi-3.5-MoE-instruct Microsoft 65.3 来源 ↗
2 Phi-3.5-mini-instruct Microsoft 61.7 来源 ↗