Benchmark

CharadesSTA

video language multimodal multimodal

Charades-STA is a benchmark dataset for temporal activity localization via language queries, extending the Charades dataset with sentence temporal annotations. It contains 12,408 training and 3,720 testing segment-sentence pairs from videos with natural language descriptions and precise temporal boundaries for localizing activities based on language queries.

语言EN
满分1
参评模型2

模型排名

名次 模型 机构 分数 来源
1 Qwen2.5 VL 32B Instruct Alibaba Cloud / Qwen Team 54.2 来源 ↗
2 Qwen2.5 VL 7B Instruct Alibaba Cloud / Qwen Team 43.6 来源 ↗