Benchmark
CharadesSTA
video
language
multimodal
multimodal
Charades-STA is a benchmark dataset for temporal activity localization via language queries, extending the Charades dataset with sentence temporal annotations. It contains 12,408 training and 3,720 testing segment-sentence pairs from videos with natural language descriptions and precise temporal boundaries for localizing activities based on language queries.
语言EN
满分1
参评模型2
模型排名
| 名次 | 模型 | 机构 | 分数 | 来源 |
|---|---|---|---|---|
| 1 | Qwen2.5 VL 32B Instruct | Alibaba Cloud / Qwen Team | 54.2 | 来源 ↗ |
| 2 | Qwen2.5 VL 7B Instruct | Alibaba Cloud / Qwen Team | 43.6 | 来源 ↗ |