Video-MME
Canonical page on the main site: sinoaihub.com/benchmarks/video-mme
Description
Video understanding benchmark spanning various video durations and domains.
Evaluations
| benchmark | model | score | model_version | metric | date | source_type | source_url |
|---|---|---|---|---|---|---|---|
| Video-MME | kimi-k3 | 90.0 | with subtitles | accuracy | 2026-07 | vendor_reported | https://github.com/MoonshotAI/Kimi-K3 |
| Video-MME | kimi-k2.5 | 87.4 | — | accuracy | — | vendor_reported | https://github.com/MoonshotAI/Kimi-K2.5 |
Methodology
Task type: Video understanding (multimodal video analysis)
Dataset size: 900 videos (254 hours total), 2,700 human-annotated question-answer pairs
Evaluation method: Video QA with subtitles and audio modalities; duration-stratified evaluation
Scoring: Accuracy (% correct); variants with/without subtitles
Relevant Models
- Kimi K3 — 90.0
Limitations
All scores are vendor-reported and not independently verified. Subtitle usage differs between evaluations.
Last Verified
2026-09-20
Sources
| source_name | source_url | source_type | last_verified | confidence |
|---|---|---|---|---|
| Kimi K3 GitHub README | https://github.com/MoonshotAI/Kimi-K3 | official | 2026-09-20 | high |
| Kimi K2.5 GitHub README | https://github.com/MoonshotAI/Kimi-K2.5 | official | 2026-09-20 | high |