multimodal benchmark
Video-MME is the first-ever comprehensive evaluation benchmark of Multi-modal Large Language Models (MLLMs) in video analysis. It features 900 videos totaling 254 hours with 2,700 human-annotated question-answer pairs across 6 primary visual domains (Knowledge, Film & Television, Sports Competition, Life Record, Multilingual, and others) and 30 subfields. The benchmark evaluates models across diverse temporal dimensions (11 seconds to 1 hour), integrates multi-modal inputs including video frames, subtitles, and audio, and uses rigorous manual labeling by expert annotators for precise assessment.
Updated Aug 7, 2026
Sorted by the source-provided rank. Higher score is better according to the registry.
| 1 | BY | 89.2% | 100.0% | 17 | C | |
| 2 | BY | 89.0% | 93.8% | 17 | C | |
| 3 | AC | 88.0% | 87.5% | 17 | C | |
| 4 | XI | 87.7% | 81.3% | 17 | C | |
| 5 | MA | 87.4% | 75.0% | 17 | C | |
| 6 | MI | 85.4% | 68.8% | 17 | C | |
| 7 | GO | 84.8% | 62.5% | 17 | C | |
| 8 | AC | 84.2% | 56.3% | 17 | C | |
| 9 | GO | 78.6% | 50.0% | 17 | C | |
| 10 | AM | 77.9% | 43.8% | 17 | C | |
| 11 | GO | 76.1% | 37.5% | 17 | C | |
| 12 | AC | 74.5% | 31.3% | 17 | C | |
| 13 | AC | 73.3% | 25.0% | 17 | C | |
| 14 | AC | 71.8% | 18.8% | 17 | C | |
| 15 | AC | 71.4% | 12.5% | 17 | C | |
| 16 | GO | 66.2% | 6.3% | 17 | C | |
| 17 | MI | 55.0% | 0.0% | 17 | C |
Top published rows on the benchmark's original scale.
The top published results on this benchmark's own scale.
Definition and scoring fields from the benchmark registry.
Video-MME is the first-ever comprehensive evaluation benchmark of Multi-modal Large Language Models (MLLMs) in video analysis. It features 900 videos totaling 254 hours with 2,700 human-annotated question-answer pairs across 6 primary visual domains (Knowledge, Film & Television, Sports Competition, Life Record, Multilingual, and others) and 30 subfields. The benchmark evaluates models across diverse temporal dimensions (11 seconds to 1 hour), integrates multi-modal inputs including video frames, subtitles, and audio, and uses rigorous manual labeling by expert annotators for precise assessment.
Scores are shown in ratio. The current registry marks this benchmark as not independently verified with evidence level B.
Source-native results are preserved. Eligibility for the overall LLMBoard score is a separate policy decision.
Common questions about Video-MME.
Seed 2.1 Pro is currently ranked first with 89.2%.
Video-MME is the first-ever comprehensive evaluation benchmark of Multi-modal Large Language Models (MLLMs) in video analysis. It features 900 videos totaling 254 hours with 2,700 human-annotated question-answer pairs across 6 primary visual domains (Knowledge, Film & Television, Sports Competition, Life Record, Multilingual, and others) and 30 subfields. The benchmark evaluates models across diverse temporal dimensions (11 seconds to 1 hour), integrates multi-modal inputs including video frames, subtitles, and audio, and uses rigorous manual labeling by expert annotators for precise assessment.
Yes. Higher values rank better for this benchmark.
17 unique published model results are currently shown.
This benchmark is preserved as source-native evidence but is not eligible for the current overall score.