reasoning benchmark
A diagnostic benchmark for very long-form video language understanding consisting of over 5000 human curated multiple choice questions based on 3-minute video clips from Ego4D, covering a broad range of natural human activities and behaviors
Updated Aug 7, 2026
Sorted by the source-provided rank. Higher score is better according to the registry.
| 1 | AC | 77.9% | 100.0% | 9 | C | |
| 2 | AC | 76.2% | 87.5% | 9 | C | |
| 3 | OP | 72.2% | 75.0% | 9 | C | |
| 4 | AM | 72.1% | 62.5% | 9 | C | |
| 5 | GO | 71.5% | 50.0% | 9 | C | |
| 6 | AM | 71.4% | 37.5% | 9 | C | |
| 7 | AC | 68.6% | 25.0% | 9 | C | |
| 8 | GO | 67.2% | 12.5% | 9 | C | |
| 9 | GO | 55.7% | 0.0% | 9 | C |
Top published rows on the benchmark's original scale.
The top published results on this benchmark's own scale.
Definition and scoring fields from the benchmark registry.
A diagnostic benchmark for very long-form video language understanding consisting of over 5000 human curated multiple choice questions based on 3-minute video clips from Ego4D, covering a broad range of natural human activities and behaviors
Scores are shown in ratio. The current registry marks this benchmark as not independently verified with evidence level B.
Source-native results are preserved. Eligibility for the overall LLMBoard score is a separate policy decision.
Common questions about EgoSchema.
Qwen2-VL-72B-Instruct is currently ranked first with 77.9%.
A diagnostic benchmark for very long-form video language understanding consisting of over 5000 human curated multiple choice questions based on 3-minute video clips from Ego4D, covering a broad range of natural human activities and behaviors
Yes. Higher values rank better for this benchmark.
9 unique published model results are currently shown.
This benchmark is preserved as source-native evidence but is not eligible for the current overall score.