llmboard.ai
Benchmarks
CompareRankings
llmboard.ai
Benchmarks
CompareRankings
HomeBenchmarksmultimodalMMMU-Pro

multimodal benchmark

MMMU-Pro

A more robust multi-discipline multimodal understanding benchmark that enhances MMMU through a three-step process: filtering text-only answerable questions, augmenting candidate options, and introducing vision-only input settings. Achieves significantly lower model performance (16.8-26.9%) compared to original MMMU, providing more rigorous evaluation that closely mimics real-world scenarios.

Updated Aug 7, 2026

Published models65
Registry coverage65
MetricScore
EvidenceB

On this page

  • Ranking
  • Distribution
  • Highlights
  • About
  • FAQ

MMMU-Pro leaderboard

Sorted by the source-provided rank. Higher score is better according to the registry.

65 rows
Columns

Show columns

1GOGemini 3.5 FlashGoogle83.6%100.0%65CAug 7, 2026
2OPGPT-5.5OpenAI83.2%98.4%65CAug 7, 2026
3OPGPT-5.6 SolOpenAI83.0%96.9%65CAug 7, 2026
4BYSeed 2.1 ProByteDance82.7%95.3%65CAug 7, 2026
5ACQwen3.8 MaxAlibaba Cloud / Qwen Team82.3%93.8%65CAug 7, 2026
6BYSeed 2.1 TurboByteDance82.2%92.2%65CAug 7, 2026
7MAKimi K3Moonshot AI81.6%90.6%65CAug 7, 2026
8GOGemini 3 FlashGoogle81.2%89.1%65CAug 7, 2026
9OPGPT-5.4OpenAI81.2%87.5%65CAug 7, 2026
10GOGemini 3 ProGoogle81.0%85.9%65CAug 7, 2026
11OPGPT-5.6 TerraOpenAI80.7%84.4%65CAug 7, 2026
12GOGemini 3.1 ProGoogle80.5%82.8%65CAug 7, 2026
13MEMuse SparkMeta80.4%81.3%65CAug 7, 2026
14MAKimi K2.6Moonshot AI80.1%79.7%65CAug 7, 2026
15OPGPT-5.2OpenAI79.5%78.1%65CAug 7, 2026
16ACQwen3.7-PlusAlibaba Cloud / Qwen Team79.0%76.6%65CAug 7, 2026
17ACQwen3.6 PlusAlibaba Cloud / Qwen Team78.8%75.0%65CAug 7, 2026
18MAKimi K2.5Moonshot AI78.5%73.4%65CAug 7, 2026
19OPGPT-5OpenAI78.4%71.9%65CAug 7, 2026
20OPGPT-5.6 LunaOpenAI78.4%70.3%65CAug 7, 2026
21MIMiniMax M3MiniMax78.1%68.8%65CAug 7, 2026
22XIMiMo-V2.5Xiaomi77.9%67.2%65CAug 7, 2026
23ANClaude Opus 4.6Anthropic77.3%65.6%65CAug 7, 2026
24GOGemma 4 31BGoogle76.9%64.1%65CAug 7, 2026
25ACQwen3.5-122B-A10BAlibaba Cloud / Qwen Team76.9%62.5%65CAug 7, 2026
26GOGemini 3.1 Flash-LiteGoogle76.8%60.9%65CAug 7, 2026
27OPGPT-5.4 miniOpenAI76.6%59.4%65CAug 7, 2026
28OPo3OpenAI76.4%57.8%65CAug 7, 2026
29OPGPT-5.5 InstantOpenAI76.0%56.3%65CAug 7, 2026
30ACQwen3.6-27BAlibaba Cloud / Qwen Team75.8%54.7%65CAug 7, 2026
31ANClaude Sonnet 4.6Anthropic75.6%53.1%65CAug 7, 2026
32ACQwen3.6-35B-A3BAlibaba Cloud / Qwen Team75.3%51.6%65CAug 7, 2026
33ACQwen3.5-35B-A3BAlibaba Cloud / Qwen Team75.1%50.0%65CAug 7, 2026
34ACQwen3.5-27BAlibaba Cloud / Qwen Team75.0%48.4%65CAug 7, 2026
35GOGemma 4 26B-A4BGoogle73.8%46.9%65CAug 7, 2026
36ACQwen3 VL 235B A22B ThinkingAlibaba Cloud / Qwen Team69.3%45.3%65CAug 7, 2026
37GOGemma 4 12BGoogle69.1%43.8%65CAug 7, 2026
38ACQwen3 VL 235B A22B InstructAlibaba Cloud / Qwen Team68.1%42.2%65CAug 7, 2026
39ACQwen3 VL 32B ThinkingAlibaba Cloud / Qwen Team68.1%40.6%65CAug 7, 2026
40OPGPT-5.4 nanoOpenAI66.1%39.1%65CAug 7, 2026
41ACQwen3 VL 32B InstructAlibaba Cloud / Qwen Team65.3%37.5%65CAug 7, 2026
42AMNova 2 ProAmazon63.5%35.9%65CAug 7, 2026
43COCommand A+Cohere63.0%34.4%65CAug 7, 2026
44ACQwen3 VL 30B A3B ThinkingAlibaba Cloud / Qwen Team63.0%32.8%65CAug 7, 2026
45AMNova 2 LiteAmazon61.8%31.3%65CAug 7, 2026
46AMNova 2 OmniAmazon61.4%29.7%65CAug 7, 2026
47ACQwen3 VL 30B A3B InstructAlibaba Cloud / Qwen Team60.4%28.1%65CAug 7, 2026
48ACQwen3 VL 8B ThinkingAlibaba Cloud / Qwen Team60.4%26.6%65CAug 7, 2026
49MAMistral Small 4Mistral AI60.0%25.0%65CAug 7, 2026
50OPGPT-4oOpenAI59.9%23.4%65CAug 7, 2026
51MELlama 4 MaverickMeta59.6%21.9%65CAug 7, 2026
52ACQwen3 VL 4B ThinkingAlibaba Cloud / Qwen Team57.0%20.3%65CAug 7, 2026
53ACQwen3 VL 8B InstructAlibaba Cloud / Qwen Team55.9%18.8%65CAug 7, 2026
54GODiffusionGemma 26B-A4BGoogle54.3%17.2%65CAug 7, 2026
55ACQwen3 VL 4B InstructAlibaba Cloud / Qwen Team53.2%15.6%65CAug 7, 2026
56GOGemma 4 E4BGoogle52.6%14.1%65CAug 7, 2026
57ACQwen2.5 VL 72B InstructAlibaba Cloud / Qwen Team51.1%12.5%65CAug 7, 2026
58ACQwen2.5 VL 32B InstructAlibaba Cloud / Qwen Team49.5%10.9%65CAug 7, 2026
59ACQwen2-VL-72B-InstructAlibaba Cloud / Qwen Team46.2%9.4%65CAug 7, 2026
60MELlama 3.2 90B InstructMeta45.2%7.8%65CAug 7, 2026
61GOGemma 4 E2BGoogle44.2%6.3%65CAug 7, 2026
62MIPhi-4-multimodal-instructMicrosoft38.5%4.7%65CAug 7, 2026
63ACQwen2.5 VL 7B InstructAlibaba Cloud / Qwen Team38.3%3.1%65CAug 7, 2026
64ACQwen2.5-Omni-7BAlibaba Cloud / Qwen Team36.6%1.6%65CAug 7, 2026
65MELlama 3.2 11B InstructMeta33.0%0.0%65CAug 7, 2026

Score distribution

Top published rows on the benchmark's original scale.

MMMU-Pro

MMMU-Pro highlights

The top published results on this benchmark's own scale.

Rank #1Gemini 3.5 Flash83.6%Rank #2GPT-5.583.2%Rank #3GPT-5.6 Sol83.0%Rank #4Seed 2.1 Pro82.7%

What is MMMU-Pro?

Definition and scoring fields from the benchmark registry.

A more robust multi-discipline multimodal understanding benchmark that enhances MMMU through a three-step process: filtering text-only answerable questions, augmenting candidate options, and introducing vision-only input settings. Achieves significantly lower model performance (16.8-26.9%) compared to original MMMU, providing more rigorous evaluation that closely mimics real-world scenarios.

Scores are shown in ratio. The current registry marks this benchmark as not independently verified with evidence level B.

Family
MMMU-Pro
Modality
multimodal
Primary category
multimodal
Score direction
higher
LLMBoard eligible
No
Evaluation key
mmmu-pro|llm-stats-current

Source-native results are preserved. Eligibility for the overall LLMBoard score is a separate policy decision.

FAQ

Common questions about MMMU-Pro.

Which model scores highest on MMMU-Pro?

Gemini 3.5 Flash is currently ranked first with 83.6%.

What does MMMU-Pro measure?

A more robust multi-discipline multimodal understanding benchmark that enhances MMMU through a three-step process: filtering text-only answerable questions, augmenting candidate options, and introducing vision-only input settings. Achieves significantly lower model performance (16.8-26.9%) compared to original MMMU, providing more rigorous evaluation that closely mimics real-world scenarios.

Is a higher score better?

Yes. Higher values rank better for this benchmark.

How many models are compared?

65 unique published model results are currently shown.

Does this benchmark affect the overall score?

This benchmark is preserved as source-native evidence but is not eligible for the current overall score.

Rankings

OverallCodingText ArenaPricing

Modalities

Image GenerationVideo GenerationSpeech-to-TextEmbeddings

Benchmarks

All BenchmarksReasoningMathCoding

Vendors

All VendorsOpenAIAnthropicGoogle
llmboard.aiCopyright 2026 llmboard.ai