llmboard.ai
Benchmarks
CompareRankings
llmboard.ai
Benchmarks
CompareRankings
HomeBenchmarksgeneralMLVU-M

general benchmark

MLVU-M

MLVU-M benchmark

Updated Aug 7, 2026

Published models8
Registry coverage8
MetricScore
EvidenceB

On this page

  • Ranking
  • Distribution
  • Highlights
  • About
  • FAQ

MLVU-M leaderboard

Sorted by the source-provided rank. Higher score is better according to the registry.

8 rows
Columns

Show columns

1ACQwen3 VL 32B InstructAlibaba Cloud / Qwen Team82.1%100.0%8CAug 7, 2026
2ACQwen3 VL 30B A3B InstructAlibaba Cloud / Qwen Team81.3%85.7%8CAug 7, 2026
3ACQwen3 VL 30B A3B ThinkingAlibaba Cloud / Qwen Team78.9%71.4%8CAug 7, 2026
4ACQwen3 VL 8B InstructAlibaba Cloud / Qwen Team78.1%57.1%8CAug 7, 2026
5ACQwen3 VL 4B ThinkingAlibaba Cloud / Qwen Team75.7%42.9%8CAug 7, 2026
6ACQwen3 VL 4B InstructAlibaba Cloud / Qwen Team75.3%28.6%8CAug 7, 2026
7ACQwen3 VL 8B ThinkingAlibaba Cloud / Qwen Team75.1%14.3%8CAug 7, 2026
8ACQwen2.5 VL 72B InstructAlibaba Cloud / Qwen Team74.6%0.0%8CAug 7, 2026

Score distribution

Top published rows on the benchmark's original scale.

MLVU-M

MLVU-M highlights

The top published results on this benchmark's own scale.

Rank #1Qwen3 VL 32B Instruct82.1%Rank #2Qwen3 VL 30B A3B Instruct81.3%Rank #3Qwen3 VL 30B A3B Thinking78.9%Rank #4Qwen3 VL 8B Instruct78.1%

What is MLVU-M?

Definition and scoring fields from the benchmark registry.

MLVU-M benchmark

Scores are shown in ratio. The current registry marks this benchmark as not independently verified with evidence level B.

Family
MLVU-M
Modality
text
Primary category
general
Score direction
higher
LLMBoard eligible
No
Evaluation key
mlvu-m|llm-stats-current

Source-native results are preserved. Eligibility for the overall LLMBoard score is a separate policy decision.

FAQ

Common questions about MLVU-M.

Which model scores highest on MLVU-M?

Qwen3 VL 32B Instruct is currently ranked first with 82.1%.

What does MLVU-M measure?

MLVU-M benchmark

Is a higher score better?

Yes. Higher values rank better for this benchmark.

How many models are compared?

8 unique published model results are currently shown.

Does this benchmark affect the overall score?

This benchmark is preserved as source-native evidence but is not eligible for the current overall score.

Rankings

OverallCodingText ArenaPricing

Modalities

Image GenerationVideo GenerationSpeech-to-TextEmbeddings

Benchmarks

All BenchmarksReasoningMathCoding

Vendors

All VendorsOpenAIAnthropicGoogle
llmboard.aiCopyright 2026 llmboard.ai