A more robust and challenging multi-task language understanding benchmark that extends MMLU by expanding multiple-choice options from 4 to 10, eliminating trivial questions, and focusing on reasoning-intensive tasks. Features over 12,000 curated questions across 14 domains and causes a 16-33% accuracy drop compared to original MMLU.
Updated Aug 7, 2026
Published models100
Registry coverage129
MetricScore
EvidenceB
MMLU-Pro leaderboard
Sorted by the source-provided rank. Higher score is better according to the registry.
Definition and scoring fields from the benchmark registry.
A more robust and challenging multi-task language understanding benchmark that extends MMLU by expanding multiple-choice options from 4 to 10, eliminating trivial questions, and focusing on reasoning-intensive tasks. Features over 12,000 curated questions across 14 domains and causes a 16-33% accuracy drop compared to original MMLU.
Scores are shown in ratio. The current registry marks this benchmark as not independently verified with evidence level B.
Family
MMLU-Pro
Modality
text
Primary category
math
Score direction
higher
LLMBoard eligible
Yes
Evaluation key
mmlu-pro|llm-stats-current
Source-native results are preserved. Eligibility for the overall LLMBoard score is a separate policy decision.
FAQ
Common questions about MMLU-Pro.
Which model scores highest on MMLU-Pro?
Qwen3.7 Max is currently ranked first with 89.6%.
What does MMLU-Pro measure?
A more robust and challenging multi-task language understanding benchmark that extends MMLU by expanding multiple-choice options from 4 to 10, eliminating trivial questions, and focusing on reasoning-intensive tasks. Features over 12,000 curated questions across 14 domains and causes a 16-33% accuracy drop compared to original MMLU.
Is a higher score better?
Yes. Higher values rank better for this benchmark.
How many models are compared?
100 unique published model results are currently shown.
Does this benchmark affect the overall score?
This benchmark is marked as eligible for the current LLMBoard capability methodology.