Benchmark category
Published math evaluations and source-native model rankings. Each benchmark keeps its original scale and methodology.
Data as of 2026-08-07
Select a benchmark to inspect model-level results, evidence fields and scoring direction.
| MMLU-Pro | MMLU-Pro | text | Score | 129 | featured | B | Yes |
| AIME 2025 | AIME 2025 | text | Score | 114 | featured | C | Yes |
| MMLU | MMLU | text | Score | 100 | featured | B | Yes |
| Humanity's Last Exam | Humanity's Last Exam | multimodal | Score | 92 | featured | B | No |
| MATH | MATH | text | Score | 71 | featured | B | Yes |
| AIME 2024 | AIME 2024 | text | Score | 53 | featured | B | Yes |
| MMMLU | MMMLU | text | Score | 49 | featured | B | Yes |
| GSM8k | GSM8k | text | Score | 48 | featured | B | Yes |
| MMLU-Redux | MMLU-Redux | text | Score | 48 | featured | B | Yes |
| MathVista | MathVista | multimodal | Score | 39 | featured | B | No |
| LiveBench | LiveBench | text | Score | 38 | featured | B | No |
| SuperGPQA | SuperGPQA | text | Score | 34 | featured | B | Yes |
| HMMT 2025 | HMMT 2025 | text | Score | 33 | featured | B | Yes |
| MATH-500 | MATH-500 | text | Score | 32 | featured | C | Yes |
| MMLU-ProX | MMLU-ProX | text | Score | 32 | featured | B | Yes |
| MathVision | MathVision | multimodal | Score | 32 | featured | B | No |
| MGSM | MGSM | text | Score | 31 | featured | B | Yes |
| DROP | DROP | text | Score | 30 | featured | B | Yes |
| HMMT25 | HMMT25 | text | Score | 25 | featured | B | Yes |
| MathVista-Mini | MathVista-Mini | multimodal | Score | 23 | featured | B | No |
| PolyMATH | PolyMATH | multimodal | Score | 23 | featured | B | No |
| BIG-Bench Hard | BIG-Bench Hard | text | Score | 21 | featured | B | Yes |
| IMO-AnswerBench | IMO-AnswerBench | text | Score | 19 | featured | B | Yes |
| SciCode | SciCode | text | Score | 18 | featured | B | Yes |
| AIME 2026 | AIME 2026 | text | Score | 17 | featured | B | Yes |
| FrontierMath | FrontierMath | text | Score | 17 | featured | B | Yes |
| CodeForces | CodeForces | text | Score | 16 | featured | B | Yes |
| LiveBench 20241125 | LiveBench 20241125 | text | Score | 14 | featured | B | No |
| HiddenMath | HiddenMath | text | Score | 13 | featured | B | Yes |
| BBH | BBH | text | Score | 12 | featured | B | Yes |
| HMMT Feb 26 | HMMT Feb 26 | text | Score | 11 | featured | B | Yes |
| AGIEval | AGIEval | text | Score | 10 | featured | B | Yes |
High-coverage benchmarks with at least two published model results.
How this category is assembled on llmboard.ai.
This page groups benchmarks whose primary or display category matches math. It does not average incompatible metrics into a new category score.
Open an individual benchmark to inspect score direction, evidence level, participant count and source-native results.
Category membership is derived from the current benchmark registry response.