llmboard.ai
Benchmarks
CompareRankings
llmboard.ai
Benchmarks
CompareRankings
HomeBenchmarksMath

Benchmark category

Math Benchmarks

Published math evaluations and source-native model rankings. Each benchmark keeps its original scale and methodology.

Data as of 2026-08-07

Benchmarks84
Highest coverage129
Score eligible25

Math benchmark registry

Select a benchmark to inspect model-level results, evidence fields and scoring direction.

32 rows
Columns

Show columns

MMLU-ProMMLU-ProtextScore129featuredBYes
AIME 2025AIME 2025textScore114featuredCYes
MMLUMMLUtextScore100featuredBYes
Humanity's Last ExamHumanity's Last ExammultimodalScore92featuredBNo
MATHMATHtextScore71featuredBYes
AIME 2024AIME 2024textScore53featuredBYes
MMMLUMMMLUtextScore49featuredBYes
GSM8kGSM8ktextScore48featuredBYes
MMLU-ReduxMMLU-ReduxtextScore48featuredBYes
MathVistaMathVistamultimodalScore39featuredBNo
LiveBenchLiveBenchtextScore38featuredBNo
SuperGPQASuperGPQAtextScore34featuredBYes
HMMT 2025HMMT 2025textScore33featuredBYes
MATH-500MATH-500textScore32featuredCYes
MMLU-ProXMMLU-ProXtextScore32featuredBYes
MathVisionMathVisionmultimodalScore32featuredBNo
MGSMMGSMtextScore31featuredBYes
DROPDROPtextScore30featuredBYes
HMMT25HMMT25textScore25featuredBYes
MathVista-MiniMathVista-MinimultimodalScore23featuredBNo
PolyMATHPolyMATHmultimodalScore23featuredBNo
BIG-Bench HardBIG-Bench HardtextScore21featuredBYes
IMO-AnswerBenchIMO-AnswerBenchtextScore19featuredBYes
SciCodeSciCodetextScore18featuredBYes
AIME 2026AIME 2026textScore17featuredBYes
FrontierMathFrontierMathtextScore17featuredBYes
CodeForcesCodeForcestextScore16featuredBYes
LiveBench 20241125LiveBench 20241125textScore14featuredBNo
HiddenMathHiddenMathtextScore13featuredBYes
BBHBBHtextScore12featuredBYes
HMMT Feb 26HMMT Feb 26textScore11featuredBYes
AGIEvalAGIEvaltextScore10featuredBYes

Top math result sets

High-coverage benchmarks with at least two published model results.

MMLU-Pro

View benchmark

AIME 2025

View benchmark

MMLU

View benchmark

MATH

View benchmark

What are math benchmarks?

How this category is assembled on llmboard.ai.

This page groups benchmarks whose primary or display category matches math. It does not average incompatible metrics into a new category score.

Open an individual benchmark to inspect score direction, evidence level, participant count and source-native results.

Category membership is derived from the current benchmark registry response.

Rankings

OverallCodingText ArenaPricing

Modalities

Image GenerationVideo GenerationSpeech-to-TextEmbeddings

Benchmarks

All BenchmarksReasoningMathCoding

Vendors

All VendorsOpenAIAnthropicGoogle
llmboard.aiCopyright 2026 llmboard.ai