llmboard.ai
Benchmarks
CompareRankings
llmboard.ai
Benchmarks
CompareRankings
HomeBenchmarksmathMATH-500

math benchmark

MATH-500

MATH-500 is a subset of the MATH dataset containing 500 challenging competition mathematics problems from AMC 10, AMC 12, AIME, and other mathematics competitions. Each problem includes full step-by-step solutions and spans multiple difficulty levels across seven mathematical subjects including Prealgebra, Algebra, Number Theory, Counting and Probability, Geometry, Intermediate Algebra, and Precalculus.

Updated Aug 7, 2026

Published models32
Registry coverage32
MetricScore
EvidenceC

On this page

  • Ranking
  • Distribution
  • Highlights
  • About
  • FAQ

MATH-500 leaderboard

Sorted by the source-provided rank. Higher score is better according to the registry.

32 rows
Columns

Show columns

1MELongCat-Flash-ThinkingMeituan99.2%100.0%32CAug 7, 2026
2SASarvam-105BSarvam AI98.6%96.8%32CAug 7, 2026
3ZAGLM-4.5Zhipu AI98.2%93.5%32CAug 7, 2026
4ZAGLM-4.5-AirZhipu AI98.1%90.3%32CAug 7, 2026
5NVNemotron Nano 9B v2NVIDIA97.8%87.1%32CAug 7, 2026
6MAKimi K2 InstructMoonshot AI97.4%83.9%32CAug 7, 2026
7MAKimi K2-Instruct-0905Moonshot AI97.4%80.7%32CAug 7, 2026
8NVLlama 3.1 Nemotron Ultra 253B v1NVIDIA97.0%77.4%32CAug 7, 2026
9SASarvam-30BSarvam AI97.0%74.2%32CAug 7, 2026
10MELongCat-Flash-LiteMeituan96.8%71.0%32CAug 7, 2026
11MIMiniMax M1 80KMiniMax96.8%67.7%32CAug 7, 2026
12NVLlama-3.3 Nemotron Super 49B v1NVIDIA96.6%64.5%32CAug 7, 2026
13MELongCat-Flash-ChatMeituan96.4%61.3%32CAug 7, 2026
14ANClaude 3.7 SonnetAnthropic96.2%58.1%32CAug 7, 2026
15MAKimi-k1.5Moonshot AI96.2%54.8%32CAug 7, 2026
16MIMiniMax M1 40KMiniMax96.0%51.6%32CAug 7, 2026
17DEDeepSeek R1 ZeroDeepSeek95.9%48.4%32CAug 7, 2026
18NVLlama 3.1 Nemotron Nano 8B V1NVIDIA95.4%45.2%32CAug 7, 2026
19MIPhi 4 Mini ReasoningMicrosoft94.6%41.9%32CAug 7, 2026
20DEDeepSeek R1 Distill Llama 70BDeepSeek94.5%38.7%32CAug 7, 2026
21DEDeepSeek R1 Distill Qwen 32BDeepSeek94.3%35.5%32CAug 7, 2026
22DEDeepSeek-V3 0324DeepSeek94.0%32.3%32CAug 7, 2026
23DEDeepSeek R1 Distill Qwen 14BDeepSeek93.9%29.0%32CAug 7, 2026
24DEDeepSeek R1 Distill Qwen 7BDeepSeek92.8%25.8%32CAug 7, 2026
25ACQwQ-32BAlibaba Cloud / Qwen Team90.6%22.6%32CAug 7, 2026
26ACQwQ-32B-PreviewAlibaba Cloud / Qwen Team90.6%19.4%32CAug 7, 2026
27DEDeepSeek-V3DeepSeek90.2%16.1%32CAug 7, 2026
28OPo1-miniOpenAI90.0%12.9%32CAug 7, 2026
29DEDeepSeek R1 Distill Llama 8BDeepSeek89.1%9.7%32CAug 7, 2026
30DEDeepSeek R1 Distill Qwen 1.5BDeepSeek83.9%6.5%32CAug 7, 2026
31IBGranite 3.3 8B BaseIBM69.0%3.2%32CAug 7, 2026
32IBGranite 3.3 8B InstructIBM69.0%0.0%32CAug 7, 2026

Score distribution

Top published rows on the benchmark's original scale.

MATH-500

MATH-500 highlights

The top published results on this benchmark's own scale.

Rank #1LongCat-Flash-Thinking99.2%Rank #2Sarvam-105B98.6%Rank #3GLM-4.598.2%Rank #4GLM-4.5-Air98.1%

What is MATH-500?

Definition and scoring fields from the benchmark registry.

MATH-500 is a subset of the MATH dataset containing 500 challenging competition mathematics problems from AMC 10, AMC 12, AIME, and other mathematics competitions. Each problem includes full step-by-step solutions and spans multiple difficulty levels across seven mathematical subjects including Prealgebra, Algebra, Number Theory, Counting and Probability, Geometry, Intermediate Algebra, and Precalculus.

Scores are shown in ratio. The current registry marks this benchmark as not independently verified with evidence level C.

Family
MATH-500
Modality
text
Primary category
math
Score direction
higher
LLMBoard eligible
Yes
Evaluation key
math-500|llm-stats-current

Source-native results are preserved. Eligibility for the overall LLMBoard score is a separate policy decision.

FAQ

Common questions about MATH-500.

Which model scores highest on MATH-500?

LongCat-Flash-Thinking is currently ranked first with 99.2%.

What does MATH-500 measure?

MATH-500 is a subset of the MATH dataset containing 500 challenging competition mathematics problems from AMC 10, AMC 12, AIME, and other mathematics competitions. Each problem includes full step-by-step solutions and spans multiple difficulty levels across seven mathematical subjects including Prealgebra, Algebra, Number Theory, Counting and Probability, Geometry, Intermediate Algebra, and Precalculus.

Is a higher score better?

Yes. Higher values rank better for this benchmark.

How many models are compared?

32 unique published model results are currently shown.

Does this benchmark affect the overall score?

This benchmark is marked as eligible for the current LLMBoard capability methodology.

Rankings

OverallCodingText ArenaPricing

Modalities

Image GenerationVideo GenerationSpeech-to-TextEmbeddings

Benchmarks

All BenchmarksReasoningMathCoding

Vendors

All VendorsOpenAIAnthropicGoogle
llmboard.aiCopyright 2026 llmboard.ai