llmboard.ai
Benchmarks
CompareRankings
llmboard.ai
Benchmarks
CompareRankings
HomeBenchmarksmathAIME 2025

math benchmark

AIME 2025

All 30 problems from the 2025 American Invitational Mathematics Examination (AIME I and AIME II), testing olympiad-level mathematical reasoning with integer answers from 000-999. Used as an AI benchmark to evaluate large language models' ability to solve complex mathematical problems requiring multi-step logical deductions and structured symbolic reasoning.

Updated Aug 7, 2026

Published models100
Registry coverage114
MetricScore
EvidenceC

On this page

  • Ranking
  • Distribution
  • Highlights
  • About
  • FAQ

AIME 2025 leaderboard

Sorted by the source-provided rank. Higher score is better according to the registry.

100 rows
Columns

Show columns

1GOGemini 3 ProGoogle100.0%100.0%114CAug 7, 2026
2OPGPT-5.2OpenAI100.0%99.1%114CAug 7, 2026
3OPGPT-5.2 ProOpenAI100.0%98.2%114CAug 7, 2026
4XAGrok-4 HeavyxAI100.0%97.3%114CAug 7, 2026
5MAKimi K2-Thinking-0905Moonshot AI100.0%96.5%114CAug 7, 2026
6ANClaude Opus 4.6Anthropic99.8%95.6%114CAug 7, 2026
7GOGemini 3 FlashGoogle99.7%94.7%114CAug 7, 2026
8OPGPT-5.1 HighOpenAI99.6%93.8%114CAug 7, 2026
9MELongCat-Flash-Thinking-2601Meituan99.6%92.9%114CAug 7, 2026
10NVNemotron 3 Nano (30B A3B)NVIDIA99.2%92.0%114CAug 7, 2026
11OPGPT OSS 20B HighOpenAI98.7%91.2%114CAug 7, 2026
12OPGPT-5.1 MediumOpenAI98.4%90.3%114CAug 7, 2026
13BYSeed 2.0 ProByteDance98.3%89.4%114CAug 7, 2026
14STStep-3.5-FlashStepFun97.3%88.5%114CAug 7, 2026
15MIMAI-Thinking-1Microsoft97.0%87.6%114CAug 7, 2026
16OPGPT-5.1 Codex HighOpenAI96.7%86.7%114CAug 7, 2026
17SASarvam-105BSarvam AI96.7%85.8%114CAug 7, 2026
18SASarvam-30BSarvam AI96.7%85.0%114CAug 7, 2026
19MAKimi K2.5Moonshot AI96.1%84.1%114CAug 7, 2026
20DEDeepSeek-V3.2-SpecialeDeepSeek96.0%83.2%114CAug 7, 2026
21ZAGLM-4.7Zhipu AI95.7%82.3%114CAug 7, 2026
22OPGPT-5OpenAI94.6%81.4%114CAug 7, 2026
23OPGPT-5 HighOpenAI94.6%80.5%114CAug 7, 2026
24XIMiMo-V2-FlashXiaomi94.1%79.7%114CAug 7, 2026
25OPGPT-5.1OpenAI94.0%78.8%114CAug 7, 2026
26OPGPT-5.1 InstantOpenAI94.0%77.9%114CAug 7, 2026
27OPGPT-5.1 ThinkingOpenAI94.0%77.0%114CAug 7, 2026
28ZAGLM-4.6Zhipu AI93.9%76.1%114CAug 7, 2026
29XAGrok-3xAI93.3%75.2%114CAug 7, 2026
30DEDeepSeek-V3.2 (Thinking)DeepSeek93.1%74.3%114CAug 7, 2026
31DEDeepSeek-V3.2DeepSeek93.1%73.5%114CAug 7, 2026
32BYSeed 2.0 LiteByteDance93.0%72.6%114CAug 7, 2026
33LAK-EXAONE-236B-A23BLG AI Research92.8%71.7%114CAug 7, 2026
34OPo4-miniOpenAI92.7%70.8%114CAug 7, 2026
35OPGPT OSS 120B HighOpenAI92.5%69.9%114CAug 7, 2026
36AMNova 2 ProAmazon92.3%69.0%114CAug 7, 2026
37ACQwen3-235B-A22B-Thinking-2507Alibaba Cloud / Qwen Team92.3%68.1%114CAug 7, 2026
38AMNova 2 OmniAmazon92.1%67.3%114CAug 7, 2026
39XAGrok 4 FastxAI92.0%66.4%114CAug 7, 2026
40XAGrok-4xAI91.7%65.5%114CAug 7, 2026
41ZAGLM-4.7-FlashZhipu AI91.6%64.6%114CAug 7, 2026
42OPGPT-5 miniOpenAI91.1%63.7%114CAug 7, 2026
43INMercury 2Inception91.1%62.8%114CAug 7, 2026
44AMNova 2 LiteAmazon91.0%62.0%114CAug 7, 2026
45XAGrok-3 MinixAI90.8%61.1%114CAug 7, 2026
46MELongCat-Flash-ThinkingMeituan90.6%60.2%114CAug 7, 2026
47NVNemotron 3 Super (120B A12B)NVIDIA90.2%59.3%114CAug 7, 2026
48COCommand A+Cohere90.0%58.4%114CAug 7, 2026
49ACQwen3 VL 235B A22B ThinkingAlibaba Cloud / Qwen Team89.7%57.5%114CAug 7, 2026
50DEDeepSeek-V3.2-ExpDeepSeek89.3%56.6%114CAug 7, 2026
51OPGPT-5 MediumOpenAI88.9%55.8%114CAug 7, 2026
52GOGemini 2.5 Pro Preview 06-05Google88.0%54.9%114CAug 7, 2026
53ACQwen3-Next-80B-A3B-ThinkingAlibaba Cloud / Qwen Team87.8%54.0%114CAug 7, 2026
54STStep3-VL-10BStepFun87.7%53.1%114CAug 7, 2026
55DEDeepSeek-R1-0528DeepSeek87.5%52.2%114CAug 7, 2026
56ANClaude Sonnet 4.5Anthropic87.0%51.3%114CAug 7, 2026
57BAERNIE 5.0Baidu87.0%50.4%114CAug 7, 2026
58OPo3OpenAI86.4%49.6%114CAug 7, 2026
59MAMistral Medium 3.5Mistral AI86.3%48.7%114CAug 7, 2026
60OPGPT-5 nanoOpenAI85.2%47.8%114CAug 7, 2026
61MAMinistral 3 (14B Reasoning 2512)Mistral AI85.0%46.9%114CAug 7, 2026
62MAMistral Small 4Mistral AI83.8%46.0%114CAug 7, 2026
63ACQwen3 VL 32B ThinkingAlibaba Cloud / Qwen Team83.7%45.1%114CAug 7, 2026
64ACQwen3 VL 30B A3B ThinkingAlibaba Cloud / Qwen Team83.1%44.3%114CAug 7, 2026
65GOGemini 2.5 ProGoogle83.0%43.4%114CAug 7, 2026
66ACQwen3 MaxAlibaba Cloud / Qwen Team81.6%42.5%114CAug 7, 2026
67ACQwen3 235B A22BAlibaba Cloud / Qwen Team81.5%41.6%114CAug 7, 2026
68OPGPT-5.5 InstantOpenAI81.2%40.7%114CAug 7, 2026
69MIMiniMax M2.1MiniMax81.0%39.8%114CAug 7, 2026
70ANClaude Haiku 4.5Anthropic80.7%38.9%114CAug 7, 2026
71ACQwen3 VL 8B ThinkingAlibaba Cloud / Qwen Team80.3%38.0%114CAug 7, 2026
72MAMinistral 3 (8B Reasoning 2512)Mistral AI78.7%37.2%114CAug 7, 2026
73OPMiniCPM-SALAOpenBMB78.3%36.3%114CAug 7, 2026
74ANClaude Opus 4.1Anthropic78.0%35.4%114CAug 7, 2026
75MIMiniMax M2MiniMax78.0%34.5%114CAug 7, 2026
76MIPhi 4 Reasoning PlusMicrosoft78.0%33.6%114CAug 7, 2026
77MIMiniMax M1 80KMiniMax76.9%32.7%114CAug 7, 2026
78ANClaude Opus 4Anthropic75.5%31.9%114CAug 7, 2026
79ACQwen3 VL 235B A22B InstructAlibaba Cloud / Qwen Team74.7%31.0%114CAug 7, 2026
80MIMiniMax M1 40KMiniMax74.6%30.1%114CAug 7, 2026
81ACQwen3 VL 4B ThinkingAlibaba Cloud / Qwen Team74.5%29.2%114CAug 7, 2026
82ACQwen3 32BAlibaba Cloud / Qwen Team72.9%28.3%114CAug 7, 2026
83NVLlama 3.1 Nemotron Ultra 253B v1NVIDIA72.5%27.4%114CAug 7, 2026
84MAMin istral 3 (3B Reasoning 2512)Mistral AI72.1%26.6%114CAug 7, 2026
85NVNemotron Nano 9B v2NVIDIA72.1%25.7%114CAug 7, 2026
86GOGemini 2.5 FlashGoogle72.0%24.8%114CAug 7, 2026
87ACQwen3 30B A3BAlibaba Cloud / Qwen Team70.9%23.9%114CAug 7, 2026
88ANClaude Sonnet 4Anthropic70.5%23.0%114CAug 7, 2026
89ACQwen3-235B-A22B-Instruct-2507Alibaba Cloud / Qwen Team70.3%22.1%114CAug 7, 2026
90ACQwen3-Next-80B-A3B-InstructAlibaba Cloud / Qwen Team69.5%21.2%114CAug 7, 2026
91ACQwen3 VL 30B A3B InstructAlibaba Cloud / Qwen Team69.3%20.4%114CAug 7, 2026
92ACQwen3 VL 32B InstructAlibaba Cloud / Qwen Team66.2%19.5%114CAug 7, 2026
93MAMagistral MediumMistral AI64.9%18.6%114CAug 7, 2026
94MELongCat-Flash-LiteMeituan63.2%17.7%114CAug 7, 2026
95MIPhi 4 ReasoningMicrosoft62.9%16.8%114CAug 7, 2026
96MAMagistral Small 2506Mistral AI62.8%15.9%114CAug 7, 2026
97MELongCat-Flash-ChatMeituan61.3%15.0%114CAug 7, 2026
98NVLlama-3.3 Nemotron Super 49B v1NVIDIA58.4%14.2%114CAug 7, 2026
99ANClaude 3.7 SonnetAnthropic54.8%13.3%114CAug 7, 2026
100DEDeepSeek-V3.1DeepSeek49.8%12.4%114CAug 7, 2026

Score distribution

Top published rows on the benchmark's original scale.

AIME 2025

AIME 2025 highlights

The top published results on this benchmark's own scale.

Rank #1Gemini 3 Pro100.0%Rank #2GPT-5.2100.0%Rank #3GPT-5.2 Pro100.0%Rank #4Grok-4 Heavy100.0%

What is AIME 2025?

Definition and scoring fields from the benchmark registry.

All 30 problems from the 2025 American Invitational Mathematics Examination (AIME I and AIME II), testing olympiad-level mathematical reasoning with integer answers from 000-999. Used as an AI benchmark to evaluate large language models' ability to solve complex mathematical problems requiring multi-step logical deductions and structured symbolic reasoning.

Scores are shown in ratio. The current registry marks this benchmark as not independently verified with evidence level C.

Family
AIME 2025
Modality
text
Primary category
math
Score direction
higher
LLMBoard eligible
Yes
Evaluation key
aime-2025|llm-stats-current

Source-native results are preserved. Eligibility for the overall LLMBoard score is a separate policy decision.

FAQ

Common questions about AIME 2025.

Which model scores highest on AIME 2025?

Gemini 3 Pro is currently ranked first with 100.0%.

What does AIME 2025 measure?

All 30 problems from the 2025 American Invitational Mathematics Examination (AIME I and AIME II), testing olympiad-level mathematical reasoning with integer answers from 000-999. Used as an AI benchmark to evaluate large language models' ability to solve complex mathematical problems requiring multi-step logical deductions and structured symbolic reasoning.

Is a higher score better?

Yes. Higher values rank better for this benchmark.

How many models are compared?

100 unique published model results are currently shown.

Does this benchmark affect the overall score?

This benchmark is marked as eligible for the current LLMBoard capability methodology.

Rankings

OverallCodingText ArenaPricing

Modalities

Image GenerationVideo GenerationSpeech-to-TextEmbeddings

Benchmarks

All BenchmarksReasoningMathCoding

Vendors

All VendorsOpenAIAnthropicGoogle
llmboard.aiCopyright 2026 llmboard.ai