llmboard.ai
Benchmarks
CompareRankings
llmboard.ai
Benchmarks
CompareRankings
HomeBenchmarksmathGSM8k

math benchmark

GSM8k

Grade School Math 8K, a dataset of 8.5K high-quality linguistically diverse grade school math word problems requiring multi-step reasoning and elementary arithmetic operations.

Updated Aug 7, 2026

Published models48
Registry coverage48
MetricScore
EvidenceB

On this page

  • Ranking
  • Distribution
  • Highlights
  • About
  • FAQ

GSM8k leaderboard

Sorted by the source-provided rank. Higher score is better according to the registry.

48 rows
Columns

Show columns

1XIMiMo-V2.5-ProXiaomi99.6%100.0%48CAug 7, 2026
2MAKimi K2 InstructMoonshot AI97.3%97.9%48CAug 7, 2026
3OPo1OpenAI97.1%95.7%48CAug 7, 2026
4OPGPT-4.5OpenAI97.0%93.6%48CAug 7, 2026
5MELlama 3.1 405B InstructMeta96.8%91.5%48CAug 7, 2026
6ANClaude 3.5 SonnetAnthropic96.4%89.4%48CAug 7, 2026
7ANClaude 3.5 SonnetAnthropic96.4%87.2%48CAug 7, 2026
8GOGemma 3 27BGoogle95.9%85.1%48CAug 7, 2026
9ACQwen2.5 32B InstructAlibaba Cloud / Qwen Team95.9%83.0%48CAug 7, 2026
10ACQwen2.5 72B InstructAlibaba Cloud / Qwen Team95.8%80.8%48CAug 7, 2026
11DEDeepSeek-V2.5DeepSeek95.1%78.7%48CAug 7, 2026
12ANClaude 3 OpusAnthropic95.0%76.6%48CAug 7, 2026
13AMNova ProAmazon94.8%74.5%48CAug 7, 2026
14ACQwen2.5 14B InstructAlibaba Cloud / Qwen Team94.8%72.3%48CAug 7, 2026
15AMNova LiteAmazon94.5%70.2%48CAug 7, 2026
16GOGemma 3 12BGoogle94.4%68.1%48CAug 7, 2026
17ACQwen3 235B A22BAlibaba Cloud / Qwen Team94.4%66.0%48CAug 7, 2026
18MAMistral Large 2Mistral AI93.0%63.8%48CAug 7, 2026
19ANClaude 3 SonnetAnthropic92.3%61.7%48CAug 7, 2026
20AMNova MicroAmazon92.3%59.6%48CAug 7, 2026
21MAKimi K2 BaseMoonshot AI92.1%57.5%48CAug 7, 2026
22ACQwen2.5 7B InstructAlibaba Cloud / Qwen Team91.6%55.3%48CAug 7, 2026
23NVLlama 3.1 Nemotron 70B InstructNVIDIA91.4%53.2%48CAug 7, 2026
24ACQwen2.5-Coder 32B InstructAlibaba Cloud / Qwen Team91.1%51.1%48CAug 7, 2026
25ACQwen2 72B InstructAlibaba Cloud / Qwen Team91.1%48.9%48CAug 7, 2026
26GOGemini 1.5 ProGoogle90.8%46.8%48CAug 7, 2026
27XAGrok-1.5xAI90.0%44.7%48CAug 7, 2026
28GOGemma 3 4BGoogle89.2%42.5%48CAug 7, 2026
29ANClaude 3 HaikuAnthropic88.9%40.4%48CAug 7, 2026
30MIPhi-3.5-MoE-instructMicrosoft88.7%38.3%48CAug 7, 2026
31ACQwen2.5-Omni-7BAlibaba Cloud / Qwen Team88.7%36.2%48CAug 7, 2026
32MIPhi 4 MiniMicrosoft88.6%34.0%48CAug 7, 2026
33ALJamba 1.5 LargeAI21 Labs87.0%31.9%48CAug 7, 2026
34GOGemini 1.5 FlashGoogle86.2%29.8%48CAug 7, 2026
35MIPhi-3.5-mini-instructMicrosoft86.2%27.7%48CAug 7, 2026
36ACQwen2.5-Coder 7B InstructAlibaba Cloud / Qwen Team83.9%25.5%48CAug 7, 2026
37ACQwen2 7B InstructAlibaba Cloud / Qwen Team82.3%23.4%48CAug 7, 2026
38IBGranite 3.3 8B InstructIBM80.9%21.3%48CAug 7, 2026
39MAMistral Small 3 24B BaseMistral AI80.7%19.1%48CAug 7, 2026
40MELlama 3.2 3B InstructMeta77.7%17.0%48CAug 7, 2026
41ALJamba 1.5 MiniAI21 Labs75.8%14.9%48CAug 7, 2026
42GOGemma 2 27BGoogle74.0%12.8%48CAug 7, 2026
43COCommand R+Cohere70.7%10.6%48CAug 7, 2026
44IBIBM Granite 4.0 Tiny PreviewIBM70.1%8.5%48CAug 7, 2026
45GOGemma 2 9BGoogle68.6%6.4%48CAug 7, 2026
46GOGemma 3 1BGoogle62.8%4.3%48CAug 7, 2026
47IBGranite 3.3 8B BaseIBM59.0%2.1%48CAug 7, 2026
48BAERNIE 4.5Baidu25.2%0.0%48CAug 7, 2026

Score distribution

Top published rows on the benchmark's original scale.

GSM8k

GSM8k highlights

The top published results on this benchmark's own scale.

Rank #1MiMo-V2.5-Pro99.6%Rank #2Kimi K2 Instruct97.3%Rank #3o197.1%Rank #4GPT-4.597.0%

What is GSM8k?

Definition and scoring fields from the benchmark registry.

Grade School Math 8K, a dataset of 8.5K high-quality linguistically diverse grade school math word problems requiring multi-step reasoning and elementary arithmetic operations.

Scores are shown in ratio. The current registry marks this benchmark as not independently verified with evidence level B.

Family
GSM8k
Modality
text
Primary category
math
Score direction
higher
LLMBoard eligible
Yes
Evaluation key
gsm8k|llm-stats-current

Source-native results are preserved. Eligibility for the overall LLMBoard score is a separate policy decision.

FAQ

Common questions about GSM8k.

Which model scores highest on GSM8k?

MiMo-V2.5-Pro is currently ranked first with 99.6%.

What does GSM8k measure?

Grade School Math 8K, a dataset of 8.5K high-quality linguistically diverse grade school math word problems requiring multi-step reasoning and elementary arithmetic operations.

Is a higher score better?

Yes. Higher values rank better for this benchmark.

How many models are compared?

48 unique published model results are currently shown.

Does this benchmark affect the overall score?

This benchmark is marked as eligible for the current LLMBoard capability methodology.

Rankings

OverallCodingText ArenaPricing

Modalities

Image GenerationVideo GenerationSpeech-to-TextEmbeddings

Benchmarks

All BenchmarksReasoningMathCoding

Vendors

All VendorsOpenAIAnthropicGoogle
llmboard.aiCopyright 2026 llmboard.ai