llmboard.ai
Benchmarks
CompareRankings
llmboard.ai
Benchmarks
CompareRankings
HomeBenchmarksmathMGSM

math benchmark

MGSM

MGSM (Multilingual Grade School Math) is a benchmark of grade-school math problems. Contains 250 grade-school math problems manually translated from the GSM8K dataset into ten typologically diverse languages: Spanish, French, German, Russian, Chinese, Japanese, Thai, Swahili, Bengali, and Telugu. Evaluates multilingual mathematical reasoning capabilities.

Updated Aug 7, 2026

Published models31
Registry coverage31
MetricScore
EvidenceB

On this page

  • Ranking
  • Distribution
  • Highlights
  • About
  • FAQ

MGSM leaderboard

Sorted by the source-provided rank. Higher score is better according to the registry.

31 rows
Columns

Show columns

1MELlama 4 MaverickMeta92.3%100.0%31CAug 7, 2026
2OPo3-miniOpenAI92.0%96.7%31CAug 7, 2026
3ANClaude 3.5 SonnetAnthropic91.6%93.3%31CAug 7, 2026
4ANClaude 3.5 SonnetAnthropic91.6%90.0%31CAug 7, 2026
5MELlama 3.3 70B InstructMeta91.1%86.7%31CAug 7, 2026
6OPo1-previewOpenAI90.8%83.3%31CAug 7, 2026
7ANClaude 3 OpusAnthropic90.7%80.0%31CAug 7, 2026
8MELlama 4 ScoutMeta90.6%76.7%31CAug 7, 2026
9OPGPT-4oOpenAI90.5%73.3%31CAug 7, 2026
10OPo1OpenAI89.3%70.0%31CAug 7, 2026
11OPGPT-4 TurboOpenAI88.5%66.7%31CAug 7, 2026
12GOGemini 1.5 ProGoogle87.5%63.3%31CAug 7, 2026
13OPGPT-4o miniOpenAI87.0%60.0%31CAug 7, 2026
14MELlama 3.2 90B InstructMeta86.9%56.7%31CAug 7, 2026
15ANClaude 3.5 HaikuAnthropic85.6%53.3%31CAug 7, 2026
16ACQwen3 235B A22BAlibaba Cloud / Qwen Team83.5%50.0%31CAug 7, 2026
17ANClaude 3 SonnetAnthropic83.5%46.7%31CAug 7, 2026
18GOGemini 1.5 FlashGoogle82.6%43.3%31CAug 7, 2026
19MIPhi 4Microsoft80.6%40.0%31CAug 7, 2026
20ANClaude 3 HaikuAnthropic75.1%36.7%31CAug 7, 2026
21OPGPT-4OpenAI74.5%33.3%31CAug 7, 2026
22MELlama 3.2 11B InstructMeta68.9%30.0%31CAug 7, 2026
23GOGemma 3n E4B InstructedGoogle67.0%26.7%31CAug 7, 2026
24MIPhi 4 MiniMicrosoft63.9%23.3%31CAug 7, 2026
25GOGemma 3n E4B Instructed LiteRT PreviewGoogle60.7%20.0%31CAug 7, 2026
26MIPhi-3.5-MoE-instructMicrosoft58.7%16.7%31CAug 7, 2026
27MELlama 3.2 3B InstructMeta58.2%13.3%31CAug 7, 2026
28OPGPT-3.5 TurboOpenAI56.3%10.0%31BAug 7, 2026
29GOGemma 3n E2B InstructedGoogle53.1%6.7%31CAug 7, 2026
30GOGemma 3n E2B Instructed LiteRT (Preview)Google53.1%3.3%31CAug 7, 2026
31MIPhi-3.5-mini-instructMicrosoft47.9%0.0%31CAug 7, 2026

Score distribution

Top published rows on the benchmark's original scale.

MGSM

MGSM highlights

The top published results on this benchmark's own scale.

Rank #1Llama 4 Maverick92.3%Rank #2o3-mini92.0%Rank #3Claude 3.5 Sonnet91.6%Rank #4Claude 3.5 Sonnet91.6%

What is MGSM?

Definition and scoring fields from the benchmark registry.

MGSM (Multilingual Grade School Math) is a benchmark of grade-school math problems. Contains 250 grade-school math problems manually translated from the GSM8K dataset into ten typologically diverse languages: Spanish, French, German, Russian, Chinese, Japanese, Thai, Swahili, Bengali, and Telugu. Evaluates multilingual mathematical reasoning capabilities.

Scores are shown in ratio. The current registry marks this benchmark as not independently verified with evidence level B.

Family
MGSM
Modality
text
Primary category
math
Score direction
higher
LLMBoard eligible
Yes
Evaluation key
mgsm|llm-stats-current

Source-native results are preserved. Eligibility for the overall LLMBoard score is a separate policy decision.

FAQ

Common questions about MGSM.

Which model scores highest on MGSM?

Llama 4 Maverick is currently ranked first with 92.3%.

What does MGSM measure?

MGSM (Multilingual Grade School Math) is a benchmark of grade-school math problems. Contains 250 grade-school math problems manually translated from the GSM8K dataset into ten typologically diverse languages: Spanish, French, German, Russian, Chinese, Japanese, Thai, Swahili, Bengali, and Telugu. Evaluates multilingual mathematical reasoning capabilities.

Is a higher score better?

Yes. Higher values rank better for this benchmark.

How many models are compared?

31 unique published model results are currently shown.

Does this benchmark affect the overall score?

This benchmark is marked as eligible for the current LLMBoard capability methodology.

Rankings

OverallCodingText ArenaPricing

Modalities

Image GenerationVideo GenerationSpeech-to-TextEmbeddings

Benchmarks

All BenchmarksReasoningMathCoding

Vendors

All VendorsOpenAIAnthropicGoogle
llmboard.aiCopyright 2026 llmboard.ai