llmboard.ai
Benchmarks
CompareRankings
llmboard.ai
Benchmarks
CompareRankings
HomeBenchmarksreasoningMBPP

reasoning benchmark

MBPP

MBPP (Mostly Basic Python Problems) is a benchmark of 974 crowd-sourced Python programming problems designed to be solvable by entry-level programmers. Each problem consists of a task description, code solution, and 3 automated test cases covering programming fundamentals and standard library functionality.

Updated Aug 11, 2026

Models33
Model coverage33
MetricScore
EvidenceB

On this page

  • Ranking
  • Distribution
  • Highlights
  • About
  • FAQ

MBPP Ranking

Higher score ranks better on this benchmark.

33 rows
Columns

Show columns

1SASarvam-30BSarvam AI0.927 points100.0%33CAug 11, 2026
2NVLlama-3.3 Nemotron Super 49B v1NVIDIA0.913 points96.9%33CAug 11, 2026
3ACQwen2.5-Coder 32B InstructAlibaba Cloud / Qwen Team0.902 points93.8%33CAug 11, 2026
4OPMiniCPM-SALAOpenBMB0.891 points90.6%33CAug 11, 2026
5ACQwen2.5 72B InstructAlibaba Cloud / Qwen Team0.882 points87.5%33CAug 11, 2026
6NVLlama 3.1 Nemotron Nano 8B V1NVIDIA0.846 points84.4%33CAug 11, 2026
7ACQwen2.5 32B InstructAlibaba Cloud / Qwen Team0.84 points81.3%33CAug 11, 2026
8ACQwen2.5 VL 32B InstructAlibaba Cloud / Qwen Team0.84 points78.1%33CAug 11, 2026
9ACQwen2.5-Coder 7B InstructAlibaba Cloud / Qwen Team0.835 points75.0%33CAug 11, 2026
10ACQwen2.5 14B InstructAlibaba Cloud / Qwen Team0.82 points71.9%33CAug 11, 2026
11ACQwen3 235B A22BAlibaba Cloud / Qwen Team0.814 points68.8%33CAug 11, 2026
12MIPhi-3.5-MoE-instructMicrosoft0.808 points65.6%33CAug 11, 2026
13ACQwen2 72B InstructAlibaba Cloud / Qwen Team0.802 points62.5%33CAug 11, 2026
14ACQwen2.5 7B InstructAlibaba Cloud / Qwen Team0.792 points59.4%33CAug 11, 2026
15MACodestral-22BMistral AI0.782 points56.3%33CAug 11, 2026
16MELlama 4 MaverickMeta0.776 points53.1%33CAug 11, 2026
17GOGemini DiffusionGoogle0.76 points50.0%33CAug 11, 2026
18MAMistral Small 3.1 24B InstructMistral AI0.747 points46.9%33CAug 11, 2026
19GOGemma 3 27BGoogle0.744 points43.8%33CAug 11, 2026
20ACQwen2.5-Omni-7BAlibaba Cloud / Qwen Team0.732 points40.6%33CAug 11, 2026
21GOGemma 3 12BGoogle0.73 points37.5%33CAug 11, 2026
22MAMistral Small 3 24B BaseMistral AI0.696 points34.4%33CAug 11, 2026
23MIPhi-3.5-mini-instructMicrosoft0.696 points31.3%33CAug 11, 2026
24MELlama 4 ScoutMeta0.678 points28.1%33CAug 11, 2026
25ACQwen2 7B InstructAlibaba Cloud / Qwen Team0.672 points25.0%33CAug 11, 2026
26GOGemma 3n E4B InstructedGoogle0.636 points21.9%33CAug 11, 2026
27GOGemma 3n E4B Instructed LiteRT PreviewGoogle0.636 points18.8%33CAug 11, 2026
28GOGemma 3 4BGoogle0.632 points15.6%33CAug 11, 2026
29GOGemma 2 27BGoogle0.626 points12.5%33CAug 11, 2026
30GOGemma 3n E2B InstructedGoogle0.566 points9.4%33CAug 11, 2026
31GOGemma 3n E2B Instructed LiteRT (Preview)Google0.566 points6.3%33CAug 11, 2026
32GOGemma 2 9BGoogle0.524 points3.1%33CAug 11, 2026
33GOGemma 3 1BGoogle0.352 points0.0%33CAug 11, 2026

MBPP Score Distribution

A closer view of the leading scores on this benchmark.

MBPP

MBPP Highlights

The leading models and scores on this benchmark.

Rank #1Sarvam-30B0.927 pointsRank #2Llama-3.3 Nemotron Super 49B v10.913 pointsRank #3Qwen2.5-Coder 32B Instruct0.902 pointsRank #4MiniCPM-SALA0.891 points

What is MBPP?

What MBPP measures and how its scores work.

MBPP (Mostly Basic Python Problems) is a benchmark of 974 crowd-sourced Python programming problems designed to be solvable by entry-level programmers. Each problem consists of a task description, code solution, and 3 automated test cases covering programming fundamentals and standard library functionality.

Scores are shown in points. This benchmark is not independently verified and has an evidence level of B.

Family
MBPP
Modality
text
Primary category
reasoning
Score direction
higher
LLMBoard eligible
Yes
Evaluation key
mbpp|llm-stats-current

Benchmark scores retain their original unit. Overall score eligibility is shown separately.

FAQ

Common questions about MBPP.

Which model scores highest on MBPP?

Sarvam-30B is currently ranked first with 0.927 points.

What does MBPP measure?

MBPP (Mostly Basic Python Problems) is a benchmark of 974 crowd-sourced Python programming problems designed to be solvable by entry-level programmers. Each problem consists of a task description, code solution, and 3 automated test cases covering programming fundamentals and standard library functionality.

Is a higher score better?

Yes. Higher values rank better for this benchmark.

How many models are compared?

33 model results are currently shown.

Does this benchmark affect the overall score?

Yes. This benchmark can contribute to the current LLMBoard capability score.

Rankings

OverallCodingText ArenaPricing

Modalities

All ModelsImage GenerationImage EditingVideo GenerationImage-to-VideoVideo EditingText-to-SpeechSpeech-to-TextEmbeddings

Benchmarks

All BenchmarksReasoningMathCoding

Vendors

All VendorsOpenAIAnthropicGoogle
llmboard.aiCopyright 2026 llmboard.ai