llmboard.ai
Benchmarks
CompareRankings
llmboard.ai
Benchmarks
CompareRankings
HomeBenchmarksreasoningMBPP

reasoning benchmark

MBPP

MBPP (Mostly Basic Python Problems) is a benchmark of 974 crowd-sourced Python programming problems designed to be solvable by entry-level programmers. Each problem consists of a task description, code solution, and 3 automated test cases covering programming fundamentals and standard library functionality.

Updated Aug 7, 2026

Published models33
Registry coverage33
MetricScore
EvidenceB

On this page

  • Ranking
  • Distribution
  • Highlights
  • About
  • FAQ

MBPP leaderboard

Sorted by the source-provided rank. Higher score is better according to the registry.

33 rows
Columns

Show columns

1SASarvam-30BSarvam AI0.927 points100.0%37CAug 7, 2026
2NVLlama-3.3 Nemotron Super 49B v1NVIDIA0.913 points97.2%37CAug 7, 2026
3ACQwen2.5-Coder 32B InstructAlibaba Cloud / Qwen Team0.902 points94.4%37CAug 7, 2026
4OPMiniCPM-SALAOpenBMB0.891 points91.7%37CAug 7, 2026
5ACQwen2.5 72B InstructAlibaba Cloud / Qwen Team0.882 points88.9%37CAug 7, 2026
6NVLlama 3.1 Nemotron Nano 8B V1NVIDIA0.846 points86.1%37CAug 7, 2026
7ACQwen2.5 32B InstructAlibaba Cloud / Qwen Team0.84 points83.3%37CAug 7, 2026
8ACQwen2.5 VL 32B InstructAlibaba Cloud / Qwen Team0.84 points80.6%37CAug 7, 2026
9ACQwen2.5-Coder 7B InstructAlibaba Cloud / Qwen Team0.835 points77.8%37CAug 7, 2026
10ACQwen2.5 14B InstructAlibaba Cloud / Qwen Team0.82 points75.0%37CAug 7, 2026
11ACQwen3 235B A22BAlibaba Cloud / Qwen Team0.814 points72.2%37CAug 7, 2026
12MIPhi-3.5-MoE-instructMicrosoft0.808 points69.4%37CAug 7, 2026
13ACQwen2 72B InstructAlibaba Cloud / Qwen Team0.802 points66.7%37CAug 7, 2026
14ACQwen2.5 7B InstructAlibaba Cloud / Qwen Team0.792 points63.9%37CAug 7, 2026
15MACodestral-22BMistral AI0.782 points61.1%37CAug 7, 2026
16MELlama 4 MaverickMeta0.776 points58.3%37CAug 7, 2026
17GOGemini DiffusionGoogle0.76 points55.6%37CAug 7, 2026
18MAMistral Small 3.1 24B InstructMistral AI0.747 points52.8%37CAug 7, 2026
19GOGemma 3 27BGoogle0.744 points50.0%37CAug 7, 2026
21ACQwen2.5-Omni-7BAlibaba Cloud / Qwen Team0.732 points44.4%37CAug 7, 2026
22GOGemma 3 12BGoogle0.73 points41.7%37CAug 7, 2026
23MAMistral Small 3 24B BaseMistral AI0.696 points38.9%37CAug 7, 2026
24MIPhi-3.5-mini-instructMicrosoft0.696 points36.1%37CAug 7, 2026
25MELlama 4 ScoutMeta0.678 points33.3%37CAug 7, 2026
27ACQwen2 7B InstructAlibaba Cloud / Qwen Team0.672 points27.8%37CAug 7, 2026
28GOGemma 3n E4B InstructedGoogle0.636 points25.0%37CAug 7, 2026
29GOGemma 3n E4B Instructed LiteRT PreviewGoogle0.636 points22.2%37CAug 7, 2026
30GOGemma 3 4BGoogle0.632 points19.4%37CAug 7, 2026
32GOGemma 2 27BGoogle0.626 points13.9%37CAug 7, 2026
33GOGemma 3n E2B InstructedGoogle0.566 points11.1%37CAug 7, 2026
34GOGemma 3n E2B Instructed LiteRT (Preview)Google0.566 points8.3%37CAug 7, 2026
35GOGemma 2 9BGoogle0.524 points5.6%37CAug 7, 2026
37GOGemma 3 1BGoogle0.352 points0.0%37CAug 7, 2026

Score distribution

Top published rows on the benchmark's original scale.

MBPP

MBPP highlights

The top published results on this benchmark's own scale.

Rank #1Sarvam-30B0.927 pointsRank #2Llama-3.3 Nemotron Super 49B v10.913 pointsRank #3Qwen2.5-Coder 32B Instruct0.902 pointsRank #4MiniCPM-SALA0.891 points

What is MBPP?

Definition and scoring fields from the benchmark registry.

MBPP (Mostly Basic Python Problems) is a benchmark of 974 crowd-sourced Python programming problems designed to be solvable by entry-level programmers. Each problem consists of a task description, code solution, and 3 automated test cases covering programming fundamentals and standard library functionality.

Scores are shown in points. The current registry marks this benchmark as not independently verified with evidence level B.

Family
MBPP
Modality
text
Primary category
reasoning
Score direction
higher
LLMBoard eligible
Yes
Evaluation key
mbpp|llm-stats-current

Source-native results are preserved. Eligibility for the overall LLMBoard score is a separate policy decision.

FAQ

Common questions about MBPP.

Which model scores highest on MBPP?

Sarvam-30B is currently ranked first with 0.927 points.

What does MBPP measure?

MBPP (Mostly Basic Python Problems) is a benchmark of 974 crowd-sourced Python programming problems designed to be solvable by entry-level programmers. Each problem consists of a task description, code solution, and 3 automated test cases covering programming fundamentals and standard library functionality.

Is a higher score better?

Yes. Higher values rank better for this benchmark.

How many models are compared?

33 unique published model results are currently shown.

Does this benchmark affect the overall score?

This benchmark is marked as eligible for the current LLMBoard capability methodology.

Rankings

OverallCodingText ArenaPricing

Modalities

Image GenerationVideo GenerationSpeech-to-TextEmbeddings

Benchmarks

All BenchmarksReasoningMathCoding

Vendors

All VendorsOpenAIAnthropicGoogle
llmboard.aiCopyright 2026 llmboard.ai