llmboard.ai
Benchmarks
CompareRankings
llmboard.ai
Benchmarks
CompareRankings
HomeBenchmarksreasoningMBPP+

reasoning benchmark

MBPP+

MBPP+ is an enhanced version of MBPP (Mostly Basic Python Problems) with significantly more test cases (35x) for more rigorous evaluation. MBPP is a benchmark of 974 crowd-sourced Python programming problems designed to be solvable by entry-level programmers, covering programming fundamentals and standard library functionality.

Updated Aug 7, 2026

Published models4
Registry coverage4
MetricScore
EvidenceB

On this page

  • Ranking
  • Distribution
  • Highlights
  • About
  • FAQ

MBPP+ leaderboard

Sorted by the source-provided rank. Higher score is better according to the registry.

4 rows
Columns

Show columns

20XIMiMo-V2.5-ProXiaomi74.1%47.2%37CAug 7, 2026
26ACQwen2.5 32B InstructAlibaba Cloud / Qwen Team67.2%30.6%37CAug 7, 2026
31ACQwen2.5 14B InstructAlibaba Cloud / Qwen Team63.2%16.7%37CAug 7, 2026
36BAERNIE 4.5Baidu40.2%2.8%37CAug 7, 2026

Score distribution

Top published rows on the benchmark's original scale.

MBPP+

MBPP+ highlights

The top published results on this benchmark's own scale.

Rank #20MiMo-V2.5-Pro74.1%Rank #26Qwen2.5 32B Instruct67.2%Rank #31Qwen2.5 14B Instruct63.2%Rank #36ERNIE 4.540.2%

What is MBPP+?

Definition and scoring fields from the benchmark registry.

MBPP+ is an enhanced version of MBPP (Mostly Basic Python Problems) with significantly more test cases (35x) for more rigorous evaluation. MBPP is a benchmark of 974 crowd-sourced Python programming problems designed to be solvable by entry-level programmers, covering programming fundamentals and standard library functionality.

Scores are shown in ratio. The current registry marks this benchmark as not independently verified with evidence level B.

Family
MBPP+
Modality
text
Primary category
reasoning
Score direction
higher
LLMBoard eligible
No
Evaluation key
mbpp+|llm-stats-current

Source-native results are preserved. Eligibility for the overall LLMBoard score is a separate policy decision.

FAQ

Common questions about MBPP+.

Which model scores highest on MBPP+?

MiMo-V2.5-Pro is currently ranked first with 74.1%.

What does MBPP+ measure?

MBPP+ is an enhanced version of MBPP (Mostly Basic Python Problems) with significantly more test cases (35x) for more rigorous evaluation. MBPP is a benchmark of 974 crowd-sourced Python programming problems designed to be solvable by entry-level programmers, covering programming fundamentals and standard library functionality.

Is a higher score better?

Yes. Higher values rank better for this benchmark.

How many models are compared?

4 unique published model results are currently shown.

Does this benchmark affect the overall score?

This benchmark is preserved as source-native evidence but is not eligible for the current overall score.

Rankings

OverallCodingText ArenaPricing

Modalities

Image GenerationVideo GenerationSpeech-to-TextEmbeddings

Benchmarks

All BenchmarksReasoningMathCoding

Vendors

All VendorsOpenAIAnthropicGoogle
llmboard.aiCopyright 2026 llmboard.ai