llmboard.ai
Benchmarks
CompareRankings
llmboard.ai
Benchmarks
CompareRankings
HomeBenchmarksreasoningHumanEval

reasoning benchmark

HumanEval

A benchmark that measures functional correctness for synthesizing programs from docstrings, consisting of 164 original programming problems assessing language comprehension, algorithms, and simple mathematics

Updated Aug 7, 2026

Published models66
Registry coverage66
MetricScore
EvidenceB

On this page

  • Ranking
  • Distribution
  • Highlights
  • About
  • FAQ

HumanEval leaderboard

Sorted by the source-provided rank. Higher score is better according to the registry.

66 rows
Columns

Show columns

1OPMiniCPM-SALAOpenBMB95.1%100.0%76CAug 7, 2026
2MAKimi K2 0905Moonshot AI94.5%98.7%76CAug 7, 2026
3ANClaude 3.5 SonnetAnthropic93.7%97.3%76CAug 7, 2026
4OPGPT-5OpenAI93.4%96.0%76CAug 7, 2026
5MAKimi K2 InstructMoonshot AI93.3%94.7%76CAug 7, 2026
7ACQwen2.5-Coder 32B InstructAlibaba Cloud / Qwen Team92.7%92.0%76CAug 7, 2026
8OPo1-miniOpenAI92.4%90.7%76CAug 7, 2026
10SASarvam-30BSarvam AI92.1%88.0%76CAug 7, 2026
11ANClaude 3.5 SonnetAnthropic92.0%86.7%76CAug 7, 2026
12MAMistral Large 2Mistral AI92.0%85.3%76CAug 7, 2026
13ACQwen2.5 VL 32B InstructAlibaba Cloud / Qwen Team91.5%84.0%76CAug 7, 2026
14OPGPT-4oOpenAI90.2%82.7%76CAug 7, 2026
15IBGranite 3.3 8B BaseIBM89.7%81.3%76CAug 7, 2026
16IBGranite 3.3 8B InstructIBM89.7%80.0%76CAug 7, 2026
17GOGemini DiffusionGoogle89.6%78.7%76CAug 7, 2026
18DEDeepSeek-V2.5DeepSeek89.0%77.3%76CAug 7, 2026
19MELlama 3.1 405B InstructMeta89.0%76.0%76CAug 7, 2026
20AMNova ProAmazon89.0%74.7%76CAug 7, 2026
21MELongCat-Flash-ChatMeituan88.4%73.3%76CAug 7, 2026
22MAMistral Small 3.1 24B InstructMistral AI88.4%72.0%76CAug 7, 2026
23XAGrok-2xAI88.4%70.7%76CAug 7, 2026
24MELlama 3.3 70B InstructMeta88.4%69.3%76CAug 7, 2026
25ACQwen2.5 32B InstructAlibaba Cloud / Qwen Team88.4%68.0%76CAug 7, 2026
26ACQwen2.5-Coder 7B InstructAlibaba Cloud / Qwen Team88.4%66.7%76CAug 7, 2026
27ANClaude 3.5 HaikuAnthropic88.1%65.3%76CAug 7, 2026
28OPo1OpenAI88.1%64.0%76CAug 7, 2026
29OPGPT-4.5OpenAI88.0%62.7%76CAug 7, 2026
30GOGemma 3 27BGoogle87.8%61.3%76CAug 7, 2026
31OPGPT-4o miniOpenAI87.2%60.0%76CAug 7, 2026
32OPGPT-4 TurboOpenAI87.1%58.7%76CAug 7, 2026
33ACQwen2.5 72B InstructAlibaba Cloud / Qwen Team86.6%57.3%76CAug 7, 2026
36ACQwen2 72B InstructAlibaba Cloud / Qwen Team86.0%53.3%76CAug 7, 2026
37XAGrok-2 minixAI85.7%52.0%76CAug 7, 2026
38GOGemma 3 12BGoogle85.4%50.7%76CAug 7, 2026
39AMNova LiteAmazon85.4%49.3%76CAug 7, 2026
40ANClaude 3 OpusAnthropic84.9%48.0%76CAug 7, 2026
41MAMistral Small 3 24B InstructMistral AI84.8%46.7%76CAug 7, 2026
42ACQwen2.5 7B InstructAlibaba Cloud / Qwen Team84.8%45.3%76CAug 7, 2026
43GOGemini 1.5 ProGoogle84.1%44.0%76CAug 7, 2026
44ACQwen2.5 14B InstructAlibaba Cloud / Qwen Team83.5%42.7%76CAug 7, 2026
46MIPhi 4Microsoft82.6%40.0%76CAug 7, 2026
47IBIBM Granite 4.0 Tiny PreviewIBM82.4%38.7%76CAug 7, 2026
48MACodestral-22BMistral AI81.1%37.3%76CAug 7, 2026
49AMNova MicroAmazon81.1%36.0%76CAug 7, 2026
50MELlama 3.1 70B InstructMeta80.5%34.7%76CAug 7, 2026
51ACQwen2 7B InstructAlibaba Cloud / Qwen Team79.9%33.3%76CAug 7, 2026
52ACQwen2.5-Omni-7BAlibaba Cloud / Qwen Team78.7%32.0%76CAug 7, 2026
54ANClaude 3 HaikuAnthropic75.9%29.3%76CAug 7, 2026
56GOGemma 3n E4B InstructedGoogle75.0%26.7%76CAug 7, 2026
57GOGemma 3n E4B Instructed LiteRT PreviewGoogle75.0%25.3%76CAug 7, 2026
58GOGemini 1.5 FlashGoogle74.3%24.0%76CAug 7, 2026
59XAGrok-1.5xAI74.1%22.7%76CAug 7, 2026
60ANClaude 3 SonnetAnthropic73.0%21.3%76CAug 7, 2026
61MELlama 3.1 8B InstructMeta72.6%20.0%76CAug 7, 2026
62MAPixtral-12BMistral AI72.0%18.7%76CAug 7, 2026
63GOGemma 3 4BGoogle71.3%17.3%76CAug 7, 2026
64MIPhi-3.5-MoE-instructMicrosoft70.7%16.0%76CAug 7, 2026
65OPGPT-3.5 TurboOpenAI68.0%14.7%76BAug 7, 2026
66OPGPT-4OpenAI67.0%13.3%76CAug 7, 2026
67GOGemma 3n E2B InstructedGoogle66.5%12.0%76CAug 7, 2026
68GOGemma 3n E2B Instructed LiteRT (Preview)Google66.5%10.7%76CAug 7, 2026
69MIPhi-3.5-mini-instructMicrosoft62.8%9.3%76CAug 7, 2026
71GOGemma 2 27BGoogle51.8%6.7%76CAug 7, 2026
73GOGemma 3 1BGoogle41.5%4.0%76CAug 7, 2026
74GOGemma 2 9BGoogle40.2%2.7%76CAug 7, 2026
75MAMinistral 8B InstructMistral AI34.8%1.3%76CAug 7, 2026

Score distribution

Top published rows on the benchmark's original scale.

HumanEval

HumanEval highlights

The top published results on this benchmark's own scale.

Rank #1MiniCPM-SALA95.1%Rank #2Kimi K2 090594.5%Rank #3Claude 3.5 Sonnet93.7%Rank #4GPT-593.4%

What is HumanEval?

Definition and scoring fields from the benchmark registry.

A benchmark that measures functional correctness for synthesizing programs from docstrings, consisting of 164 original programming problems assessing language comprehension, algorithms, and simple mathematics

Scores are shown in ratio. The current registry marks this benchmark as not independently verified with evidence level B.

Family
HumanEval
Modality
text
Primary category
reasoning
Score direction
higher
LLMBoard eligible
Yes
Evaluation key
humaneval|llm-stats-current

Source-native results are preserved. Eligibility for the overall LLMBoard score is a separate policy decision.

FAQ

Common questions about HumanEval.

Which model scores highest on HumanEval?

MiniCPM-SALA is currently ranked first with 95.1%.

What does HumanEval measure?

A benchmark that measures functional correctness for synthesizing programs from docstrings, consisting of 164 original programming problems assessing language comprehension, algorithms, and simple mathematics

Is a higher score better?

Yes. Higher values rank better for this benchmark.

How many models are compared?

66 unique published model results are currently shown.

Does this benchmark affect the overall score?

This benchmark is marked as eligible for the current LLMBoard capability methodology.

Rankings

OverallCodingText ArenaPricing

Modalities

Image GenerationVideo GenerationSpeech-to-TextEmbeddings

Benchmarks

All BenchmarksReasoningMathCoding

Vendors

All VendorsOpenAIAnthropicGoogle
llmboard.aiCopyright 2026 llmboard.ai