llmboard.ai
Benchmarks
CompareRankings
llmboard.ai
Benchmarks
CompareRankings
HomeBenchmarksreasoningLiveCodeBench

reasoning benchmark

LiveCodeBench

LiveCodeBench is a holistic and contamination-free evaluation benchmark for large language models for code. It continuously collects new problems from programming contests (LeetCode, AtCoder, CodeForces) and evaluates four different scenarios: code generation, self-repair, code execution, and test output prediction. Problems are annotated with release dates to enable evaluation on unseen problems released after a model's training cutoff.

Updated Aug 7, 2026

Published models73
Registry coverage73
MetricScore
EvidenceC

On this page

  • Ranking
  • Distribution
  • Highlights
  • About
  • FAQ

LiveCodeBench leaderboard

Sorted by the source-provided rank. Higher score is better according to the registry.

73 rows
Columns

Show columns

1DEDeepSeek-V4-Pro-MaxDeepSeek93.5%100.0%73CAug 7, 2026
2DEDeepSeek-V4-Flash-MaxDeepSeek91.6%98.6%73CAug 7, 2026
3DEDeepSeek-V3.2 (Thinking)DeepSeek83.3%97.2%73CAug 7, 2026
4DEDeepSeek-V3.2DeepSeek83.3%95.8%73CAug 7, 2026
5MIMiniMax M2MiniMax83.0%94.4%73CAug 7, 2026
6MELongCat-Flash-Thinking-2601Meituan82.8%93.1%73CAug 7, 2026
7NVNemotron 3 Super (120B A12B)NVIDIA81.2%91.7%73CAug 7, 2026
8XAGrok-3 MinixAI80.4%90.3%73CAug 7, 2026
9XAGrok 4 FastxAI80.0%88.9%73CAug 7, 2026
10XAGrok-3xAI79.4%87.5%73CAug 7, 2026
11XAGrok-4 HeavyxAI79.4%86.1%73CAug 7, 2026
12MELongCat-Flash-ThinkingMeituan79.4%84.7%73CAug 7, 2026
13XAGrok-4xAI79.0%83.3%73CAug 7, 2026
14MIMiniMax M2.1MiniMax78.0%81.9%73CAug 7, 2026
15AMNova 2 ProAmazon74.6%80.6%73CAug 7, 2026
16DEDeepSeek-V3.2-ExpDeepSeek74.1%79.2%73CAug 7, 2026
17DEDeepSeek-R1-0528DeepSeek73.3%77.8%73CAug 7, 2026
18ZAGLM-4.5Zhipu AI72.9%76.4%73CAug 7, 2026
19NVNemotron Nano 9B v2NVIDIA71.1%75.0%73CAug 7, 2026
20AMNova 2 LiteAmazon71.0%73.6%73CAug 7, 2026
21ZAGLM-4.5-AirZhipu AI70.7%72.2%73CAug 7, 2026
22ACQwen3 235B A22BAlibaba Cloud / Qwen Team70.7%70.8%73CAug 7, 2026
23GOGemini 2.5 Pro Preview 06-05Google69.0%69.4%73CAug 7, 2026
24INMercury 2Inception67.0%68.1%73CAug 7, 2026
25NVLlama 3.1 Nemotron Ultra 253B v1NVIDIA66.3%66.7%73CAug 7, 2026
26ACQwen3 32BAlibaba Cloud / Qwen Team65.7%65.3%73CAug 7, 2026
27MIMiniMax M1 80KMiniMax65.0%63.9%73CAug 7, 2026
28MAMinistral 3 (14B Reasoning 2512)Mistral AI64.6%62.5%73CAug 7, 2026
29MAMistral Small 4Mistral AI63.6%61.1%73CAug 7, 2026
30ACQwQ-32BAlibaba Cloud / Qwen Team63.4%59.7%73CAug 7, 2026
31ACQwen3 30B A3BAlibaba Cloud / Qwen Team62.6%58.3%73CAug 7, 2026
32MIMiniMax M1 40KMiniMax62.3%56.9%73CAug 7, 2026
33MAMinistral 3 (8B Reasoning 2512)Mistral AI61.6%55.6%73CAug 7, 2026
34DEDeepSeek R1 Distill Llama 70BDeepSeek57.5%54.2%73CAug 7, 2026
35DEDeepSeek R1 Distill Qwen 32BDeepSeek57.2%52.8%73CAug 7, 2026
36DEDeepSeek-V3.1DeepSeek56.4%51.4%73CAug 7, 2026
37ACQwen2.5 72B InstructAlibaba Cloud / Qwen Team55.5%50.0%73CAug 7, 2026
38MAMin istral 3 (3B Reasoning 2512)Mistral AI54.8%48.6%73CAug 7, 2026
39MIPhi 4 ReasoningMicrosoft53.8%47.2%73CAug 7, 2026
40MAKimi K2-Instruct-0905Moonshot AI53.7%45.8%73CAug 7, 2026
41DEDeepSeek R1 Distill Qwen 14BDeepSeek53.1%44.4%73CAug 7, 2026
42MIPhi 4 Reasoning PlusMicrosoft53.1%43.1%73CAug 7, 2026
43MAMagistral Small 2506Mistral AI51.3%41.7%73CAug 7, 2026
44MAMagistral MediumMistral AI50.3%40.3%73CAug 7, 2026
45DEDeepSeek R1 ZeroDeepSeek50.0%38.9%73CAug 7, 2026
46ACQwQ-32B-PreviewAlibaba Cloud / Qwen Team50.0%37.5%73CAug 7, 2026
47DEDeepSeek-V3 0324DeepSeek49.2%36.1%73CAug 7, 2026
48MELongCat-Flash-ChatMeituan48.0%34.7%73CAug 7, 2026
49MELlama 4 MaverickMeta43.4%33.3%73CAug 7, 2026
50DEDeepSeek R1 Distill Llama 8BDeepSeek39.6%31.9%73CAug 7, 2026
51DEDeepSeek R1 Distill Qwen 7BDeepSeek37.6%30.6%73CAug 7, 2026
52DEDeepSeek-V3DeepSeek37.6%29.2%73CAug 7, 2026
53GOGemini 2.0 FlashGoogle35.1%27.8%73CAug 7, 2026
54MAMistral Large 3 (675B Instruct 2512 Eagle)Mistral AI34.4%26.4%73CAug 7, 2026
55MAMistral Large 3 (675B Base)Mistral AI34.4%25.0%73CAug 7, 2026
56MAMistral Large 3 (675B Instruct 2512 NVFP4)Mistral AI34.4%23.6%73CAug 7, 2026
57MAMistral Large 3 (675B Instruct 2512)Mistral AI34.4%22.2%73CAug 7, 2026
58GOGemini 2.5 Flash-LiteGoogle33.7%20.8%73CAug 7, 2026
59MELlama 4 ScoutMeta32.8%19.4%73CAug 7, 2026
60ACQwen2.5-Coder 32B InstructAlibaba Cloud / Qwen Team31.4%18.1%73CAug 7, 2026
61GOGemini DiffusionGoogle30.9%16.7%73CAug 7, 2026
62GOGemma 3 27BGoogle29.7%15.3%73CAug 7, 2026
63ACQwen2.5 7B InstructAlibaba Cloud / Qwen Team28.7%13.9%73CAug 7, 2026
64ACQwen2 7B InstructAlibaba Cloud / Qwen Team26.6%12.5%73CAug 7, 2026
65GOGemma 3 12BGoogle24.6%11.1%73CAug 7, 2026
66ACQwen2.5-Coder 7B InstructAlibaba Cloud / Qwen Team18.2%9.7%73CAug 7, 2026
67DEDeepSeek R1 Distill Qwen 1.5BDeepSeek16.9%8.3%73CAug 7, 2026
68GOGemma 3n E2B InstructedGoogle13.2%6.9%73CAug 7, 2026
69GOGemma 3n E2B Instructed LiteRT (Preview)Google13.2%5.6%73CAug 7, 2026
70GOGemma 3n E4B InstructedGoogle13.2%4.2%73CAug 7, 2026
71GOGemma 3n E4B Instructed LiteRT PreviewGoogle13.2%2.8%73CAug 7, 2026
72GOGemma 3 4BGoogle12.6%1.4%73CAug 7, 2026
73GOGemma 3 1BGoogle1.9%0.0%73CAug 7, 2026

Score distribution

Top published rows on the benchmark's original scale.

LiveCodeBench

LiveCodeBench highlights

The top published results on this benchmark's own scale.

Rank #1DeepSeek-V4-Pro-Max93.5%Rank #2DeepSeek-V4-Flash-Max91.6%Rank #3DeepSeek-V3.2 (Thinking)83.3%Rank #4DeepSeek-V3.283.3%

What is LiveCodeBench?

Definition and scoring fields from the benchmark registry.

LiveCodeBench is a holistic and contamination-free evaluation benchmark for large language models for code. It continuously collects new problems from programming contests (LeetCode, AtCoder, CodeForces) and evaluates four different scenarios: code generation, self-repair, code execution, and test output prediction. Problems are annotated with release dates to enable evaluation on unseen problems released after a model's training cutoff.

Scores are shown in ratio. The current registry marks this benchmark as not independently verified with evidence level C.

Family
LiveCodeBench
Modality
text
Primary category
reasoning
Score direction
higher
LLMBoard eligible
Yes
Evaluation key
livecodebench|llm-stats-current

Source-native results are preserved. Eligibility for the overall LLMBoard score is a separate policy decision.

FAQ

Common questions about LiveCodeBench.

Which model scores highest on LiveCodeBench?

DeepSeek-V4-Pro-Max is currently ranked first with 93.5%.

What does LiveCodeBench measure?

LiveCodeBench is a holistic and contamination-free evaluation benchmark for large language models for code. It continuously collects new problems from programming contests (LeetCode, AtCoder, CodeForces) and evaluates four different scenarios: code generation, self-repair, code execution, and test output prediction. Problems are annotated with release dates to enable evaluation on unseen problems released after a model's training cutoff.

Is a higher score better?

Yes. Higher values rank better for this benchmark.

How many models are compared?

73 unique published model results are currently shown.

Does this benchmark affect the overall score?

This benchmark is marked as eligible for the current LLMBoard capability methodology.

Rankings

OverallCodingText ArenaPricing

Modalities

Image GenerationVideo GenerationSpeech-to-TextEmbeddings

Benchmarks

All BenchmarksReasoningMathCoding

Vendors

All VendorsOpenAIAnthropicGoogle
llmboard.aiCopyright 2026 llmboard.ai