llmboard.ai
Benchmarks
CompareRankings
llmboard.ai
Benchmarks
CompareRankings
HomeBenchmarksreasoningSWE-Bench Verified

reasoning benchmark

SWE-Bench Verified

A verified subset of 500 software engineering problems from real GitHub issues, validated by human annotators for evaluating language models' ability to resolve real-world coding issues by generating patches for Python codebases.

Updated Aug 7, 2026

Published models100
Registry coverage104
MetricScore
EvidenceC

On this page

  • Ranking
  • Distribution
  • Highlights
  • About
  • FAQ

SWE-Bench Verified leaderboard

Sorted by the source-provided rank. Higher score is better according to the registry.

100 rows
Columns

Show columns

1ANClaude Fable 5Anthropic95.0%100.0%104CAug 7, 2026
2ANClaude Mythos PreviewAnthropic93.9%99.0%104CAug 7, 2026
3ANClaude Opus 4.8Anthropic88.6%98.1%104CAug 7, 2026
4ANClaude Opus 4.7Anthropic87.6%97.1%104CAug 7, 2026
5ANClaude Sonnet 5Anthropic85.2%96.1%104CAug 7, 2026
6ANClaude Opus 4.5Anthropic80.9%95.2%104CAug 7, 2026
7ANClaude Opus 4.6Anthropic80.8%94.2%104CAug 7, 2026
8DEDeepSeek-V4-Pro-MaxDeepSeek80.6%93.2%104CAug 7, 2026
9GOGemini 3.1 ProGoogle80.6%92.2%104CAug 7, 2026
10MIMiniMax M3MiniMax80.5%91.3%104CAug 7, 2026
11ACQwen3.7 MaxAlibaba Cloud / Qwen Team80.4%90.3%104CAug 7, 2026
12MAKimi K2.6Moonshot AI80.2%89.3%104CAug 7, 2026
13MIMiniMax M2.5MiniMax80.2%88.3%104CAug 7, 2026
14OPGPT-5.2OpenAI80.0%87.4%104CAug 7, 2026
15ANClaude Sonnet 4.6Anthropic79.6%86.4%104CAug 7, 2026
16DEDeepSeek-V4-Flash-MaxDeepSeek79.0%85.4%104CAug 7, 2026
17XIMiMo-V2.5-ProXiaomi78.9%84.5%104CAug 7, 2026
18ACQwen3.6 PlusAlibaba Cloud / Qwen Team78.8%83.5%104CAug 7, 2026
19GOGemini 3 FlashGoogle78.0%82.5%104CAug 7, 2026
20TEHy3Tencent78.0%81.5%104CAug 7, 2026
21XIMiMo-V2-ProXiaomi78.0%80.6%104CAug 7, 2026
22ZAGLM-5Zhipu AI77.8%79.6%104CAug 7, 2026
23ACQwen3.7-PlusAlibaba Cloud / Qwen Team77.7%78.6%104CAug 7, 2026
24MAMistral Medium 3.5Mistral AI77.6%77.7%104CAug 7, 2026
25MEMuse SparkMeta77.4%76.7%104CAug 7, 2026
26ACQwen3.6-27BAlibaba Cloud / Qwen Team77.2%75.7%104CAug 7, 2026
27MAKimi K2.5Moonshot AI76.8%74.8%104CAug 7, 2026
28BYSeed 2.0 ProByteDance76.5%73.8%104CAug 7, 2026
29ACQwen3.5-397B-A17BAlibaba Cloud / Qwen Team76.4%72.8%104CAug 7, 2026
30OPGPT-5.1OpenAI76.3%71.8%104CAug 7, 2026
31OPGPT-5.1 InstantOpenAI76.3%70.9%104CAug 7, 2026
32OPGPT-5.1 ThinkingOpenAI76.3%69.9%104CAug 7, 2026
33GOGemini 3 ProGoogle76.2%68.9%104CAug 7, 2026
34OPGPT-5OpenAI74.9%68.0%104CAug 7, 2026
35XIMiMo-V2-OmniXiaomi74.8%67.0%104CAug 7, 2026
36ANClaude Opus 4.1Anthropic74.5%66.0%104CAug 7, 2026
37OPGPT-5 CodexOpenAI74.5%65.0%104CAug 7, 2026
38STStep-3.5-FlashStepFun74.4%64.1%104CAug 7, 2026
39ZAGLM-4.7Zhipu AI73.8%63.1%104CAug 7, 2026
40OPGPT-5.1 CodexOpenAI73.7%62.1%104CAug 7, 2026
41MIMAI-Thinking-1Microsoft73.5%61.2%104CAug 7, 2026
42BYSeed 2.0 LiteByteDance73.5%60.2%104CAug 7, 2026
43XIMiMo-V2-FlashXiaomi73.4%59.2%104CAug 7, 2026
44ACQwen3.6-35B-A3BAlibaba Cloud / Qwen Team73.4%58.3%104CAug 7, 2026
45ANClaude Haiku 4.5Anthropic73.3%57.3%104CAug 7, 2026
46DEDeepSeek-V3.2 (Thinking)DeepSeek73.1%56.3%104CAug 7, 2026
47DEDeepSeek-V3.2DeepSeek73.1%55.3%104CAug 7, 2026
48DEDeepSeek-V3.2-SpecialeDeepSeek73.1%54.4%104CAug 7, 2026
49ANClaude Sonnet 4Anthropic72.7%53.4%104CAug 7, 2026
50ANClaude Opus 4Anthropic72.5%52.4%104CAug 7, 2026
51ACQwen3.5-27BAlibaba Cloud / Qwen Team72.4%51.5%104CAug 7, 2026
52ACQwen3.5-122B-A10BAlibaba Cloud / Qwen Team72.0%50.5%104CAug 7, 2026
53MIMAI-Code-1-FlashMicrosoft71.6%49.5%104CAug 7, 2026
54MAKimi K2-Thinking-0905Moonshot AI71.3%48.5%104CAug 7, 2026
55XAGrok Code Fast 1xAI70.8%47.6%104CAug 7, 2026
56NVNemotron 3 Ultra (550B A55B)NVIDIA70.7%46.6%104CAug 7, 2026
57ANClaude 3.7 SonnetAnthropic70.3%45.6%104CAug 7, 2026
58MELongCat-Flash-Thinking-2601Meituan70.0%44.7%104CAug 7, 2026
59AMNova 2 ProAmazon70.0%43.7%104CAug 7, 2026
60ACQwen3-Coder 480B A35B InstructAlibaba Cloud / Qwen Team69.6%42.7%104CAug 7, 2026
61ACQwen3 MaxAlibaba Cloud / Qwen Team69.6%41.8%104CAug 7, 2026
62MIMiniMax M2MiniMax69.4%40.8%104CAug 7, 2026
63ACQwen3.5-35B-A3BAlibaba Cloud / Qwen Team69.2%39.8%104CAug 7, 2026
64OPo3OpenAI69.1%38.8%104CAug 7, 2026
65OPo4-miniOpenAI68.1%37.9%104CAug 7, 2026
66ZAGLM-4.6Zhipu AI68.0%36.9%104CAug 7, 2026
67DEDeepSeek-V3.2-ExpDeepSeek67.8%35.9%104CAug 7, 2026
68CONorth Mini Code 1.0Cohere67.6%35.0%104CAug 7, 2026
69GOGemini 2.5 Pro Preview 06-05Google67.2%34.0%104CAug 7, 2026
70MIMiniMax M2.1MiniMax67.0%33.0%104CAug 7, 2026
71DEDeepSeek-V3.1DeepSeek66.0%32.0%104CAug 7, 2026
72MAKimi K2-Instruct-0905Moonshot AI65.8%31.1%104CAug 7, 2026
73AMNova 2 LiteAmazon64.5%30.1%104CAug 7, 2026
74ZAGLM-4.5Zhipu AI64.2%29.1%104CAug 7, 2026
75GOGemini 2.5 ProGoogle63.2%28.2%104CAug 7, 2026
76MADevstral MediumMistral AI61.6%27.2%104CAug 7, 2026
77GOGemini 2.5 FlashGoogle60.4%26.2%104CAug 7, 2026
78MELongCat-Flash-ChatMeituan60.4%25.2%104CAug 7, 2026
79MELongCat-Flash-ThinkingMeituan59.4%24.3%104CAug 7, 2026
80ZAGLM-4.7-FlashZhipu AI59.2%23.3%104CAug 7, 2026
81ZAGLM-4.5-AirZhipu AI57.6%22.3%104CAug 7, 2026
82MIMiniMax M1 80KMiniMax56.0%21.4%104CAug 7, 2026
83MIMiniMax M1 40KMiniMax55.6%20.4%104CAug 7, 2026
84OPGPT-4.1OpenAI54.6%19.4%104CAug 7, 2026
85MELongCat-Flash-LiteMeituan54.4%18.4%104CAug 7, 2026
86NVNemotron 3 Super (120B A12B)NVIDIA53.7%17.5%104CAug 7, 2026
87MADevstral Small 1.1Mistral AI53.6%16.5%104CAug 7, 2026
88OPo3-miniOpenAI49.3%15.5%104CAug 7, 2026
89ANClaude 3.5 SonnetAnthropic49.0%14.6%104CAug 7, 2026
90SASarvam-105BSarvam AI45.0%13.6%104CAug 7, 2026
91DEDeepSeek-R1-0528DeepSeek44.6%12.6%104CAug 7, 2026
92DEDeepSeek-V3DeepSeek42.0%11.7%104CAug 7, 2026
93OPo1-previewOpenAI41.3%10.7%104CAug 7, 2026
94OPo1OpenAI41.0%9.7%104CAug 7, 2026
95ANClaude 3.5 HaikuAnthropic40.6%8.7%104CAug 7, 2026
96NVNemotron 3 Nano (30B A3B)NVIDIA38.8%7.8%104CAug 7, 2026
97OPGPT-4.5OpenAI38.0%6.8%104CAug 7, 2026
98SASarvam-30BSarvam AI34.0%5.8%104CAug 7, 2026
99OPGPT-4oOpenAI33.2%4.8%104CAug 7, 2026
100GOGemini 2.5 Flash-LiteGoogle31.6%3.9%104CAug 7, 2026

Score distribution

Top published rows on the benchmark's original scale.

SWE-Bench Verified

SWE-Bench Verified highlights

The top published results on this benchmark's own scale.

Rank #1Claude Fable 595.0%Rank #2Claude Mythos Preview93.9%Rank #3Claude Opus 4.888.6%Rank #4Claude Opus 4.787.6%

What is SWE-Bench Verified?

Definition and scoring fields from the benchmark registry.

A verified subset of 500 software engineering problems from real GitHub issues, validated by human annotators for evaluating language models' ability to resolve real-world coding issues by generating patches for Python codebases.

Scores are shown in ratio. The current registry marks this benchmark as not independently verified with evidence level C.

Family
SWE-Bench Verified
Modality
text
Primary category
reasoning
Score direction
higher
LLMBoard eligible
No
Evaluation key
swe-bench-verified|llm-stats-current

Source-native results are preserved. Eligibility for the overall LLMBoard score is a separate policy decision.

FAQ

Common questions about SWE-Bench Verified.

Which model scores highest on SWE-Bench Verified?

Claude Fable 5 is currently ranked first with 95.0%.

What does SWE-Bench Verified measure?

A verified subset of 500 software engineering problems from real GitHub issues, validated by human annotators for evaluating language models' ability to resolve real-world coding issues by generating patches for Python codebases.

Is a higher score better?

Yes. Higher values rank better for this benchmark.

How many models are compared?

100 unique published model results are currently shown.

Does this benchmark affect the overall score?

This benchmark is preserved as source-native evidence but is not eligible for the current overall score.

Rankings

OverallCodingText ArenaPricing

Modalities

Image GenerationVideo GenerationSpeech-to-TextEmbeddings

Benchmarks

All BenchmarksReasoningMathCoding

Vendors

All VendorsOpenAIAnthropicGoogle
llmboard.aiCopyright 2026 llmboard.ai