llmboard.ai
Benchmarks
CompareRankings
llmboard.ai
Benchmarks
CompareRankings
HomeBenchmarksgeneralAider-Polyglot

general benchmark

Aider-Polyglot

A coding benchmark that evaluates LLMs on 225 challenging Exercism programming exercises across C++, Go, Java, JavaScript, Python, and Rust. Models receive two attempts to solve each problem, with test error feedback provided after the first attempt if it fails. The benchmark measures both initial problem-solving ability and capacity to edit code based on error feedback, providing an end-to-end evaluation of code generation and editing capabilities across multiple programming languages.

Updated Aug 7, 2026

Published models22
Registry coverage22
MetricScore
EvidenceB

On this page

  • Ranking
  • Distribution
  • Highlights
  • About
  • FAQ

Aider-Polyglot leaderboard

Sorted by the source-provided rank. Higher score is better according to the registry.

22 rows
Columns

Show columns

1OPGPT-5OpenAI88.0%100.0%22CAug 7, 2026
2GOGemini 2.5 Pro Preview 06-05Google82.2%95.2%22CAug 7, 2026
3OPo3OpenAI81.3%90.5%22CAug 7, 2026
4GOGemini 2.5 ProGoogle76.5%85.7%22CAug 7, 2026
5DEDeepSeek-V3.2-ExpDeepSeek74.5%81.0%22CAug 7, 2026
6DEDeepSeek-R1-0528DeepSeek71.6%76.2%22CAug 7, 2026
7OPo4-miniOpenAI68.9%71.4%22CAug 7, 2026
8DEDeepSeek-V3.1DeepSeek68.4%66.7%22CAug 7, 2026
9OPo3-miniOpenAI66.7%61.9%22CAug 7, 2026
10GOGemini 2.5 FlashGoogle61.9%57.1%22CAug 7, 2026
11ACQwen3-Coder 480B A35B InstructAlibaba Cloud / Qwen Team61.8%52.4%22CAug 7, 2026
12MAKimi K2 InstructMoonshot AI60.0%47.6%22CAug 7, 2026
13MAKimi K2-Instruct-0905Moonshot AI60.0%42.9%22CAug 7, 2026
14ACQwen3-235B-A22B-Instruct-2507Alibaba Cloud / Qwen Team57.3%38.1%22CAug 7, 2026
15OPGPT-4.1OpenAI51.6%33.3%22CAug 7, 2026
16ACQwen3-Next-80B-A3B-InstructAlibaba Cloud / Qwen Team49.8%28.6%22CAug 7, 2026
17DEDeepSeek-V3DeepSeek49.6%23.8%22CAug 7, 2026
18MAMagistral MediumMistral AI47.1%19.1%22CAug 7, 2026
19OPGPT-4.1 miniOpenAI34.7%14.3%22CAug 7, 2026
20OPGPT-4oOpenAI30.7%9.5%22CAug 7, 2026
21GOGemini 2.5 Flash-LiteGoogle26.7%4.8%22CAug 7, 2026
22OPGPT-4.1 nanoOpenAI9.8%0.0%22CAug 7, 2026

Score distribution

Top published rows on the benchmark's original scale.

Aider-Polyglot

Aider-Polyglot highlights

The top published results on this benchmark's own scale.

Rank #1GPT-588.0%Rank #2Gemini 2.5 Pro Preview 06-0582.2%Rank #3o381.3%Rank #4Gemini 2.5 Pro76.5%

What is Aider-Polyglot?

Definition and scoring fields from the benchmark registry.

A coding benchmark that evaluates LLMs on 225 challenging Exercism programming exercises across C++, Go, Java, JavaScript, Python, and Rust. Models receive two attempts to solve each problem, with test error feedback provided after the first attempt if it fails. The benchmark measures both initial problem-solving ability and capacity to edit code based on error feedback, providing an end-to-end evaluation of code generation and editing capabilities across multiple programming languages.

Scores are shown in ratio. The current registry marks this benchmark as not independently verified with evidence level B.

Family
Aider-Polyglot
Modality
text
Primary category
general
Score direction
higher
LLMBoard eligible
Yes
Evaluation key
aider-polyglot|llm-stats-current

Source-native results are preserved. Eligibility for the overall LLMBoard score is a separate policy decision.

FAQ

Common questions about Aider-Polyglot.

Which model scores highest on Aider-Polyglot?

GPT-5 is currently ranked first with 88.0%.

What does Aider-Polyglot measure?

A coding benchmark that evaluates LLMs on 225 challenging Exercism programming exercises across C++, Go, Java, JavaScript, Python, and Rust. Models receive two attempts to solve each problem, with test error feedback provided after the first attempt if it fails. The benchmark measures both initial problem-solving ability and capacity to edit code based on error feedback, providing an end-to-end evaluation of code generation and editing capabilities across multiple programming languages.

Is a higher score better?

Yes. Higher values rank better for this benchmark.

How many models are compared?

22 unique published model results are currently shown.

Does this benchmark affect the overall score?

This benchmark is marked as eligible for the current LLMBoard capability methodology.

Rankings

OverallCodingText ArenaPricing

Modalities

Image GenerationVideo GenerationSpeech-to-TextEmbeddings

Benchmarks

All BenchmarksReasoningMathCoding

Vendors

All VendorsOpenAIAnthropicGoogle
llmboard.aiCopyright 2026 llmboard.ai