general benchmark
A challenging multi-language coding benchmark that evaluates models' code editing abilities across C++, Go, Java, JavaScript, Python, and Rust. Contains 225 of Exercism's most difficult programming problems, selected as problems that were solved by 3 or fewer out of 7 top coding models. The benchmark focuses on code editing tasks and measures both correctness of solutions and proper edit format usage. Designed to re-calibrate evaluation scales so top models score between 5-50%.
Updated Aug 7, 2026
Sorted by the source-provided rank. Higher score is better according to the registry.
| 1 | DE | 79.7% | 100.0% | 10 | C | |
| 2 | GO | 72.7% | 88.9% | 10 | C | |
| 3 | OP | 60.4% | 77.8% | 10 | C | |
| 4 | OP | 58.2% | 66.7% | 10 | C | |
| 5 | GO | 56.7% | 55.6% | 10 | C | |
| 6 | OP | 52.9% | 44.4% | 10 | C | |
| 7 | OP | 44.9% | 33.3% | 10 | C | |
| 8 | OP | 31.6% | 22.2% | 10 | C | |
| 9 | OP | 18.2% | 11.1% | 10 | C | |
| 10 | OP | 6.2% | 0.0% | 10 | C |
Top published rows on the benchmark's original scale.
The top published results on this benchmark's own scale.
Definition and scoring fields from the benchmark registry.
A challenging multi-language coding benchmark that evaluates models' code editing abilities across C++, Go, Java, JavaScript, Python, and Rust. Contains 225 of Exercism's most difficult programming problems, selected as problems that were solved by 3 or fewer out of 7 top coding models. The benchmark focuses on code editing tasks and measures both correctness of solutions and proper edit format usage. Designed to re-calibrate evaluation scales so top models score between 5-50%.
Scores are shown in ratio. The current registry marks this benchmark as not independently verified with evidence level B.
Source-native results are preserved. Eligibility for the overall LLMBoard score is a separate policy decision.
Common questions about Aider-Polyglot Edit.
DeepSeek-V3 is currently ranked first with 79.7%.
A challenging multi-language coding benchmark that evaluates models' code editing abilities across C++, Go, Java, JavaScript, Python, and Rust. Contains 225 of Exercism's most difficult programming problems, selected as problems that were solved by 3 or fewer out of 7 top coding models. The benchmark focuses on code editing tasks and measures both correctness of solutions and proper edit format usage. Designed to re-calibrate evaluation scales so top models score between 5-50%.
Yes. Higher values rank better for this benchmark.
10 unique published model results are currently shown.
This benchmark is marked as eligible for the current LLMBoard capability methodology.