reasoning benchmark
FrontierCode 1.1 evaluates whether coding-agent changes are mergeable, using unit tests, maintainer-defined rubrics, and verifiers. Runs flagged for unfair internet use receive a zero score.
Updated Aug 7, 2026
Sorted by the source-provided rank. Higher score is better according to the registry.
| 1 | AN | 53.5% | 100.0% | 15 | B | |
| 2 | AN | 53.4% | 92.9% | 15 | B | |
| 3 | OP | 47.5% | 85.7% | 15 | B | |
| 4 | AN | 46.5% | 78.6% | 15 | B | |
| 5 | OP | 43.0% | 71.4% | 15 | B | |
| 6 | AN | 42.7% | 64.3% | 15 | B | |
| 7 | XA | 42.4% | 57.1% | 15 | B | |
| 8 | OP | 41.3% | 50.0% | 15 | B | |
| 9 | OP | 39.8% | 42.9% | 15 | B | |
| 10 | AN | 38.5% | 35.7% | 15 | B | |
| 11 | MA | 30.1% | 28.6% | 15 | B | |
| 12 | ZA | 24.5% | 21.4% | 15 | B | |
| 13 | DE | 17.6% | 14.3% | 15 | B | |
| 14 | MI | 14.7% | 7.1% | 15 | B | |
| 15 | AC | 10.2% | 0.0% | 15 | B |
Top published rows on the benchmark's original scale.
The top published results on this benchmark's own scale.
Definition and scoring fields from the benchmark registry.
FrontierCode 1.1 evaluates whether coding-agent changes are mergeable, using unit tests, maintainer-defined rubrics, and verifiers. Runs flagged for unfair internet use receive a zero score.
Scores are shown in ratio. The current registry marks this benchmark as not independently verified with evidence level B.
Source-native results are preserved. Eligibility for the overall LLMBoard score is a separate policy decision.
Common questions about FrontierCode 1.1.
Claude Fable 5 is currently ranked first with 53.5%.
FrontierCode 1.1 evaluates whether coding-agent changes are mergeable, using unit tests, maintainer-defined rubrics, and verifiers. Runs flagged for unfair internet use receive a zero score.
Yes. Higher values rank better for this benchmark.
15 unique published model results are currently shown.
This benchmark is marked as eligible for the current LLMBoard capability methodology.