reasoning benchmark
Terminal-Bench is a benchmark for testing AI agents in real terminal environments. It evaluates how well agents can handle real-world, end-to-end tasks autonomously, including compiling code, training models, setting up servers, system administration, security tasks, data science workflows, and cybersecurity vulnerabilities. The benchmark consists of a dataset of ~100 hand-crafted, human-verified tasks and an execution harness that connects language models to a terminal sandbox.
Updated Aug 7, 2026
Sorted by the source-provided rank. Higher score is better according to the registry.
| 1 | AN | 50.0% | 100.0% | 25 | C | |
| 2 | MI | 47.9% | 95.8% | 25 | C | |
| 3 | MA | 47.1% | 91.7% | 25 | C | |
| 4 | MI | 46.3% | 87.5% | 25 | C | |
| 5 | AN | 43.3% | 83.3% | 25 | C | |
| 6 | AM | 41.3% | 79.2% | 25 | C | |
| 7 | AN | 41.0% | 75.0% | 25 | C | |
| 8 | ZA | 40.5% | 70.8% | 25 | C | |
| 9 | ME | 39.5% | 66.7% | 25 | C | |
| 10 | AN | 39.2% | 62.5% | 25 | C | |
| 11 | DE | 37.7% | 58.3% | 25 | C | |
| 12 | ZA | 37.5% | 54.2% | 25 | C | |
| 13 | AN | 35.5% | 50.0% | 25 | C | |
| 14 | AN | 35.2% | 45.8% | 25 | C | |
| 15 | ME | 33.8% | 41.7% | 25 | C | |
| 16 | ZA | 33.3% | 37.5% | 25 | C | |
| 17 | AM | 32.5% | 33.3% | 25 | C | |
| 18 | DE | 31.3% | 29.2% | 25 | C | |
| 19 | XI | 30.5% | 25.0% | 25 | C | |
| 20 | ZA | 30.0% | 20.8% | 25 | C | |
| 21 | MA | 30.0% | 16.7% | 25 | C | |
| 22 | NV | 25.8% | 12.5% | 25 | C | |
| 23 | MA | 25.0% | 8.3% | 25 | C | |
| 24 | NV | 8.5% | 4.2% | 25 | C | |
| 25 | DE | 5.7% | 0.0% | 25 | C |
Top published rows on the benchmark's original scale.
The top published results on this benchmark's own scale.
Definition and scoring fields from the benchmark registry.
Terminal-Bench is a benchmark for testing AI agents in real terminal environments. It evaluates how well agents can handle real-world, end-to-end tasks autonomously, including compiling code, training models, setting up servers, system administration, security tasks, data science workflows, and cybersecurity vulnerabilities. The benchmark consists of a dataset of ~100 hand-crafted, human-verified tasks and an execution harness that connects language models to a terminal sandbox.
Scores are shown in ratio. The current registry marks this benchmark as not independently verified with evidence level B.
Source-native results are preserved. Eligibility for the overall LLMBoard score is a separate policy decision.
Common questions about Terminal-Bench.
Claude Sonnet 4.5 is currently ranked first with 50.0%.
Terminal-Bench is a benchmark for testing AI agents in real terminal environments. It evaluates how well agents can handle real-world, end-to-end tasks autonomously, including compiling code, training models, setting up servers, system administration, security tasks, data science workflows, and cybersecurity vulnerabilities. The benchmark consists of a dataset of ~100 hand-crafted, human-verified tasks and an execution harness that connects language models to a terminal sandbox.
Yes. Higher values rank better for this benchmark.
25 unique published model results are currently shown.
This benchmark is preserved as source-native evidence but is not eligible for the current overall score.