reasoning benchmark
MT-Bench is a challenging multi-turn benchmark that measures the ability of large language models to engage in coherent, informative, and engaging conversations. It uses strong LLMs as judges for scalable and explainable evaluation of multi-turn dialogue capabilities.
Updated Aug 7, 2026
Sorted by the source-provided rank. Higher score is better according to the registry.
| 1 | AC | 0.935 points | 100.0% | 12 | C | |
| 2 | NV | 0.917 points | 90.9% | 12 | C | |
| 3 | DE | 0.902 points | 81.8% | 12 | C | |
| 4 | NR | 8.99 points | 72.7% | 12 | C | |
| 5 | AC | 0.875 points | 63.6% | 12 | C | |
| 6 | MA | 0.863 points | 54.5% | 12 | C | |
| 7 | AC | 0.841 points | 45.5% | 12 | C | |
| 8 | MA | 0.835 points | 36.4% | 12 | C | |
| 9 | MA | 0.83 points | 27.3% | 12 | C | |
| 10 | NV | 0.81 points | 18.2% | 12 | C | |
| 11 | MA | 0.768 points | 9.1% | 12 | C | |
| 12 | NV | 0.09 points | 0.0% | 12 | C |
Top published rows on the benchmark's original scale.
The top published results on this benchmark's own scale.
Definition and scoring fields from the benchmark registry.
MT-Bench is a challenging multi-turn benchmark that measures the ability of large language models to engage in coherent, informative, and engaging conversations. It uses strong LLMs as judges for scalable and explainable evaluation of multi-turn dialogue capabilities.
Scores are shown in points. The current registry marks this benchmark as not independently verified with evidence level B.
Source-native results are preserved. Eligibility for the overall LLMBoard score is a separate policy decision.
Common questions about MT-Bench.
Qwen2.5 72B Instruct is currently ranked first with 0.935 points.
MT-Bench is a challenging multi-turn benchmark that measures the ability of large language models to engage in coherent, informative, and engaging conversations. It uses strong LLMs as judges for scalable and explainable evaluation of multi-turn dialogue capabilities.
Yes. Higher values rank better for this benchmark.
12 unique published model results are currently shown.
This benchmark is marked as eligible for the current LLMBoard capability methodology.