math benchmark
LiveBench is a challenging, contamination-limited LLM benchmark that addresses test set contamination by releasing new questions monthly based on recently-released datasets, arXiv papers, news articles, and IMDb movie synopses. It comprises tasks across math, coding, reasoning, language, instruction following, and data analysis with verifiable, objective ground-truth answers.
Updated Aug 7, 2026
Sorted by the source-provided rank. Higher score is better according to the registry.
| 1 | AC | 79.6% | 100.0% | 14 | C | |
| 2 | AC | 78.4% | 92.3% | 14 | C | |
| 3 | AC | 76.6% | 84.6% | 14 | C | |
| 4 | AC | 75.8% | 76.9% | 14 | C | |
| 5 | AC | 75.4% | 69.2% | 14 | C | |
| 6 | AC | 74.8% | 61.5% | 14 | C | |
| 7 | AC | 74.7% | 53.9% | 14 | C | |
| 8 | AC | 72.2% | 46.1% | 14 | C | |
| 9 | AC | 72.1% | 38.5% | 14 | C | |
| 10 | AC | 69.8% | 30.8% | 14 | C | |
| 11 | AC | 68.4% | 23.1% | 14 | C | |
| 12 | AC | 65.4% | 15.4% | 14 | C | |
| 13 | AC | 62.0% | 7.7% | 14 | C | |
| 14 | AC | 60.9% | 0.0% | 14 | C |
Top published rows on the benchmark's original scale.
The top published results on this benchmark's own scale.
Definition and scoring fields from the benchmark registry.
LiveBench is a challenging, contamination-limited LLM benchmark that addresses test set contamination by releasing new questions monthly based on recently-released datasets, arXiv papers, news articles, and IMDb movie synopses. It comprises tasks across math, coding, reasoning, language, instruction following, and data analysis with verifiable, objective ground-truth answers.
Scores are shown in ratio. The current registry marks this benchmark as not independently verified with evidence level B.
Source-native results are preserved. Eligibility for the overall LLMBoard score is a separate policy decision.
Common questions about LiveBench 20241125.
Qwen3 VL 235B A22B Thinking is currently ranked first with 79.6%.
LiveBench is a challenging, contamination-limited LLM benchmark that addresses test set contamination by releasing new questions monthly based on recently-released datasets, arXiv papers, news articles, and IMDb movie synopses. It comprises tasks across math, coding, reasoning, language, instruction following, and data analysis with verifiable, objective ground-truth answers.
Yes. Higher values rank better for this benchmark.
14 unique published model results are currently shown.
This benchmark is preserved as source-native evidence but is not eligible for the current overall score.