Benchmark category
Published reasoning evaluations and source-native model rankings. Each benchmark keeps its original scale and methodology.
Data as of 2026-08-07
Select a benchmark to inspect model-level results, evidence fields and scoring direction.
| GPQA | GPQA | text | Score | 233 | featured | C | Yes |
| MMLU-Pro | MMLU-Pro | text | Score | 129 | featured | B | Yes |
| AIME 2025 | AIME 2025 | text | Score | 114 | featured | C | Yes |
| SWE-Bench Verified | SWE-Bench Verified | text | Score | 104 | featured | C | No |
| MMLU | MMLU | text | Score | 100 | featured | B | Yes |
| Humanity's Last Exam | Humanity's Last Exam | multimodal | Score | 92 | featured | B | No |
| LiveCodeBench | LiveCodeBench | text | Score | 73 | featured | C | Yes |
| MATH | MATH | text | Score | 71 | featured | B | Yes |
| HumanEval | HumanEval | text | Score | 66 | featured | B | Yes |
| MMMU-Pro | MMMU-Pro | multimodal | Score | 65 | featured | B | No |
| MMMU | MMMU | multimodal | Score | 63 | featured | B | No |
| BrowseComp | BrowseComp | text | Score | 58 | featured | B | Yes |
| AIME 2024 | AIME 2024 | text | Score | 53 | featured | B | Yes |
| LiveCodeBench v6 | LiveCodeBench v6 | text | Score | 53 | featured | B | Yes |
| MMMLU | MMMLU | text | Score | 49 | featured | B | Yes |
| Terminal-Bench 2.0 | Terminal-Bench 2.0 | text | Score | 49 | featured | C | No |
| GSM8k | GSM8k | text | Score | 48 | featured | B | Yes |
| MMLU-Redux | MMLU-Redux | text | Score | 48 | featured | B | Yes |
| CharXiv-R | CharXiv-R | multimodal | Score | 47 | featured | B | No |
| SimpleQA | SimpleQA | text | Score | 46 | featured | B | Yes |
| SWE-Bench Pro | SWE-Bench Pro | text | Score | 44 | featured | B | No |
| LiveBench | LiveBench | text | Score | 38 | featured | B | No |
| Tau2 Telecom | Tau2 Telecom | text | Score | 35 | featured | B | No |
| ARC-C | ARC-C | text | Score | 34 | featured | B | Yes |
| SuperGPQA | SuperGPQA | text | Score | 34 | featured | B | Yes |
| SWE-bench Multilingual | SWE-bench Multilingual | text | Score | 34 | featured | B | No |
| MBPP | MBPP | text | Score | 33 | featured | B | Yes |
| MATH-500 | MATH-500 | text | Score | 32 | featured | C | Yes |
| MMLU-ProX | MMLU-ProX | text | Score | 32 | featured | B | Yes |
| AI2D | AI2D | multimodal | Score | 32 | featured | B | No |
| MGSM | MGSM | text | Score | 31 | featured | B | Yes |
| Toolathlon | Toolathlon | text | Score | 31 | featured | B | Yes |
| DROP | DROP | text | Score | 30 | featured | B | Yes |
| MCP Atlas | MCP Atlas | text | Score | 30 | featured | B | Yes |
| Multi-Challenge | Multi-Challenge | text | Score | 29 | featured | C | Yes |
| HellaSwag | HellaSwag | text | Score | 27 | featured | B | Yes |
| Arena Hard | Arena Hard | text | Score | 26 | featured | B | No |
| Finance Agent v2 | Finance Agent v2 | text | Score | 26 | featured | B | No |
| Tau2 Retail | Tau2 Retail | text | Score | 26 | featured | B | No |
| VideoMMMU | VideoMMMU | multimodal | Score | 26 | featured | B | No |
| TAU-bench Retail | TAU-bench Retail | text | Score | 25 | featured | B | No |
| Terminal-Bench | Terminal-Bench | text | Score | 25 | featured | B | No |
| ChartQA | ChartQA | multimodal | Score | 24 | featured | B | No |
| t2-bench | t2-bench | text | Score | 23 | featured | B | Yes |
| ERQA | ERQA | multimodal | Score | 23 | featured | C | No |
| PolyMATH | PolyMATH | multimodal | Score | 23 | featured | B | No |
| TAU-bench Airline | TAU-bench Airline | text | Score | 23 | featured | B | No |
| Tau2 Airline | Tau2 Airline | text | Score | 23 | featured | B | No |
| Winogrande | Winogrande | text | Score | 22 | featured | B | Yes |
| MMStar | MMStar | multimodal | Score | 22 | featured | B | No |
| BIG-Bench Hard | BIG-Bench Hard | text | Score | 21 | featured | B | Yes |
| MRCR v2 (8-needle) | MRCR v2 (8-needle) | text | Score | 21 | featured | C | Yes |
| Multi-IF | Multi-IF | text | Score | 20 | featured | B | Yes |
| BFCL-v3 | BFCL-v3 | text | Score | 19 | featured | B | Yes |
| IMO-AnswerBench | IMO-AnswerBench | text | Score | 19 | featured | B | Yes |
| C-Eval | C-Eval | text | Score | 18 | featured | B | Yes |
| SciCode | SciCode | text | Score | 18 | featured | B | Yes |
| TriviaQA | TriviaQA | text | Score | 18 | featured | B | Yes |
| TruthfulQA | TruthfulQA | text | Score | 18 | featured | B | Yes |
| MMBench-V1.1 | MMBench-V1.1 | multimodal | Score | 18 | featured | B | No |
| AIME 2026 | AIME 2026 | text | Score | 17 | featured | B | Yes |
| FrontierMath | FrontierMath | text | Score | 17 | featured | B | Yes |
| LongBench v2 | LongBench v2 | text | Score | 17 | featured | C | Yes |
| MVBench | MVBench | multimodal | Score | 17 | featured | B | No |
| Terminal-Bench 2.1 | Terminal-Bench 2.1 | text | Score | 17 | featured | B | No |
| Video-MME | Video-MME | multimodal | Score | 17 | featured | B | No |
| CodeForces | CodeForces | text | Score | 16 | featured | B | Yes |
| ARC-AGI v2 | ARC-AGI v2 | multimodal | Score | 16 | featured | B | No |
| Arena-Hard v2 | Arena-Hard v2 | text | Score | 16 | featured | B | No |
| CharXiv-D | CharXiv-D | multimodal | Score | 16 | featured | B | No |
| Hallusion Bench | Hallusion Bench | multimodal | Score | 16 | featured | B | No |
| OmniDocBench 1.5 | OmniDocBench 1.5 | multimodal | Score | 16 | featured | B | No |
| FrontierCode 1.1 | FrontierCode 1.1 | text | Score | 15 | featured | B | Yes |
| AA-LCR | AA-LCR | text | Score | 15 | featured | B | No |
| Global-MMLU-Lite | Global-MMLU-Lite | text | Score | 14 | featured | B | Yes |
| LiveBench 20241125 | LiveBench 20241125 | text | Score | 14 | featured | B | No |
| BrowseComp-zh | BrowseComp-zh | text | Score | 13 | featured | B | Yes |
| FACTS Grounding | FACTS Grounding | text | Score | 13 | featured | C | Yes |
| Global PIQA | Global PIQA | text | Score | 13 | featured | B | Yes |
| HiddenMath | HiddenMath | text | Score | 13 | featured | B | Yes |
| BLINK | BLINK | multimodal | Score | 13 | featured | B | No |
| Legal Agent Benchmark | Legal Agent Benchmark | text | Score | 13 | featured | B | No |
| BBH | BBH | text | Score | 12 | featured | B | Yes |
| MT-Bench | MT-Bench | text | Score | 12 | featured | B | Yes |
| MedXpertQA | MedXpertQA | multimodal | Score | 12 | featured | B | No |
| BFCL | BFCL | text | Score | 11 | featured | B | Yes |
| BIG-Bench Extra Hard | BIG-Bench Extra Hard | text | Score | 11 | featured | B | Yes |
| Graphwalks BFS <128k | Graphwalks BFS <128k | text | Score | 11 | featured | B | Yes |
| Graphwalks BFS >128k | Graphwalks BFS >128k | text | Score | 11 | featured | B | Yes |
| Graphwalks parents <128k | Graphwalks parents <128k | text | Score | 11 | featured | B | Yes |
| HMMT Feb 26 | HMMT Feb 26 | text | Score | 11 | featured | B | Yes |
| PIQA | PIQA | text | Score | 11 | featured | B | Yes |
| MMMU (val) | MMMU (val) | multimodal | Score | 11 | featured | B | No |
| MuirBench | MuirBench | multimodal | Score | 11 | featured | B | No |
| AGIEval | AGIEval | text | Score | 10 | featured | B | Yes |
| BoolQ | BoolQ | text | Score | 10 | featured | B | Yes |
| COLLIE | COLLIE | text | Score | 10 | featured | B | Yes |
| HumanEval+ | HumanEval+ | text | Score | 10 | featured | B | Yes |
| VITA-Bench | VITA-Bench | text | Score | 10 | featured | B | Yes |
High-coverage benchmarks with at least two published model results.
How this category is assembled on llmboard.ai.
This page groups benchmarks whose primary or display category matches reasoning. It does not average incompatible metrics into a new category score.
Open an individual benchmark to inspect score direction, evidence level, participant count and source-native results.
Category membership is derived from the current benchmark registry response.