llmboard.ai
Benchmarks
CompareRankings
llmboard.ai
Benchmarks
CompareRankings
HomeBenchmarks

Evaluation registry

AI Benchmarks

Browse published benchmark definitions, capability categories and source-native model result pages.

Data as of 2026-08-07

Benchmarks661
Categories35
Score eligible72
With model coverage658

Benchmark highlights

A quick view of registry breadth and current model coverage.

Highest model coverageGPQA233 modelsLargest categoryreasoning245 benchmarks
Scoring inputs72Eligible benchmark definitions
Registry total66135 primary categories

Capability categories

Open a category to compare its benchmark registry and leading result sets.

35 rows
Columns

Show columns

reasoning24534104
multimodal152065
math8425129
agents49019
safety16011
code1504
general15431
language11223
physics73233
spatial reasoning7026
healthcare609
productivity603
summarization605
structured output5165
image to text4022
long context4052
speech to text406
uncategorized4012
audio201
finance201
image-generation201
instruction following2128
3d101
creativity1113
factuality101
frontend development103
legal1115
medical101
memory101
psychology109
question answering101
research101
robotics101
video101
vision1016

All benchmarks

Open any benchmark to inspect its leaderboard, scoring direction, evidence fields and programmatic FAQ.

145 rows
Columns

Show columns

GPQAphysicsGPQAScore233featuredCYes
MMLU-PromathMMLU-ProScore129featuredBYes
AIME 2025mathAIME 2025Score114featuredCYes
SWE-Bench VerifiedreasoningSWE-Bench VerifiedScore104featuredCNo
MMLUmathMMLUScore100featuredBYes
Humanity's Last ExammathHumanity's Last ExamScore92featuredBNo
LiveCodeBenchreasoningLiveCodeBenchScore73featuredCYes
MATHmathMATHScore71featuredBYes
HumanEvalreasoningHumanEvalScore66featuredBYes
IFEvalstructured outputIFEvalScore65featuredBYes
MMMU-PromultimodalMMMU-ProScore65featuredBNo
MMMUmultimodalMMMUScore63featuredBNo
BrowseCompreasoningBrowseCompScore58featuredBYes
AIME 2024mathAIME 2024Score53featuredBYes
LiveCodeBench v6reasoningLiveCodeBench v6Score53featuredBYes
nolimalong contextnolimaScore52featuredCNo
MMMLUmathMMMLUScore49featuredBYes
Terminal-Bench 2.0reasoningTerminal-Bench 2.0Score49featuredCNo
GSM8kmathGSM8kScore48featuredBYes
MMLU-ReduxmathMMLU-ReduxScore48featuredBYes
CharXiv-RmultimodalCharXiv-RScore47featuredBNo
SimpleQAreasoningSimpleQAScore46featuredBYes
SWE-Bench ProreasoningSWE-Bench ProScore44featuredBNo
MathVistamathMathVistaScore39featuredBNo
LiveBenchmathLiveBenchScore38featuredBNo
Tau2 TelecomreasoningTau2 TelecomScore35featuredBNo
ARC-CreasoningARC-CScore34featuredBYes
SuperGPQAmathSuperGPQAScore34featuredBYes
SWE-bench MultilingualreasoningSWE-bench MultilingualScore34featuredBNo
HMMT 2025mathHMMT 2025Score33featuredBYes
MBPPreasoningMBPPScore33featuredBYes
MATH-500mathMATH-500Score32featuredCYes
MMLU-ProXmathMMLU-ProXScore32featuredBYes
AI2DmultimodalAI2DScore32featuredBNo
MathVisionmathMathVisionScore32featuredBNo
MGSMmathMGSMScore31featuredBYes
ToolathlonreasoningToolathlonScore31featuredBYes
IncludegeneralIncludeScore31featuredBNo
DROPmathDROPScore30featuredBYes
MCP AtlasreasoningMCP AtlasScore30featuredBYes
Multi-ChallengereasoningMulti-ChallengeScore29featuredCYes
IFBenchinstruction followingIFBenchScore28featuredBYes
HellaSwagreasoningHellaSwagScore27featuredBYes
Arena HardreasoningArena HardScore26featuredBNo
DocVQAmultimodalDocVQAScore26featuredBNo
Finance Agent v2reasoningFinance Agent v2Score26featuredBNo
RealWorldQAspatial reasoningRealWorldQAScore26featuredBNo
Tau2 RetailreasoningTau2 RetailScore26featuredBNo
VideoMMMUmultimodalVideoMMMUScore26featuredBNo
HMMT25mathHMMT25Score25featuredBYes
TAU-bench RetailreasoningTAU-bench RetailScore25featuredBNo
Terminal-BenchreasoningTerminal-BenchScore25featuredBNo
ChartQAmultimodalChartQAScore24featuredBNo
LVBenchmultimodalLVBenchScore24featuredBNo
ScreenSpot PromultimodalScreenSpot ProScore24featuredCNo
t2-benchreasoningt2-benchScore23featuredBYes
WMT24++languageWMT24++Score23featuredBYes
ERQAreasoningERQAScore23featuredCNo
MathVista-MinimathMathVista-MiniScore23featuredBNo
PolyMATHmathPolyMATHScore23featuredBNo
TAU-bench AirlinereasoningTAU-bench AirlineScore23featuredBNo
Tau2 AirlinereasoningTau2 AirlineScore23featuredBNo
Aider-PolyglotgeneralAider-PolyglotScore22featuredBYes
WinograndereasoningWinograndeScore22featuredBYes
MMStarmultimodalMMStarScore22featuredBNo
OCRBenchimage to textOCRBenchScore22featuredBNo
OSWorld-VerifiedmultimodalOSWorld-VerifiedScore22featuredBNo
BIG-Bench HardmathBIG-Bench HardScore21featuredBYes
MRCR v2 (8-needle)reasoningMRCR v2 (8-needle)Score21featuredCYes
Multi-IFreasoningMulti-IFScore20featuredBYes
OSWorldmultimodalOSWorldScore20featuredBNo
BFCL-v3reasoningBFCL-v3Score19featuredBYes
IMO-AnswerBenchmathIMO-AnswerBenchScore19featuredBYes
DeepSWE 1.1agentsDeepSWE 1.1Score19featuredBNo
C-EvalreasoningC-EvalScore18featuredBYes
SciCodemathSciCodeScore18featuredBYes
TriviaQAreasoningTriviaQAScore18featuredBYes
TruthfulQAreasoningTruthfulQAScore18featuredBYes
CC-OCRmultimodalCC-OCRScore18featuredBNo
MMBench-V1.1multimodalMMBench-V1.1Score18featuredBNo
AIME 2026mathAIME 2026Score17featuredBYes
FrontierMathmathFrontierMathScore17featuredBYes
LongBench v2reasoningLongBench v2Score17featuredCYes
MM-MT-BenchmultimodalMM-MT-BenchScore17featuredBNo
MVBenchmultimodalMVBenchScore17featuredBNo
Terminal-Bench 2.1reasoningTerminal-Bench 2.1Score17featuredBNo
Video-MMEmultimodalVideo-MMEScore17featuredBNo
CodeForcesmathCodeForcesScore16featuredBYes
ARC-AGI v2reasoningARC-AGI v2Score16featuredBNo
Arena-Hard v2reasoningArena-Hard v2Score16featuredBNo
CharXiv-DmultimodalCharXiv-DScore16featuredBNo
Hallusion BenchreasoningHallusion BenchScore16featuredBNo
ODinWvisionODinWScore16featuredBNo
OmniDocBench 1.5multimodalOmniDocBench 1.5Score16featuredBNo
ScreenSpotmultimodalScreenSpotScore16featuredBNo
FrontierCode 1.1reasoningFrontierCode 1.1Score15featuredBYes
WritingBenchlegalWritingBenchScore15featuredBYes
AA-LCRreasoningAA-LCRScore15featuredBNo
FrontierSWEagentsFrontierSWEScore15featuredBNo
TextVQAmultimodalTextVQAScore15featuredBNo
Global-MMLU-LitereasoningGlobal-MMLU-LiteScore14featuredBYes
LiveBench 20241125mathLiveBench 20241125Score14featuredBNo
NL2RepoagentsNL2RepoScore14featuredBNo
BrowseComp-zhreasoningBrowseComp-zhScore13featuredBYes
Creative Writing v3creativityCreative Writing v3Score13featuredBYes
FACTS GroundingreasoningFACTS GroundingScore13featuredCYes
Global PIQAphysicsGlobal PIQAScore13featuredBYes
HiddenMathmathHiddenMathScore13featuredBYes
MultiPL-ElanguageMultiPL-EScore13featuredBYes
BFCL-V4agentsBFCL-V4Score13featuredBNo
BLINKmultimodalBLINKScore13featuredBNo
Claw-EvalagentsClaw-EvalScore13featuredBNo
Legal Agent BenchmarkreasoningLegal Agent BenchmarkScore13featuredBNo
SimpleVQAmultimodalSimpleVQAScore13featuredBNo
BBHmathBBHScore12featuredBYes
MT-BenchreasoningMT-BenchScore12featuredBYes
CharadesSTAmultimodalCharadesSTAScore12featuredBNo
InfoVQAtestmultimodalInfoVQAtestScore12featuredBNo
MedXpertQAmultimodalMedXpertQAScore12featuredBNo
OCRBench-V2 (en)image to textOCRBench-V2 (en)Score12featuredBNo
OptimBenchuncategorizedOptimBenchScore12featuredCNo
BFCLreasoningBFCLScore11featuredBYes
BIG-Bench Extra HardreasoningBIG-Bench Extra HardScore11featuredBYes
Graphwalks BFS <128kreasoningGraphwalks BFS <128kScore11featuredBYes
Graphwalks BFS >128kreasoningGraphwalks BFS >128kScore11featuredBYes
Graphwalks parents <128kreasoningGraphwalks parents <128kScore11featuredBYes
HMMT Feb 26mathHMMT Feb 26Score11featuredBYes
MAXIFEgeneralMAXIFEScore11featuredBYes
NOVA-63generalNOVA-63Score11featuredBYes
PIQAphysicsPIQAScore11featuredBYes
CyberGymsafetyCyberGymScore11featuredBNo
DocVQAtestmultimodalDocVQAtestScore11featuredBNo
MMMU (val)multimodalMMMU (val)Score11featuredBNo
MuirBenchmultimodalMuirBenchScore11featuredBNo
OCRBench-V2 (zh)image to textOCRBench-V2 (zh)Score11featuredBNo
AGIEvalmathAGIEvalScore10featuredBYes
Aider-Polyglot EditgeneralAider-Polyglot EditScore10featuredBYes
BoolQreasoningBoolQScore10featuredBYes
COLLIEreasoningCOLLIEScore10featuredBYes
HumanEval+reasoningHumanEval+Score10featuredBYes
VITA-BenchreasoningVITA-BenchScore10featuredBYes
DeepSWEagentsDeepSWEScore10featuredBNo
MLVUmultimodalMLVUScore10featuredBNo
VideoMME w sub.multimodalVideoMME w sub.Score10featuredBNo
VideoMME w/o sub.multimodalVideoMME w/o sub.Score10featuredBNo

Benchmark values stay on their source-native scales. Every registry row uses its unique Benchmark handle for the detail link.

Rankings

OverallCodingText ArenaPricing

Modalities

Image GenerationVideo GenerationSpeech-to-TextEmbeddings

Benchmarks

All BenchmarksReasoningMathCoding

Vendors

All VendorsOpenAIAnthropicGoogle
llmboard.aiCopyright 2026 llmboard.ai