llmboard.ai
Benchmarks
CompareRankings
llmboard.ai
Benchmarks
CompareRankings
HomeBenchmarks

Evaluation directory

AI Benchmarks

Browse benchmark definitions, capability categories and model rankings.

Data as of 2026-08-11

Benchmarks668
Categories34
Score eligible331
With model coverage665

Benchmark highlights

A quick view of benchmark breadth and current model coverage.

Highest model coverageGPQA234 modelsLargest categoryreasoning171 benchmarks
Scoring inputs331Eligible benchmark definitions
Total benchmarks66834 primary categories

Capability categories

Open a category to compare its benchmarks and leading model results.

34 rows
Columns

Show columns

reasoning17195105
multimodal1306666
math6639114
language5534129
long context532052
agents501820
safety18911
code1514
general15931
image to text121026
knowledge1213
legal10634
instruction following8365
physics73234
spatial reasoning7626
healthcare649
productivity623
uncategorized4012
memory301
structured output317
audio201
finance201
image-generation201
3d101
creativity1113
data analysis101
factuality101
frontend development113
psychology119
question answering101
research101
robotics101
video101
vision1116

All benchmarks

Open any benchmark to inspect its leaderboard, scoring direction, evidence fields and programmatic FAQ.

145 rows
Columns

Show columns

GPQAphysicsGPQAScore234featuredCYes
MMLU-ProlanguageMMLU-ProScore129featuredBYes
AIME 2025mathAIME 2025Score114featuredCYes
SWE-Bench VerifiedreasoningSWE-Bench VerifiedScore105featuredCYes
MMLUlanguageMMLUScore100featuredBYes
Humanity's Last ExammathHumanity's Last ExamScore93featuredBYes
LiveCodeBenchreasoningLiveCodeBenchScore73featuredCYes
MATHmathMATHScore71featuredBYes
HumanEvalreasoningHumanEvalScore66featuredBYes
MMMU-PromultimodalMMMU-ProScore66featuredBYes
IFEvalinstruction followingIFEvalScore65featuredBYes
MMMUmultimodalMMMUScore63featuredBYes
BrowseCompreasoningBrowseCompScore58featuredBYes
AIME 2024mathAIME 2024Score53featuredBYes
LiveCodeBench v6reasoningLiveCodeBench v6Score53featuredBYes
nolimalong contextnolimaScore52featuredCNo
MMMLUlanguageMMMLUScore49featuredBYes
Terminal-Bench 2.0reasoningTerminal-Bench 2.0Score49featuredCYes
CharXiv-RmultimodalCharXiv-RScore48featuredBYes
GSM8kmathGSM8kScore48featuredBYes
MMLU-ReduxlanguageMMLU-ReduxScore48featuredBYes
SimpleQAreasoningSimpleQAScore46featuredBYes
SWE-Bench ProreasoningSWE-Bench ProScore45featuredBYes
MathVistamathMathVistaScore39featuredBYes
LiveBenchmathLiveBenchScore38featuredBYes
Tau2 TelecomreasoningTau2 TelecomScore35featuredBYes
ARC-CreasoningARC-CScore34featuredBYes
SuperGPQAlegalSuperGPQAScore34featuredBYes
SWE-bench MultilingualreasoningSWE-bench MultilingualScore34featuredBYes
HMMT 2025mathHMMT 2025Score33featuredBYes
MBPPreasoningMBPPScore33featuredBYes
AI2DmultimodalAI2DScore32featuredBYes
MATH-500mathMATH-500Score32featuredCYes
MathVisionmathMathVisionScore32featuredBYes
MMLU-ProXlanguageMMLU-ProXScore32featuredBYes
IncludegeneralIncludeScore31featuredBYes
MCP AtlasreasoningMCP AtlasScore31featuredBYes
MGSMmathMGSMScore31featuredBYes
ToolathlonreasoningToolathlonScore31featuredBYes
DROPmathDROPScore30featuredBYes
IFBenchinstruction followingIFBenchScore29featuredBYes
Multi-ChallengereasoningMulti-ChallengeScore29featuredCYes
HellaSwagreasoningHellaSwagScore27featuredBYes
Arena HardreasoningArena HardScore26featuredBYes
DocVQAimage to textDocVQAScore26featuredBYes
Finance Agent v2reasoningFinance Agent v2Score26featuredBYes
RealWorldQAspatial reasoningRealWorldQAScore26featuredBYes
Tau2 RetailreasoningTau2 RetailScore26featuredBYes
VideoMMMUmultimodalVideoMMMUScore26featuredBYes
HMMT25mathHMMT25Score25featuredBYes
ScreenSpot PromultimodalScreenSpot ProScore25featuredCYes
TAU-bench RetailreasoningTAU-bench RetailScore25featuredBYes
Terminal-BenchreasoningTerminal-BenchScore25featuredBYes
ChartQAmultimodalChartQAScore24featuredBYes
LVBenchlong contextLVBenchScore24featuredBYes
ERQAreasoningERQAScore23featuredCYes
MathVista-MinimathMathVista-MiniScore23featuredBYes
OSWorld-VerifiedmultimodalOSWorld-VerifiedScore23featuredBYes
PolyMATHmathPolyMATHScore23featuredBYes
t2-benchreasoningt2-benchScore23featuredBYes
TAU-bench AirlinereasoningTAU-bench AirlineScore23featuredBYes
Tau2 AirlinereasoningTau2 AirlineScore23featuredBYes
WMT24++languageWMT24++Score23featuredBYes
Aider-PolyglotgeneralAider-PolyglotScore22featuredBYes
MMStarmultimodalMMStarScore22featuredBYes
OCRBenchimage to textOCRBenchScore22featuredBYes
WinograndelanguageWinograndeScore22featuredBYes
BIG-Bench HardlanguageBIG-Bench HardScore21featuredBYes
MRCR v2 (8-needle)long contextMRCR v2 (8-needle)Score21featuredCYes
DeepSWE 1.1agentsDeepSWE 1.1Score20featuredBYes
Multi-IFinstruction followingMulti-IFScore20featuredBYes
OSWorldmultimodalOSWorldScore20featuredBYes
BFCL-v3reasoningBFCL-v3Score19featuredBYes
IMO-AnswerBenchmathIMO-AnswerBenchScore19featuredBYes
SciCodemathSciCodeScore19featuredBYes
Terminal-Bench 2.1reasoningTerminal-Bench 2.1Score19featuredBYes
AIME 2026mathAIME 2026Score18featuredBYes
C-EvalreasoningC-EvalScore18featuredBYes
CC-OCRmultimodalCC-OCRScore18featuredBYes
MMBench-V1.1multimodalMMBench-V1.1Score18featuredBYes
TriviaQAreasoningTriviaQAScore18featuredBYes
TruthfulQAlegalTruthfulQAScore18featuredBYes
FrontierMathmathFrontierMathScore17featuredBYes
LongBench v2long contextLongBench v2Score17featuredCYes
MM-MT-BenchmultimodalMM-MT-BenchScore17featuredBYes
MVBenchmultimodalMVBenchScore17featuredBYes
OmniDocBench 1.5multimodalOmniDocBench 1.5Score17featuredBYes
Video-MMEmultimodalVideo-MMEScore17featuredBYes
AA-LCRlong contextAA-LCRScore16featuredBYes
ARC-AGI v2reasoningARC-AGI v2Score16featuredBYes
Arena-Hard v2reasoningArena-Hard v2Score16featuredBYes
CharXiv-DmultimodalCharXiv-DScore16featuredBYes
CodeForcesmathCodeForcesScore16featuredBYes
Hallusion BenchreasoningHallusion BenchScore16featuredBYes
ODinWvisionODinWScore16featuredBYes
ScreenSpotmultimodalScreenSpotScore16featuredBYes
FrontierCode 1.1reasoningFrontierCode 1.1Score15featuredBYes
FrontierSWEagentsFrontierSWEScore15featuredBYes
TextVQAimage to textTextVQAScore15featuredBYes
WritingBenchlegalWritingBenchScore15featuredBYes
Global-MMLU-LitelanguageGlobal-MMLU-LiteScore14featuredBYes
LiveBench 20241125mathLiveBench 20241125Score14featuredBYes
NL2RepoagentsNL2RepoScore14featuredBYes
BFCL-V4agentsBFCL-V4Score13featuredBYes
BLINKmultimodalBLINKScore13featuredBYes
BrowseComp-zhreasoningBrowseComp-zhScore13featuredBYes
Claw-EvalagentsClaw-EvalScore13featuredBYes
Creative Writing v3creativityCreative Writing v3Score13featuredBYes
FACTS GroundingreasoningFACTS GroundingScore13featuredCYes
Global PIQAphysicsGlobal PIQAScore13featuredBYes
HiddenMathmathHiddenMathScore13featuredBYes
Legal Agent BenchmarklegalLegal Agent BenchmarkScore13featuredBYes
MultiPL-ElanguageMultiPL-EScore13featuredBYes
SimpleVQAimage to textSimpleVQAScore13featuredBYes
BBHlanguageBBHScore12featuredBYes
CharadesSTAlanguageCharadesSTAScore12featuredBYes
InfoVQAtestmultimodalInfoVQAtestScore12featuredBYes
MedXpertQAmultimodalMedXpertQAScore12featuredBYes
MT-BenchreasoningMT-BenchScore12featuredBYes
OCRBench-V2 (en)image to textOCRBench-V2 (en)Score12featuredBYes
OptimBenchuncategorizedOptimBenchScore12featuredCNo
BFCLreasoningBFCLScore11featuredBYes
BIG-Bench Extra HardlanguageBIG-Bench Extra HardScore11featuredBYes
CyberGymsafetyCyberGymScore11featuredBYes
DocVQAtestmultimodalDocVQAtestScore11featuredBYes
Graphwalks BFS <128kreasoningGraphwalks BFS <128kScore11featuredBYes
Graphwalks BFS >128klong contextGraphwalks BFS >128kScore11featuredBYes
Graphwalks parents <128kreasoningGraphwalks parents <128kScore11featuredBYes
HMMT Feb 26mathHMMT Feb 26Score11featuredBYes
MAXIFEgeneralMAXIFEScore11featuredBYes
MMMU (val)multimodalMMMU (val)Score11featuredBYes
MuirBenchmultimodalMuirBenchScore11featuredBYes
NOVA-63generalNOVA-63Score11featuredBYes
OCRBench-V2 (zh)image to textOCRBench-V2 (zh)Score11featuredBYes
PIQAphysicsPIQAScore11featuredBYes
AGIEvallegalAGIEvalScore10featuredBYes
Aider-Polyglot EditgeneralAider-Polyglot EditScore10featuredBYes
BoolQlanguageBoolQScore10featuredBYes
COLLIElanguageCOLLIEScore10featuredBYes
DeepSWEagentsDeepSWEScore10featuredBYes
HumanEval+reasoningHumanEval+Score10featuredBYes
MLVUlong contextMLVUScore10featuredBYes
VideoMME w sub.multimodalVideoMME w sub.Score10featuredBYes
VideoMME w/o sub.multimodalVideoMME w/o sub.Score10featuredBYes
VITA-BenchreasoningVITA-BenchScore10featuredBYes

Benchmark values retain their original scales. Open a row to see its model results and scoring direction.

Rankings

OverallCodingText ArenaPricing

Modalities

All ModelsImage GenerationImage EditingVideo GenerationImage-to-VideoVideo EditingText-to-SpeechSpeech-to-TextEmbeddings

Benchmarks

All BenchmarksReasoningMathCoding

Vendors

All VendorsOpenAIAnthropicGoogle
llmboard.aiCopyright 2026 llmboard.ai