llmboard.ai
Benchmarks
CompareRankings
llmboard.ai
Benchmarks
CompareRankings
HomeBenchmarksAgents

Benchmark category

Agents Benchmarks

Published agents evaluations and source-native model rankings. Each benchmark keeps its original scale and methodology.

Data as of 2026-08-07

Benchmarks156
Highest coverage58
Score eligible7

Agents benchmark registry

Select a benchmark to inspect model-level results, evidence fields and scoring direction.

22 rows
Columns

Show columns

BrowseCompBrowseComptextScore58featuredBYes
Terminal-Bench 2.0Terminal-Bench 2.0textScore49featuredCNo
SWE-Bench ProSWE-Bench ProtextScore44featuredBNo
ToolathlonToolathlontextScore31featuredBYes
MCP AtlasMCP AtlastextScore30featuredBYes
Finance Agent v2Finance Agent v2textScore26featuredBNo
Terminal-BenchTerminal-BenchtextScore25featuredBNo
t2-bencht2-benchtextScore23featuredBYes
OSWorld-VerifiedOSWorld-VerifiedmultimodalScore22featuredBNo
OSWorldOSWorldmultimodalScore20featuredBNo
BFCL-v3BFCL-v3textScore19featuredBYes
DeepSWE 1.1DeepSWE 1.1textScore19featuredBNo
Terminal-Bench 2.1Terminal-Bench 2.1textScore17featuredBNo
FrontierCode 1.1FrontierCode 1.1textScore15featuredBYes
FrontierSWEFrontierSWEtextScore15featuredBNo
NL2RepoNL2RepotextScore14featuredBNo
BFCL-V4BFCL-V4textScore13featuredBNo
Claw-EvalClaw-EvaltextScore13featuredBNo
Legal Agent BenchmarkLegal Agent BenchmarktextScore13featuredBNo
CyberGymCyberGymtextScore11featuredBNo
VITA-BenchVITA-BenchtextScore10featuredBYes
DeepSWEDeepSWEtextScore10featuredBNo

Top agents result sets

High-coverage benchmarks with at least two published model results.

BrowseComp

View benchmark

Toolathlon

View benchmark

MCP Atlas

View benchmark

t2-bench

View benchmark

What are agents benchmarks?

How this category is assembled on llmboard.ai.

This page groups benchmarks whose primary or display category matches agents. It does not average incompatible metrics into a new category score.

Open an individual benchmark to inspect score direction, evidence level, participant count and source-native results.

Category membership is derived from the current benchmark registry response.

Rankings

OverallCodingText ArenaPricing

Modalities

Image GenerationVideo GenerationSpeech-to-TextEmbeddings

Benchmarks

All BenchmarksReasoningMathCoding

Vendors

All VendorsOpenAIAnthropicGoogle
llmboard.aiCopyright 2026 llmboard.ai