llmboard.ai
Benchmarks
CompareRankings
llmboard.ai
Benchmarks
CompareRankings
HomeBenchmarksagentsCL-bench

agents benchmark

CL-bench

CL-bench is an open-source benchmark with its own data and rubrics for evaluating models on coding and agentic tasks, scored using a setup fully aligned with the official procedure.

Updated Aug 11, 2026

Models2
Model coverage2
MetricScore
EvidenceB

On this page

  • Ranking
  • Distribution
  • Highlights
  • About
  • FAQ

CL-bench Ranking

Higher score ranks better on this benchmark.

2 rows
Columns

Show columns

1TEHy3Tencent23.8%100.0%2CAug 11, 2026
2MIMiniMax M3MiniMax20.5%0.0%2CAug 11, 2026

CL-bench Score Distribution

A closer view of the leading scores on this benchmark.

CL-bench

CL-bench Highlights

The leading models and scores on this benchmark.

Rank #1Hy323.8%Rank #2MiniMax M320.5%

What is CL-bench?

What CL-bench measures and how its scores work.

CL-bench is an open-source benchmark with its own data and rubrics for evaluating models on coding and agentic tasks, scored using a setup fully aligned with the official procedure.

Scores are shown in ratio. This benchmark is not independently verified and has an evidence level of B.

Family
CL-bench
Modality
text
Primary category
agents
Score direction
higher
LLMBoard eligible
No
Evaluation key
cl-bench|llm-stats-current

Benchmark scores retain their original unit. Overall score eligibility is shown separately.

FAQ

Common questions about CL-bench.

Which model scores highest on CL-bench?

Hy3 is currently ranked first with 23.8%.

What does CL-bench measure?

CL-bench is an open-source benchmark with its own data and rubrics for evaluating models on coding and agentic tasks, scored using a setup fully aligned with the official procedure.

Is a higher score better?

Yes. Higher values rank better for this benchmark.

How many models are compared?

2 model results are currently shown.

Does this benchmark affect the overall score?

No. This benchmark is shown for reference but does not contribute to the overall score.

Rankings

OverallCodingText ArenaPricing

Modalities

All ModelsImage GenerationImage EditingVideo GenerationImage-to-VideoVideo EditingText-to-SpeechSpeech-to-TextEmbeddings

Benchmarks

All BenchmarksReasoningMathCoding

Vendors

All VendorsOpenAIAnthropicGoogle
llmboard.aiCopyright 2026 llmboard.ai